Skip to content
Huang's Blog

QwenASR

A fast, pure-Rust, CPU-only inference engine for Qwen3-ASR speech-to-text, tuned for Apple Silicon.

Tech stackPure Rust, CPU-only — no Python, no GPU, no tensor framework, only libc
OptimizationHand-written NEON / Accelerate / AMX-aware kernels for Apple Silicon
PerformanceTranscribes a 28-second clip in 613 ms on an M5 Pro (46× realtime)
FeaturesOffline & streaming transcription, live capture with VAD, SRT/VTT subtitles, structured JSON, forced alignment
Installcargo install qwen-asr-cli
Websitegithub.com/huanglizhuo/QwenASR

QwenASR is a fast, pure-Rust, CPU-only inference engine for Qwen3-ASR speech-to-text, tuned for Apple Silicon with hand-written NEON / Accelerate / AMX-aware kernels.

Performance

It transcribes a 28-second clip in 613 ms on an M5 Pro (46× realtime), beating GPU-based MLX implementations.

Features

  • Offline and streaming transcription
  • Live capture with VAD
  • SRT/VTT subtitles
  • Structured JSON output
  • Forced alignment

Install

cargo install qwen-asr-cli

FAQ

Do I need a GPU or Python to run QwenASR?

No. It is a pure-Rust, CPU-only inference engine — no Python, no GPU, no tensor framework, only libc.

How fast is it?

It transcribes a 28-second clip in 613 ms on an M5 Pro (46× realtime), beating GPU-based MLX implementations.