QwenASR
A fast, pure-Rust, CPU-only inference engine for Qwen3-ASR speech-to-text, tuned for Apple Silicon.
| Tech stack | Pure Rust, CPU-only — no Python, no GPU, no tensor framework, only libc |
|---|
| Optimization | Hand-written NEON / Accelerate / AMX-aware kernels for Apple Silicon |
|---|
| Performance | Transcribes a 28-second clip in 613 ms on an M5 Pro (46× realtime) |
|---|
| Features | Offline & streaming transcription, live capture with VAD, SRT/VTT subtitles, structured JSON, forced alignment |
|---|
| Install | cargo install qwen-asr-cli |
|---|
| Website | github.com/huanglizhuo/QwenASR |
|---|
QwenASR is a fast, pure-Rust, CPU-only inference engine for Qwen3-ASR speech-to-text, tuned for Apple Silicon with hand-written NEON / Accelerate / AMX-aware kernels.
It transcribes a 28-second clip in 613 ms on an M5 Pro (46× realtime), beating GPU-based MLX implementations.
Features
- Offline and streaming transcription
- Live capture with VAD
- SRT/VTT subtitles
- Structured JSON output
- Forced alignment
Install
cargo install qwen-asr-cli
FAQ
Do I need a GPU or Python to run QwenASR?
No. It is a pure-Rust, CPU-only inference engine — no Python, no GPU, no tensor framework, only libc.
How fast is it?
It transcribes a 28-second clip in 613 ms on an M5 Pro (46× realtime), beating GPU-based MLX implementations.