N
Jun 18, 2026NeMo out, GGUF in: how parakeet.cpp ports NVIDIA ASR to C++
parakeet.cpp v0.1.0 ports NVIDIA's 6.05% WER ASR to C++/GGUF. Build from source, pull weights, transcribe — no NeMo needed.
Jun 18, 20266 min read0 reactions0 comments
Tag archive
parakeet.cpp v0.1.0 ports NVIDIA's 6.05% WER ASR to C++/GGUF. Build from source, pull weights, transcribe — no NeMo needed.

Building llama.cpp from source on a RISC-V board and running a local LLM inference server with TinyLlama 1.1B, achieving 8.5 tokens/second.