Kyutai’s Voice of Reason takes a spoken math question and generates an answer through a speech-native model built on GLM-4-Voice. Its two released checkpoints report 70.6% and 77.1% on GSM8K, but those figures grade the model’s decoded text stream. The published research also evaluates what listeners would hear, and the released files make “run locally” a substantial hardware project rather than a one-click laptop demo.
The useful question is where the gain comes from and which number a prospective user can expect to reproduce. The answer depends on the checkpoint, decoding settings, evaluation channel and audio pipeline. Treating the two variants as one model hides the main engineering trade-off.
One checkpoint speaks its reasoning; the other thinks between speech chunks
The direct checkpoint interleaves text and audio tokens without an additional hidden reasoning channel. Its step-by-step working is part of the answer it speaks. The STITCH checkpoint inserts unspoken reasoning chunks marked [SOPR] and [EOPR] between spoken segments. Those marked chunks are not meant to be read aloud. Kyutai’s model cards report 70.6% and 77.1%, respectively, on 1,310 GSM8K items using the written channel and a GPT-4o judge.
STITCH’s design uses time that speech playback already occupies. While one chunk plays, the model can produce reasoning tokens for the next. That can reduce the extra delay associated with thinking before speaking, but it is a scheduling claim about that pipeline, not proof that every local backend will achieve the same end-to-end latency. Audio decoding, buffering and the device running inference still matter.
A correct text answer is not automatically a correct spoken answer
The paper’s Table 2 separates scoring on the decoded text stream from scoring a transcript of the generated audio. With its specified top-k decoding, the direct SFT-and-RL model averages 65.5% on GSM8K text and 63.9% after speech transcription. The STITCH-style model averages 74.8% on text and 72.0% after transcription. That second measurement is closer to a listener’s experience, although it still depends on a separate speech-recognition model used for evaluation.
The paper reports averages over three training seeds. A footnote says removing top-k=50 raises the text scores to 70.3% and 77.1% for the released models. The current direct model card lists 70.6%, a small difference from the paper’s 70.3%. More consequentially, the direct card says its reasoning behavior was learned through reinforcement learning without a supervised fine-tuning stage, while the paper describes SFT followed by RL for its main direct model. The public descriptions do not clearly reconcile those training histories. Readers should identify the exact checkpoint revision and evaluation recipe before equating the card’s figure with the paper’s multi-seed result.
The task boundary matters too. A score on spoken grade-school math is evidence about that benchmark, not about general conversational reasoning, transcription, or production voice-assistant reliability. In the paper’s spoken TriviaQA subset, the direct model falls from the base model’s 40.6% to 34.0% after math-focused training; the STITCH-style model records 21.4%. The authors’ math improvement therefore comes with an observed general-knowledge trade-off in their evaluation.
Local use requires more than the model file
The direct model’s Hugging Face file listing shows a 19.1GB BF16 safetensors checkpoint. That is only the core weight file. The model card’s runnable recipe separately pulls the original GLM-4-Voice speech tokenizer and decoder, and says the printed script ran on one H100 with Transformers 4.47.1 and PyTorch 2.8.0. It uses CUDA and trust_remote_code=True.
That published recipe establishes a working reference environment, not a minimum consumer-GPU requirement. Available memory must cover the weights, speech components and runtime buffers. No official quantized GGUF or MLX path appears in the two released model cards, so a smaller-machine performance claim would need its own measured port. Check the inherited GLM-4-Voice license before treating the checkpoint like an unrestricted commercial component.
For an honest evaluation, record the checkpoint revision, audio tokenizer and decoder revisions, Transformers and PyTorch versions, GPU, decoding settings, and the exact scoring channel. Test both short and long spoken questions. Save generated audio as well as text so mistakes introduced during speech generation are visible. Measure time to first audio and total answer time alongside accuracy. Those observations would show whether the STITCH design pays off for the intended local setup.
Voice of Reason demonstrates a real route toward stronger speech-native math, with a meaningful gain in Kyutai’s controlled experiments. The next practical proof is a reproducible deployment report that joins the checkpoint identity, audible answers, latency and memory use on hardware developers actually have.
