Paradigma’s Limite 1B Violetto is an open-weight, one-billion-parameter model built for competition-style mathematics. The publisher reports 94.01% on AIME 2026 and 74.25% on BeyondAIME, but also documents failures on ordinary instructions. For a developer, the useful question is how to run this specialized solver faithfully and where to avoid using it.

The weights and model card are available under Apache-2.0, along with a custom vLLM plugin. The model was trained from scratch on fewer than 300 billion curated tokens, then post-trained with supervised data and reinforcement learning. It supports sequences up to 131K tokens, although a long advertised context is not a promise of low memory use or stable quality at every length.

The benchmark pattern defines the use case

Paradigma’s published evaluation table puts Violetto at 94.01% on AIME 2026, 83.62% on HMMT February 2026 and 74.25% on BeyondAIME. On ArXivMath May 2026, its listed score is 25.08%, below the table’s Qwen3.5-4B at 26.33% and Qwen3.5-9B at 31.64%. These tasks and scoring formats differ, and several comparison-model results were copied from external cards rather than rerun under one controlled harness. The table supports a strong result in certain competition formats, not a claim that one billion parameters beat larger models across mathematics.

The developer’s own examples give a clearer boundary. Asked about photosynthesis, the model answered with a cell-division problem; asked about Earth’s seasons, it produced calendar arithmetic. Paradigma says Violetto can reinterpret a prompt as a different mathematical task and should not be expected to follow general-assistant instructions. An application that routes every question to it would need a separate, reliable gate to decide whether the input is actually a suitable math problem, and a way to reject off-scope answers.

Reproduction depends on the prompt and inference path

The model card instructs users to start each new problem with the canonical system prompt and the current user message, without a long conversation history. Its recommended sampling is temperature 0.6 and top-p 0.95. A generic chat benchmark that carries prior turns, overrides sampling defaults or asks for a personality-style response measures a different setup. For a comparison, record the exact prompt template, sampling settings, answer extraction and whether the score is from one try or multiple samples.

For serving, the publisher’s locked quickstart targets Linux x86-64 with an NVIDIA GPU, Python 3.12, vLLM 0.26.0, PyTorch 2.11.0 and CUDA 13.0-compatible drivers. The Hugging Face Transformers path uses custom architecture code with trust_remote_code=True; Paradigma says it validated that path on an H100. This does not imply that an H100 is the minimum hardware, only that other combinations have not received the same published qualification.

One unusual integration detail is spelled out in the model card: SDPA works with DynamicCache and supported StaticCache variants, while FlashAttention 2 with StaticCache is rejected because that pairing produces incorrect logits for Violetto’s hybrid attention layout. FlashAttention 2 with DynamicCache is supported. A runtime that silently changes cache or attention settings can therefore make an apparently valid local comparison misleading.

At one billion BF16 parameters, raw weights are on the order of two gigabytes; actual memory must also hold runtime state and the growing context cache. The publisher has not supplied a comparable tokens-per-second or concurrency sweep in its release materials, so “high-throughput” remains a design goal rather than a verified speed advantage on a named consumer GPU. Teams considering many simultaneous math problems should measure throughput, answer quality and cost on their own prompt distribution.

Violetto is promising as a narrowly routed mathematical solver with inspectable weights and serving code. Its strongest reported scores, documented off-scope behavior and precise runtime constraints all point to the same operating rule: evaluate it on the math tasks and inference configuration it was built for, then put an independent check between its answer and any consequential use.

Last Update: October 3, 2026