onPanda is an open-source annotation interface that lets a human replace the first wrong token in a model response, then asks the model to regenerate the rest. In a 21-prompt study using Qwen3.5-35B-A3B, its median annotation time was 330 seconds, against 681 seconds for manual post-editing. That is a useful workflow result, although it does not establish that models trained on the resulting data improve.

The practical question for an alignment team is whether this technique reduces the human cost of obtaining an acceptable answer while keeping most of the answer in the model’s own distribution. The onPanda paper provides early evidence on both, with clear limits on what was tested.

The correction changes the continuation, not just the displayed answer

An annotator reads a response until its first inappropriate token. The interface shows the model’s highest-probability alternatives at that position, with 20 candidates by default. The annotator can select one or type a replacement. onPanda retains the preceding text, removes the remainder, and requests a new continuation from the corrected prefix. A single correction can therefore prevent the same mistake from propagating through later sentences.

The implementation also saves each branch of the session, including the original response, correction position, replacement and later continuation. A final accepted branch can become a supervised fine-tuning sample; earlier branches can form preference pairs. This is more specific feedback than a whole-response thumbs-up, but a collection of precise labels is not itself proof of better training outcomes.

There are technical dependencies. The inference endpoint must support continuation from an assistant-message prefix and return candidate-token log probabilities. Probability refresh for pasted or manually edited text needs prompt log probabilities as well. Agent trajectories add response-template integration so edited tool calls can be parsed and executed. A simple chat API that exposes only completed text will not support the full workflow.

What the Qwen experiment actually measured

The controlled study involved three annotators and 21 image-description prompts. All three interfaces began from rollouts by the same Qwen3.5-35B-A3B model. onPanda was compared with POTATO for manual post-editing and Argilla for ranking four candidate answers. These are complete workflow comparisons, so interface design and annotation method both contribute to the result.

WorkflowMedian time per promptMean time per promptQualified SFT answers
onPanda correction330 s516 s21 of 21
POTATO post-editing681 s711 s21 of 21
Argilla ranking336 s685 s11 of 21

The 51.5% median reduction is relative to post-editing; onPanda’s median was about the same as ranking. The denominator matters: Argilla’s annotators could choose among four candidates, but only 11 prompts yielded an answer judged suitable for supervised fine-tuning. Editing workflows continued until the answer met that bar, so their 100% coverage was built into the task design, not a surprise model capability.

The authors used perplexity under the rollout model as a proxy for how close edited answers stayed to that model’s distribution. onPanda was 0.86% above a fresh-sample baseline; post-editing was 36.31% above it. Per-prompt resampling varied by roughly 2.8%, so the onPanda difference is within that reported noise. The result supports a narrow claim about this model and these responses. It does not mean every corrected answer is literally sampled from the unmodified policy, especially where a human types a replacement the model would not have produced.

The cost case depends on the usable output

For a team paying annotators by time, median speed alone is a poor budget number: total work tracks the mean and the share of outputs usable for the intended training set. In this study, onPanda averaged about 8.6 minutes per prompt and post-editing about 11.9 minutes. Multiplying those means by 21 gives approximately 180 versus 249 annotator minutes for a hypothetical single pass at the reported averages. That arithmetic illustrates the observed time difference; it excludes setup, model inference, review and rejected data. Argilla’s 11 usable responses out of 21 make a simple seconds-per-prompt comparison especially misleading.

The paper also reports a 66.7% pairwise win rate for onPanda answers in GPT-5.5 judgments and a smaller human comparison favoring onPanda over post-editing 54.8% of the time. These are encouraging quality checks, but the prompt set is small, the study uses one rollout model, and the main quality ranking depends on an LLM judge. The authors report production annotation sessions across vision, audio and agents; they do not present a controlled agent annotation trial or a downstream model-training experiment.

For deployment, the sensible pilot is to measure accepted examples per annotator hour, actual inference spend, correction density and performance after training on the collected data. If errors are frequent rather than sparse, the repeated intervention may erase the time advantage. If data is used to train a different model, or a continually changing model, the claim of staying near the original rollout policy also weakens. Those are limitations the authors identify, and they set the boundary between a promising annotation tool and a proven alignment improvement.

Last Update: October 1, 2026