Cloudflare has released Clef, a 27B decision model, and a smaller 9B Clef-flash. Given a state and typed questions, they return probabilities over allowed answers instead of generating prose. Both can read images and are available through Workers AI; their model weights are also published under Apache-2.0. The practical choice is whether a bounded decision needs that multimodal input and whether the hosted route’s limits fit the workflow.
Clef addresses a familiar agent task: classify a support request, choose a team and score urgency without asking a general chat model to write and then parse JSON. This makes it relevant to existing Jev-compatible local decision workflows, although compatibility of a request schema does not establish equal output quality or latency.
One prefill, then schema-bound scores
Cloudflare’s model card says Clef is post-trained from Qwen3.8-27B, including its vision encoder. A joint schema head reads the backbone’s final hidden states, routes evidence to each question and assigns a logit to every allowed option. Softmax converts each question’s logits into probabilities. Cloudflare says Clef-flash uses Qwen3.5-9B. The decision stage scores options together rather than emitting a token sequence one piece at a time.
Inputs can be text or structured state, with typed questions such as yes/no, named choices and ordered scores. The released code includes a System One-compatible interface, and the Workers AI API reference accepts up to 64 questions in one request. That is useful for a multi-field triage record, but correlated mistakes can affect several fields at once. A confidence value should be calibrated on the organization’s own labeled cases before it triggers an irreversible action.
Published benchmark wins need a matched latency test
Cloudflare’s launch comparison reports Clef above Jev on several classification and routing tasks: BANKING77 macro-F1 is 94.20 versus 79.74, and BFCL case-exact accuracy is 98.47 versus 95.75. Its table also shows Jev ahead on When2Call accuracy, 80.97 versus Clef’s 72.37, and BRIGHT ranking, 47.52 versus 45.91. The useful conclusion is task-dependent rather than a universal model ranking; these are Cloudflare-run evaluations, not an independent reproduction of Clef.
The same blog lists median latency of 209.3 milliseconds for Clef, 38.8 for Clef-flash and 524.1 for Jev. Those figures should not be read as a controlled model-speed ratio. The models can run on different serving stacks and the Jev number includes a remote hosted route in the comparison. A team deciding what to put in a live path should measure the same request sizes, network location, concurrency and percentile latency through the route it plans to use. The published benchmark suite and model files make the quality comparison inspectable, but an identical deployment comparison is still a separate exercise.
The hosted multimodal route is narrower than the weights
The Workers AI page for Clef lists a 65,536-token context window and $0.24 per million input tokens. Clef-flash is listed at $0.09 per million. At 100,000 requests of 400 input tokens each, the listed input charges would be about $9.60 for Clef or $3.60 for Flash, before any other Cloudflare charges or free allocation. That arithmetic does not price local GPU ownership or operational labor.
The released Clef model card describes image and video records. The current hosted API schema documents up to four embedded PNG, JPEG or WebP images, with a 4 MiB and 16-megapixel limit per image and 8 MiB total decoded image data. It does not document a video field for the Workers AI request. Teams needing video should distinguish what the weights and local code can process from what the managed endpoint currently exposes. Cloudflare’s local example was tested on a single H200, so “open weights” also does not imply a cheap laptop deployment for the 27B version.
Clef is a substantive new option when a workflow needs several typed decisions over the same text or image state. A sound trial would label examples from the actual workload, compare Clef and Flash against the incumbent using one schema, and log accuracy, calibration, rejected inputs, cost and end-to-end latency. The published release provides enough code and documentation to start that evaluation without mistaking a vendor’s benchmark table for the result on the reader’s own data.
