In this article
Cloudflare has released Clef and Clef-flash, which MarkTechPost describes as the first models trained by its Workers AI team. Instead of generating prose, the models read an input state and a schema of typed questions, then return a probability for every permitted answer.
Both models are available on Workers AI and as Apache 2.0-licensed weights on Hugging Face. Cloudflare’s internal Decision Index run measured median latency of 209.3 ms for the 27B Clef and 38.8 ms for the 9B Clef-flash. The intended applications include routing, triage, agent guardrails and abuse detection, where software can act on probabilities without parsing generated text.
Typed outputs instead of token streams
Clef supports three question types. A noul question represents yes or no and returns the probability of yes. A choice question selects among named options, returning per-option probabilities and a confidence value. A score question applies an ordered rubric and produces a probability-weighted score.
MarkTechPost reports that one Workers AI request can contain up to 64 questions and four images. Both models have a 65,536-token context window and retain their backbones’ vision encoders, allowing the input state to include text, JSON, images or video.
The models implement the System One API used by TypeSafe AI’s Jev, which launched on September 15, 2026. According to MarkTechPost, migrating from Jev mainly requires changing the endpoint and model name.
Clef is post-trained from Qwen3.8-27B, while Clef-flash uses Qwen3.5-9B. During inference, the frozen backbone performs one prefill-only pass over the state and questions. A small transformer called the joint schema head reads the final hidden states, routes evidence to each question, permits cross-attention between fields and scores the available options. Per-question softmax then converts the logits into probabilities.
MarkTechPost says Cloudflare kept both backbones frozen while optimizing the routing head with rank-256 low-rank adapters. Training combined label-smoothed cross-entropy with Brier loss for calibration. A secondary objective called Reinforcement Learning for Calibrated Decisions gives partial credit to adjacent ordinal choices.
Vendor-reported benchmarks
Cloudflare reports that a Clef model led seven of the 10 benchmarks in its Decision Index 0.2.1 shortlist. Clef scored 94.20 macro-F1 on BANKING77 against Jev’s 79.74, and 97.43 on CLINC150+OOS against 89.27. Clef-flash reached 97.73 case-exact accuracy on the home-appliances test, compared with 52.27 for Jev.
Jev led the reported knowledge-heavy tests: 78.3 versus 48.0 on GPQA Diamond, 82.7 versus 65.9 on MMLU-Pro and 92.9 versus 73.7 on BBH. On TypeSafe AI’s workflow evaluations, as reported by MarkTechPost, Clef led three of four categories: invoice processing at 64.7 versus 61.8, customer service at 76.3 versus 76.0, and security incidents at 62.9 versus 61.7. Jev led agent-trace observability at 71.6 versus 68.5.
In Cloudflare’s threat-intelligence workflow, the company says Clef classified a domain in 2.2 seconds, compared with 4.7 seconds for gpt-oss-120b. These benchmark and performance results are vendor-reported and have not been independently replicated.
| Feature | Clef | Clef-flash | Jev | Kev-9B | Laya |
|---|---|---|---|---|---|
| Developer | Cloudflare | Cloudflare | TypeSafe AI | Jared Palmer | Convai Innovations |
| Size | 27B | 9B | Not disclosed | 9B + 45.4M LoRA | 421M |
| Backbone | Qwen3.8-27B | Qwen3.5-9B | Not disclosed | Qwen3.5-9B-Base | ModernBERT-large |
| Image input | Yes | Yes | No | No | No |
| Context | 65,536 | 65,536 | 32K, per Cloudflare | 65,536; 8,192 validated | 512, English |
| Median latency* | 209.3 ms | 38.8 ms | 524.1 ms | 51.4 ms | 5.8 ms |
| Hosted input price** | $0.24/M | $0.09/M | $0.042/M | Self-hosted | Self-hosted |
| Weights | Apache 2.0 | Apache 2.0 | Hosted API | Apache 2.0 | Apache 2.0 |
*Cloudflare’s internal Decision Index run. **Hosted prices reported in MarkTechPost’s comparison. MarkTechPost says all five models implement the System One API.
AI Mastery analysis
Removing autoregressive generation makes execution depend on the input, schema and option set rather than the length of a generated answer. That should make decision pipelines easier to constrain, but it shifts responsibility to schema design: a fixed output cannot introduce an unanticipated category or explain a low-confidence result unless the surrounding application provides another path.
This is the architectural trade-off explored in Constrained Inference Beats Generation on Speed, Cost, and Predictability. Clef’s weaker reported results on GPQA Diamond, MMLU-Pro and BBH also caution against treating a specialised decision model as a general chatbot replacement.
Self-hosting requires more than loading weights. The joint schema head, custom chat template and image and video placeholder tokens form part of the inference contract. MarkTechPost reports that the model cards list BF16 testing on a single NVIDIA H200.
Cloudflare has also announced a reinforcement-learning service for tuning Clef on private data. It will begin with the company’s forward-deployed engineers, followed by a planned self-service platform combining AI Gateway, Workers AI, Containers and a Trainer component.
Sources
Frequently asked questions
What is the difference between Clef and Clef-flash?
Clef is a 27B model based on Qwen3.8-27B, while Clef-flash is a 9B model based on Qwen3.5-9B. Cloudflare’s internal Decision Index run measured median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash.
What outputs do Cloudflare Clef models return?
Clef returns probabilities for typed `noul`, `choice`, and `score` questions rather than free-form text. A Workers AI request can contain up to 64 questions and four images.
Can Clef and Clef-flash be self-hosted?
Yes. Both models have Apache 2.0-licensed weights on Hugging Face, and MarkTechPost reports that their model cards list testing with BF16 weights on a single NVIDIA H200.
How much do Clef and Clef-flash cost per million input tokens?
MarkTechPost’s comparison lists hosted input prices of $0.24 per million tokens for Clef and $0.09 per million for Clef-flash. It lists Jev at $0.042 per million input tokens.
How large is the Clef context window?
Clef and Clef-flash each support a 65,536-token context window. They retain their backbones’ vision encoders and can process text, JSON, images, and video.
Related Reading
Perplexity's 9B Contextual Embedder Beats voyage-context-4 by 14.4 Points
Perplexity releases pplx-embed-v2-context-9b-preview, an open-weights 9B RAG embedder that scores 45.5% Answer Recall@10 and cuts storage 8× vs float32.
Saaras V4 Covers All 22 Indian Languages With a 3B Hybrid Decoder
Sarvam AI's Saaras V4 handles all 22 scheduled Indian languages plus global English via a 3B hybrid state-space decoder, five output modes, and keyterm prompting at ₹30/hour.
ThinkingCap-Qwen3.8-27B Cuts Thinking Tokens 37.2% for 0.86pp Accuracy
BottleCap AI's fine-tune of Qwen3.8-27B drops thinking tokens 37.2% across 12 benchmarks, losing just 0.86pp of macro accuracy at xhigh effort.