bytevyte
bytevyte
Language

Cloudflare Clef Decision Models Target the Costliest Step in Agent Workflows

Cloudflare Clef decision models return typed probabilities instead of prose at 38.8 ms, and input-only pricing could cut the cost of every agent branch.

Cloudflare Clef decision models

Cloudflare has shipped its first in-house trained models, and neither of them writes a sentence. The Cloudflare Clef decision models, a 27B-parameter system named Clef and a 9B sibling called Clef-flash, accept a state (text, JSON, images or video) alongside a list of typed questions, then return a probability for every permitted answer. Both went live on Workers AI on October 1, 2026, with weights published under Apache 2.0.

The difference from a chat model is structural rather than cosmetic. An LLM asked to route a support ticket generates reasoning tokens, then prose, then JSON that a downstream parser has to validate. Clef returns the decision itself, with a confidence score attached. Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, and a macro-F1 score of 97.4 on CLINC150+OOS for the larger model against 66.8 for the smaller one.

Cloudflare's pitch is that routine branching does not need a frontier model at all. If that holds outside the benchmark, the most repetitive step of an agent workflow becomes a small schema-bound classifier, and the unit economics of running agents at volume shift away from frontier-lab token demand.

What Cloudflare Clef Decision Models Replace

Clef targets three jobs inside an agent loop: classify an input, route it to a branch, or escalate it. There is no free-form output to parse and no reasoning tokens to wait for. The models are Jev-API compatible, which matters because TypeSafe's Jev set the reference point for this class roughly a month before Clef arrived, as a hosted API with proprietary weights.

The commercial design follows from the output shape. Cloudflare bills input tokens only, at $0.24 per million for Clef and $0.09 per million for Clef-flash, because there is no generated text to charge for. A frontier call that emits 200 reasoning tokens plus a JSON object bills on both sides of the meter. On a decision-heavy workload, the output half of that meter disappears.

Self-hosting is available as an alternative to the hosted endpoint. Cloudflare's AI Platform group product manager Michelle Chen put the floor at 41 GB of GPU VRAM for Clef-flash and 85 GB for Clef at single concurrency and 64k context. The weights are open under Apache 2.0 on Hugging Face, though the training datasets remain private, so the release is reproducible in deployment but not in training.

ModelParametersMedian latencyCLINC150+OOS macro-F1Input priceSelf-host VRAM
Clef27B209.3 ms97.4$0.24 / M tokens85 GB
Clef-flash9B38.8 ms66.8$0.09 / M tokens41 GB

The Trade-Offs Teams Will Have to Price

The 30.6-point F1 gap between the two models is the first hard choice. Clef-flash answers in under 40 milliseconds but misclassifies far more often on the CLINC150+OOS set, so the cheap tier is not a drop-in substitute for the expensive one on judgment-heavy tasks. Teams will either route by difficulty, sending ambiguous inputs to Clef and clear ones to flash, or accept a higher error rate and absorb the cost of a wrong branch further down the pipeline.

Latency compounds in a way a single benchmark number hides. An agent that branches twenty times inside one task waits roughly 4.2 seconds at Clef's 209.3 ms median and about 0.8 seconds at Clef-flash's 38.8 ms. That gap separates an interactive workflow from a batch one, and it explains why the smaller model exists at all despite its accuracy penalty.

Open weights reopen the build-versus-buy question for one layer of the agent stack, but the hardware requirement narrows who can take that path. A single 80 GB accelerator can host Clef-flash; running Clef at 64k context needs more memory than one such device offers. For most teams that makes the hosted endpoint the default for the Cloudflare Clef decision models, with self-hosting a niche option for regulated or air-gapped deployments.

Multimodal input extends the same pattern beyond text routing. Clef accepts images and video alongside text and JSON, so a classification task that would otherwise need a vision model plus a separate decision step can collapse into one call. Cloudflare's own threat team pointed the models at website domains, running a fetch, render and classify cycle as a single pass. The pattern generalizes to any pipeline where the branch is chosen by what a model sees rather than by what it writes.

Distribution is where Cloudflare's position differs from a pure model release. Clef runs on the same edge network that already serves the Workers executing the agent, which is why a 38.8 ms median is plausible in production rather than only on a benchmark rig. Amazon shipped Strands Decider 2B on the same day, and Ollama added a /v1/systemone endpoint for the same model class on September 28. All of them make the same argument: an agent that asks a 27B model to write JSON is paying for capability it never uses.

Where the Token Bill Moves

Frontier labs sell reasoning tokens. Every routing step moved to a classifier removes an output-token stream from the meter, and agent loops generate those steps constantly. The compression lands unevenly. Branching calls are numerous and individually small; the genuinely hard reasoning is rarer and longer. So the hit falls on request volume more than on the long-context work that carries the highest margin per call.

Volume is still where agent invoices accumulate. A support automation executing a million decisions a day looks very different when each decision costs input tokens only, at $0.09 per million, rather than a full reasoning round trip. Jev-API compatibility compounds the effect: harnesses written against Jev can swap in Clef without redesign, which lowers the cost of trying a rival and turns the decision layer into a component that gets priced down rather than a moat that holds.

The evaluation burden is the hidden line item. A team replacing frontier calls with Clef-flash has to measure the cost of a misroute, not only the cost of a token: a blocked legitimate login or a misrouted refund request costs far more than the fraction of a cent saved on the classification. That calculation, rather than the benchmark table, decides whether the cheap tier is usable, and it differs from one workflow to the next.

The limit is scope. A schema-bound model answers only the questions it was given, so unusual inputs need an escalation path to a frontier model or to a human. The frontier model keeps the tail of hard cases; the classifier takes the volume. Cloudflare also shipped a reinforcement-learning fine-tuning platform alongside the models, which gives teams a way to close the accuracy gap on their own labeled data, at the cost of building an evaluation pipeline to prove the fine-tune did not regress.

The next signal to watch is whether that price holds. Clef-flash at $0.09 per million input tokens sits far below a frontier reasoning call, but Strands Decider 2B and the Ollama systemone endpoint put the same argument to the same developers, and Apache 2.0 weights remove the technical lock-in that would otherwise protect the number. Independent head-to-head routing accuracy against Jev, outside each vendor's own benchmark set, is what will decide how much of the agent loop actually migrates.

Why this matters

The release matters less as a benchmark win than as a pricing argument. Cloudflare is claiming that a large slice of agent spend buys capability the workflow never exercises, and it has attached open weights and an input-only price to make that claim testable. If developers accept it, the decision layer of the agent stack becomes cheap infrastructure, and the first tokens to leave the frontier-lab meter are the ones carrying the least premium.

Sources

Introducing Clef: our open-source decision models, and new RL fine-tuning platform | Cloudflare Blog

Clef: Open-source decision model — Cloudflare

Introducing Clef: Cloudflare's first open-source decision models, now on Workers AI · Changelog

AI-generated image.

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.