Cloudflare released the Clef and Clef-flash decision models on Workers AI on October 1, 2026. Their weights are also available on Hugging Face under the Apache 2.0 license. Unlike a general large language model that generates open-ended text or tool calls, Clef is designed to return predefined, typed choices with probabilities. Source
Use it for classification and routing
A team can submit a support message and ask whether it is urgent, which team should handle it and how severe it is. Code can then route the request, escalate it or defer to a person when confidence is insufficient. This makes a decision model relevant when a workflow repeatedly classifies inputs within known choices, rather than producing new explanatory text. Source
The models accept images as well as text, have a 64,000-token context window and are compatible with the Jev API. Cloudflare published its own accuracy and latency benchmarks. Teams should still test representative data and consider the cost of incorrect classifications before putting a model into an automated path. Source
Choose a model and check availability
Cloudflare positions Clef for decisions that prioritize precision and Clef-flash for latency-sensitive work. Both can be called through Workers AI now, while the published weights allow local testing. Selection should account for response time, the impact of a wrong decision and the point at which a person reviews the result. Source
Cloudflare says it does not read, store or train on hosted requests or responses unless a customer uses fine-tuning. Workload-specific fine-tuning is currently offered as a hands-on service with Cloudflare engineers. A self-service platform for training and redeployment is planned later; the announcement does not specify its release date or price. Source