The News, Explained
Cloudflare released Clef-omni on October 9, 2026, a decision model that processes text, images, audio, and video together. A decision model differs from a general large language model that writes open-ended answers: it scores predefined options for classification or routing. One example is assigning a customer video with sound to preset fields such as urgency and the team that should handle it. Source
Clef-omni accepts WAV or MP3 audio and MP4 or WebM video alongside text and images in one call. It is designed to score synchronized visual and audio information in the same pipeline instead of first transcribing speech or analyzing the sound and frames with separate models. Cloudflare also published the model weights on Hugging Face. Source
Cloudflare says it built Clef-omni on the comprehension components of Qwen3-Omni-30B-A3B-Instruct and removed the speech-output components. The model scores permitted options rather than generating prose. At launch, Workers AI charges $0.15 per million input tokens. Cloudflare reports median response times of about 130 milliseconds for text, 150 milliseconds for images, a few hundred milliseconds for audio, and about 1.5 seconds for a 21-second video with sound. These are measurements from Cloudflare’s hosted environment. Source
The existing models changed as well. Hosted Clef-flash input pricing fell from $0.09 to $0.038 per million tokens. Its hosted context window—the amount of input one request can process—was reduced from 64,000 to 24,000 tokens. Cloudflare says the published weights are unchanged and were trained for a 256,000-token context when self-hosted. Standard Clef keeps its $0.24-per-million-input-token price and 64,000-token context window. Source
Hosted Clef also became faster. Across Cloudflare’s three published input-size examples, median processing time improved by roughly 1.7 to 2 times. Cloudflare attributes the gain to serving infrastructure work, including a move to SGLang, rather than new model weights. In the published benchmark tables, Clef-omni, Clef, and Clef-flash lead different tasks; no single variant has the highest score in every row. Source
OYOPICK’s Take
This update makes model selection a question of input type, input length, and the cost of a wrong decision rather than choosing the newest model by default. Clef-omni can remove integration steps when video and audio need to be classified together. Clef-flash’s lower price may fit text-heavy requests that stay within 24,000 tokens, while hosted Clef remains the option to compare when a longer context is required.
Disclosing the shorter context window alongside the price cut helps buyers see the trade-off. Cloudflare says only 0.24% of its observed requests exceeded 24,000 input tokens, but that distribution does not establish the pattern for another service. A deployment should measure its own request lengths, media sizes, and the consequences of misclassification before setting a routing rule.
A model that evaluates several media types together points toward AI systems that can work with a broader slice of real-world signals. We hope that capability develops with escalation paths that hand low-confidence or high-impact decisions to people, rather than treating automation as the goal by itself. Cloudflare’s latency and benchmark figures are vendor measurements, so accuracy, response time, and cost still need testing on representative workloads.