AI & Machine Learning

Cloudflare Releases Clef and Clef-flash Open Decision Models

Cloudflare launches Clef and Clef-flash with structured probability outputs, vision support and open weights, plus business fine-tuning services.

By Daniel Mercer Edited by Maria Konash Published: Updated:
Cloudflare Releases Clef and Clef-flash Open Decision Models
Cloudflare has released Clef and Clef-flash to make structured decisions for automated workflows. Photo: HaeB / Wikimedia Commons

Key Notes

  • Clef and Clef-flash return probabilities for predefined answers.
  • Cloudflare reports median latency of 38.8 milliseconds for Clef-flash and 209.3 milliseconds for Clef.
  • Open weights are available now, with self-service fine-tuning planned for later.

Cloudflare has released Clef and Clef-flash, two open decision models designed to help software choose between predefined options and return probability scores. The company announced the models on October 1 alongside a fine-tuning service for businesses, with a self-service training platform planned for later.

The release targets decisions inside automated workflows, such as assessing whether a support request is urgent and identifying the team that should handle it. Developers define the available choices, then use the model’s output to route work, trigger another step or send an uncertain case to a person.

Structured Decisions Instead of Generated Answers

According to its model card, Clef is a 27-billion-parameter multimodal model that processes a description of the situation and a schema of questions. It returns probabilities for the permitted answers in a single forward pass, without generating a free-form response that software must parse afterward.

The smaller Clef-flash has 9 billion parameters. Both releases support text, JSON, images and video as inputs. Their question formats cover named choices, ordered scores and true-or-false judgments, allowing one request to assess several aspects of a situation.

For a support system, that could mean evaluating department, urgency and severity together. The practical benefit is an output structure the application already understands. It still needs rules for what happens next: a confident classification can trigger routing, while a borderline result can prompt a review rather than an automatic escalation.

Vision Support and a Larger Context Window

Cloudflare’s documentation lists a 65,536-token context window for Clef and confirms vision support. The capacity is measured in tokens, rather than kilobytes. A larger input window can accommodate more documents or workflow state, though the amount of material that fits depends on its format and tokenization.

The launch follows Typesafe AI’s Jev release, which introduced its System One approach to fast, structured decisions. Typesafe emphasizes predefined outputs accompanied by confidence estimates. Cloudflare says its models are compatible with the Jev API and contrasts Clef’s vision capabilities and 64k context with Jev’s text classification and 32k window.

Compatibility may make experimentation easier for developers already using that interface, but it does not make the models interchangeable in accuracy. Teams would still need to compare results on their own inputs, especially when a classification leads to a financial, security or customer-service action.

Fast Responses, With Benchmark Trade-Offs

Cloudflare reports median latency of 38.8 milliseconds for Clef-flash and 209.3 milliseconds for Clef across its evaluation suite. Those are company-reported benchmark figures, rather than guaranteed response times for every deployment. Network conditions, input size and the rest of an application can affect the time a user experiences.

The published evaluations also show that the smaller model is not uniformly better or worse than Clef. For example, the model cards report stronger Clef-flash results on the home-appliance task, while Clef scores higher on banking-intent classification. The choice therefore depends on the workload as well as the desired balance between latency and quality.

Open Weights and Business Fine-Tuning

The weights are available on Hugging Face under the Apache 2.0 license, and hosted access is available through Workers AI. Cloudflare is initially offering workload-specific tuning through its engineering team. It plans to follow that service with a platform where customers can train and redeploy models themselves.

The release expands Cloudflare’s tooling for automated software. AIstify previously covered its Kitesurf browser for agents and its programmable wallets. Clef adds a component for selecting actions from an application’s allowed choices. Its usefulness will depend on whether those selections remain reliable on real business data, including ambiguous cases that warrant human judgment.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Enterprise Tech, News