AI systems are getting better at generating code, analyzing incidents, calling tools, and operating applications. But once AI moves from a chatbot into a production system, a different problem becomes apparent: most production workflows don’t need an AI to write an answer. They need an AI to make a decision.
Should a security alert be blocked or allowed? Should a support ticket go to billing or engineering? Is a cloud resource worth investigating? Should an automated agent execute a remediation step or ask for human approval? These are not primarily content-generation problems. They are decision-making problems.
For these use cases, sending every request through a large generative LLM can be excessive. The application often needs a bounded answer, a predictable structure, confidence information, and low latency. It doesn’t necessarily need several paragraphs of generated text.
That’s the problem Cloudflare is targeting with Clef AI, its new family of open-source decision models. Cloudflare introduced Clef and Clef-flash on October 1, 2026, describing them as models designed to evaluate application state against typed questions and return structured decisions. The models are available through Workers AI, while their weights have been released on Hugging Face under the Apache 2.0 license.
The interesting part isn’t simply that Cloudflare released another AI model. The more important development is the architectural idea behind it. Clef separates the task of evaluating a situation from the application logic that decides what should happen next.
That distinction could be particularly useful for AI agents, DevOps automation, security systems, FinOps workflows, and other production software where an AI model’s output needs to drive an actual action. Instead of asking an AI system to generate a recommendation and then parsing that response, a decision model can return a structured result that the surrounding application can evaluate and act upon.
What Is Clef AI?
Clef AI is Cloudflare’s open-source decision model designed to evaluate application state against a set of typed questions and return probabilities for the allowed answers. Rather than asking the model to generate an arbitrary response, developers define the decisions their application needs and specify the possible answers. The model then evaluates the available context and returns a structured result that software can consume directly.
Consider a production checkout system that has been failing for every customer for the past hour. Instead of asking an LLM to write a general analysis, the application could provide the current state and ask a set of focused questions: Is the incident urgent? Which team should handle it? How severe is the problem? A decision model could return a structured result such as urgent: yes, team: technical, severity: critical, along with probabilities for the possible answers.
The application can then use those results as part of its existing workflow. For example, a critical incident could be routed to the technical team, the on-call engineer could be notified, and an escalation workflow could be triggered automatically. The model is responsible for evaluating the situation, while the application remains responsible for deciding what actions are permitted and executing those actions.
This is fundamentally different from asking a conventional LLM, “Analyze this incident and tell me what we should do.” A generative LLM is optimized to produce language and can return a detailed explanation or recommendation. A decision model is designed to return a bounded result that software can interpret without relying on free-form text parsing.
That distinction becomes important in production systems. When an AI output is going to trigger an API call, route an incident, modify infrastructure, or start an automated workflow, predictable structure matters. Cloudflare describes decision models as systems designed to produce bounded, structured outputs that can be incorporated directly into workflows where a decision is required.
Why Did Cloudflare Build a Decision Model?
Large language models are extremely flexible, and that flexibility is one of their biggest strengths. However, it can also become a limitation for certain production workloads. An LLM can classify an incident, determine which team owns it, recommend an action, summarize the available context, generate a response, and even call external tools. That makes it useful for many different tasks, but not every application needs that level of open-ended generation.
Consider a simple classification problem where an application needs to determine which team should handle a request. The only valid answers might be billing, technical, or sales. In this situation, generating a paragraph of text doesn’t provide much additional value. The application needs a reliable classifier that can understand the available context and return one of the predefined options.
This is where decision models take a different approach. Cloudflare describes traditional LLMs as models suited to open-ended reasoning, text generation, and tool use, while decision models are designed around bounded, structured outputs that applications can consume directly.
The distinction leads to a useful architecture for production AI systems. LLMs can handle open-ended reasoning, decision models can handle bounded decisions, and application code can handle execution. Each layer has a defined responsibility rather than asking one model to perform every part of the workflow.
That doesn’t mean decision models replace LLMs. In many production systems, the two approaches can work alongside each other. An LLM might analyze logs and explain why an incident occurred, while a decision model determines whether the incident requires escalation. The application can then apply its policies and execute the appropriate workflow.
How Does a Decision Model Work?
The easiest way to understand Clef is to compare it with two approaches that developers already use: traditional deterministic rules and generative large language models. Each approach solves a different type of problem, and the right choice depends on how much flexibility and interpretation the application requires.
Traditional Rules
Traditional rules are straightforward and predictable. An application can monitor its current state and apply an explicit condition, such as checking whether CPU utilization has exceeded a defined threshold. If the condition is true, the application performs the predefined action.
Application State
↓
if CPU > 90%
↓
Scale cluster
This approach is reliable because the logic is explicitly defined by the engineering team. The limitation is that someone has to anticipate the conditions and encode them in advance. As the number of variables and possible situations increases, maintaining large collections of rules can become difficult.
Generative LLM
A generative LLM takes a much more flexible approach. The application provides its current state or context to the model, and the model generates a response. The application then has to interpret that response before deciding what to do.
Application State
↓
LLM
↓
Generated Response
↓
Interpretation
↓
Application Logic
This works well when the system needs explanation, reasoning, summarization, or other open-ended output. However, when the response needs to drive an automated action, relying on free-form text can introduce another layer of complexity. The application may need to parse the response, validate it, and determine whether the suggested action is actually permitted.
Clef-Style Decision Model
A decision model takes a different approach. Instead of asking the model to generate an unrestricted response, the application defines focused questions and the set of valid answers. The decision model evaluates the available context and returns a structured result.
Application State
↓
Focused Questions
↓
Decision Model
↓
Structured Decision
↓
Application Logic
This approach sits between deterministic programming and open-ended generation. The application still defines the decision space and controls what actions are possible, but the model can interpret complex context that would be difficult to capture with simple if/else rules.
For example, a traditional rule might determine that CPU usage above 90% should trigger scaling. A decision model can consider CPU utilization together with recent deployments, error rates, customer impact, and other application signals before deciding whether scaling is actually the appropriate response.
That makes the decision model particularly useful when the input is complex but the output needs to remain bounded and structured.
What Does “Typed Question” Mean?
One of the most important concepts behind Clef is the idea of typed questions. A question in Clef isn’t simply a natural-language prompt sent to a model with an instruction to generate an answer. Instead, each question is associated with an expected answer type, which defines the structure of the decision the application needs.
Cloudflare currently documents three question types for §: noul, choice, and score. A noul question represents a yes-or-no decision, such as whether an incident is urgent. A choice question asks the model to select one option from a predefined set, such as determining which team owns an incident. A score question is used when the application needs an ordered evaluation, such as assessing the severity of an incident.
| Type | Purpose | Example |
| noul | Yes/no decision | Is this incident urgent? |
| choice | Select one option | Which team owns this incident? |
| score | Ordered evaluation | How severe is the incident? |
Cloudflare’s current documentation allows up to 64 questions in a single request, with the answers returned using their corresponding question IDs. This allows an application to evaluate several related decisions against the same context rather than sending separate requests for every question.
For example, an incident-routing workflow might ask which team should handle a particular event. The application could define the valid choices as billing, technical, or sales. Clef then evaluates the incident context and selects from those predefined options.
Question: Which team should handle this incident?
Allowed choices:
– billing
– technical
– sales
The important part is that the model isn’t being asked to invent a fourth category or return an arbitrary sentence. The application has already defined the decision space, and the model’s role is to evaluate the available context and select an appropriate answer within that space.
This bounded output becomes particularly valuable when the result is going directly into software. Instead of parsing a paragraph such as “This appears to be a technical issue that should probably be handled by the engineering team,” the application receives a structured decision that can immediately be passed to routing logic, an automation workflow, or another controlled system.
Clef AI vs Generative LLMs
One of the biggest misconceptions could be Clef is that it is simply a smaller version of a conventional large language model. That’s not what makes Clef different. Its significance comes from both its architecture and its intended purpose.
A traditional LLM is generally optimized to generate language, reason over open-ended prompts, and produce responses that can take many forms. Clef, by contrast, is designed around a much narrower problem: evaluating application context against predefined questions and returning structured decisions.
The difference isn’t primarily about model size. It’s about what the model is designed to do. A smaller LLM could still generate text, summarize information, or answer questions. Clef is designed so that the output itself can become a decision inside an application or automation workflow.
This makes Clef better understood as a specialized decision layer rather than simply a smaller LLM.

Cloudflare reports that Clef is designed around a non-autoregressive decision step. Rather than generating intermediate tokens and then converting those tokens into a structured answer, the model scores valid schema choices directly. That architectural choice matters.
A production system doesn’t necessarily care what the model says. It cares what the model decides.
Clef vs Traditional Rules
If decision models return structured answers, a natural question is: why not just use if/else statements? The answer is that real-world application state is rarely as clean as the conditions used in a simple rule. Traditional rules work extremely well when the relationship between an input and an action is known in advance. The difficulty starts when a system needs to interpret several signals together before deciding what those signals actually mean.
Consider a production incident where CPU utilization has reached 91%, memory usage is 74%, the error rate is 18%, and a new deployment was made seven minutes ago. Customers are experiencing checkout failures in us-east-1, and the same service has already experienced three incidents recently.
A simple rule might look like this:
if cpu > 90:
scale_cluster()
The rule is deterministic, but CPU utilization alone doesn’t tell us why the system is under pressure or whether scaling is the right response. The increase could be caused by a problematic deployment rather than insufficient capacity. The checkout failures could also be related to a downstream payment provider. Scaling the cluster in either situation might increase infrastructure capacity without addressing the actual problem. A rollback could be more appropriate than scaling, or the incident might require investigation before any automated action is taken.
This is where a decision model can add another layer of intelligence. Instead of evaluating a single threshold, the system can consider the broader application state and answer focused questions such as whether the incident is primarily a capacity problem, whether the recent deployment is likely related, whether the deployment should be rolled back, whether human approval is required, and which remediation path is most appropriate.
The important distinction is that the decision model doesn’t replace the deterministic code that ultimately performs the action. The workflow can still use explicit policies, authorization checks, and predefined remediation functions. The model helps interpret the messy context before those deterministic steps execute.
A useful production pattern therefore looks like this:
Application State
↓
Decision Model
↓
Structured Decision
↓
Policy / Rules
↓
Controlled Action
This combination gives engineering teams the strengths of both approaches. The decision model can interpret complex context, while deterministic application logic remains responsible for enforcing policies and executing actions safely.
How Cloudflare Clef Works Under the Hood
Cloudflare says Clef uses Qwen as its base model and then post-trains the model specifically for decision-model workloads. The current Clef model is based on Qwen3.8-27B, while Clef-flash uses Qwen3.5-9B. The important point, however, isn’t simply which base model Clef uses. The more interesting difference is what happens during inference and how the model produces its output.
Traditional generative LLMs typically use autoregressive generation. Given an input, the model generates one token after another until it reaches the end of its response. If you asked a conventional LLM to assess an incident, it might generate an explanation such as “The incident appears to be serious because…” and eventually arrive at a conclusion such as severity = critical. The application then has to extract and validate that conclusion from the generated response.
Clef takes a different approach. Cloudflare describes its inference process as using a prefill-only pass through the Qwen backbone, followed by scoring the valid choices defined by the decision schema. Rather than generating a complete textual explanation and then extracting the answer, Clef evaluates the available options directly.
For example, if the application defines:
Question: How severe is this incident?
Allowed choices:
– low
– medium
– high
– critical
Clef can evaluate those predefined choices against the available application state and return the corresponding probabilities. It doesn’t need to generate a paragraph explaining the incident before reaching the final classification.
That leads to a useful engineering principle: don’t generate text when the application only needs a decision. If the next component in the system expects a bounded value, generating a long natural-language response first introduces unnecessary work and another parsing step.
Cloudflare describes Clef as using specialized attention routing and schema-bound scoring to connect evidence from the input state with the valid choices defined by the application’s schema. This allows the model to return structured probabilities without relying on free-form output parsing.
The result is a different inference pattern from a conventional generative workflow. Instead of asking the model to generate an answer and then asking application code to figure out what that answer means, the application defines the decision space up front and lets the model evaluate the available options.
Application State
↓
Typed Question
↓
Clef Decision Model
↓
Valid Choices Scored
↓
Structured Probabilities
↓
Application Logic
This architecture is particularly useful when model output needs to become an input to another software component. A security workflow might need allow, block, or review. A DevOps workflow might need rollback, scale, or investigate. A FinOps workflow might need optimize, notify-owner, or require-approval.
In each case, the application already knows the possible outcomes. What it needs from the model is the ability to interpret complex context and select the appropriate outcome. That’s the role Clef is designed to fill.
Clef and Probability
Probability is another important part of the decision-model approach because production systems often need more than a simple yes-or-no classification. Knowing which outcome the model considers most likely is useful, but understanding how strongly the model favors that outcome can help an application determine what should happen next.
Suppose a security system receives an event that needs to be classified. A traditional classifier might simply return malicious. A decision model can instead provide probabilities for each of the allowed outcomes:
malicious: 0.91
suspicious: 0.07
benign: 0.02
The application can then use those probabilities as an input to its own policy. For example, an organization might decide that events with a probability above 0.95 can be blocked automatically, events between 0.80 and 0.95 should be quarantined for further investigation, and anything below 0.80 should be allowed while still being logged.
> 0.95 → Automatically block
0.80–0.95 → Quarantine and investigate
< 0.80 → Allow but log
The exact thresholds would depend on the application, risk tolerance, and operational requirements. A financial system, production infrastructure platform, and internal development environment may all require very different thresholds.
This separation between model assessment and application policy is particularly important in enterprise systems. The model provides an evaluation of the available evidence, but it shouldn’t secretly determine what the organization considers safe enough to automate. Engineering and security teams should define those policies explicitly and decide when an automated action is acceptable, when additional checks are required, and when a human should be involved.
A useful architecture therefore keeps the responsibilities separate: the decision model evaluates the situation, while application code and policy controls determine what happens next.
Security Event
↓
Decision Model
↓
Probability / Assessment
↓
Enterprise Policy
↓
Block / Quarantine / Allow
This approach also makes automated systems easier to audit. Teams can inspect the model’s assessment, the threshold that was applied, and the policy that produced the final action rather than treating the model itself as the authority for the entire decision.
Why Calibration Matters
Probability is only useful when it has a meaningful relationship with reality. If a model assigns a 95% probability to an outcome, that should represent a substantially stronger level of confidence than a 55% prediction. Otherwise, the numbers may look precise without providing much practical value to the application consuming them.
This is why probability calibration matters in decision systems. Cloudflare says its Clef post-training process uses cross-entropy objectives and Brier loss to improve calibration. It also describes a reinforcement-learning approach called Reinforcement Learning for Calibrated Decisions (RLCD), which is intended to improve how reliably the model’s probabilities correspond to actual outcomes.
For enterprise automation, this distinction can have a direct impact on how much autonomy a workflow should be given. Consider an infrastructure workflow that receives a model assessment before taking a potentially disruptive action. Instead of treating every prediction as equally trustworthy, the workflow can use the model’s confidence as one input to its execution policy.
AI Decision
↓
Confidence Check
↓
Policy Threshold
↓
Action
For example, an organization could define a policy where high-confidence decisions are eligible for automatic execution, medium-confidence decisions require an approval step, and low-confidence decisions are escalated to an engineer.
High confidence
↓
Automatic execution
Medium confidence
↓
Request approval
Low confidence
↓
Escalate to human
The exact thresholds should be determined through testing and risk analysis rather than chosen arbitrarily. A workflow that restarts a non-critical development service can tolerate a different level of uncertainty than one that modifies production infrastructure, blocks customer traffic, or changes financial resources.
This creates a useful separation of responsibilities. The model provides an assessment, while the organization defines how much confidence is required before an action can be automated. That makes the resulting system more controllable than treating every AI prediction as equally reliable.
For production AI automation, this pattern also creates a natural path toward graduated autonomy. Teams can begin with human approval for most decisions, measure model performance, refine their thresholds, and gradually automate well-understood cases as confidence and operational evidence improve.
What Are Clef and Clef-flash?
Cloudflare launched two versions of its decision model: Clef and Clef-flash. They are designed for different operational requirements, allowing teams to choose between a larger model aimed at higher-precision decision workloads and a smaller model optimized for lower latency.
| Model | Parameters | Intended Use |
| Clef | 27B | Higher-precision decision workloads |
| Clef-flash | 9B | Faster, latency-sensitive workloads |
According to Cloudflare’s current Workers AI documentation, both models support a 64K-token context window. This gives them enough context capacity for workloads where the model needs to evaluate substantial application state before making a decision.
Clef also extends the decision-model concept beyond text and structured data. The current Workers AI documentation describes support for text, JSON, images, and video, with up to four embedded images per request within the documented limits.
Multimodal input makes the decision-model approach useful for situations where the evidence isn’t entirely contained in logs or structured application data. For example, a system could provide an image or screenshot to Clef, ask a bounded question about what it contains, and use the resulting decision inside an automated workflow.
The basic pattern could look like this:
Image / Visual Evidence
↓
Clef
↓
Visual Classification
↓
Structured Decision
↓
Workflow
This could be useful for visual inspection, security analysis, support operations, or infrastructure workflows where screenshots and other visual evidence form part of the incident context. Instead of generating a general description of an image, the application can define the specific decision it needs and use the model’s structured result as the next input to its workflow.
How Fast Is Cloudflare Clef?
Latency is one of the main reasons decision models are interesting for production systems. Cloudflare reports the following benchmark results across its evaluation runs:
| Model | Median latency | p95 latency |
| Clef | 209.3 ms | 238.6 ms |
| Clef-flash | 38.8 ms | 122.4 ms |
| Jev | 524.1 ms | 536.0 ms |
These are Cloudflare’s own reported measurements, so they should be treated as vendor benchmark results rather than universal production guarantees. Actual latency will depend on workload, network path, infrastructure, concurrency, and serving configuration. Cloudflare Blog
Cloudflare also reports that its Workers AI deployment can reduce network latency by running inference on GPUs distributed across its network. The architectural implication is more important than the benchmark number:
A decision model can potentially sit directly inside an application’s hot path.
What Can You Build With Clef AI?
The practical use cases are where this gets interesting.
1. Support Ticket Routing
Support ticket routing is a straightforward example of where a decision model can be useful. Consider a customer who submits the message, “My company has been charged twice this month.” Instead of asking a generative LLM to write a response or summarize the ticket, the application can provide the ticket context to a decision model and ask a small set of focused questions.
The workflow might ask whether the ticket is urgent, which team should handle it, and what category best describes the issue. The model can then return a structured result such as:
urgent: no
team: billing
category: duplicate-charge
Because the possible outcomes are already defined by the application, the result can be passed directly into the routing workflow. The ticket can be assigned to the billing team and categorized as a duplicate-charge issue without requiring the application to parse a natural-language response.
This approach also leaves room for additional business logic. For example, an organization could automatically escalate the ticket if the customer is a high-value account, require human review for unusually large duplicate charges, or trigger a separate payment investigation workflow. The decision model handles the classification, while the application remains responsible for enforcing business rules and executing the appropriate action.
2. Security Event Classification
Decision models can also be useful in security workflows where a system needs to evaluate several signals before determining how an event should be handled. Consider a suspicious API request. Instead of making a decision based on a single indicator, the application can collect a broader set of context, including IP reputation, request metadata, authentication history, the targeted endpoint, request frequency, geographic information, and previous security incidents.
That context can then be passed to Clef along with focused questions about the appropriate response. Rather than generating a long security analysis, the model evaluates the available evidence and returns a structured decision that the workflow can use.
For example, the security workflow might define the following possible actions:
Allow
Challenge
Rate-limit
Block
Escalate
The important point is that Clef isn’t responsible for the entire security decision. It provides an intelligence layer that helps interpret the available signals. The surrounding security system can then apply additional policies, authorization checks, thresholds, and controls before taking action.
A production workflow might therefore look like this:
API Request
↓
Collect Security Context
↓
Clef Decision Model
↓
Structured Assessment
↓
Security Policy
↓
Allow / Challenge / Rate-limit / Block / Escalate
This separation is important for security engineering. The model helps evaluate a potentially suspicious request, but deterministic security controls remain responsible for enforcing the final policy. That makes the decision model one component of a broader security architecture rather than an autonomous replacement for established security controls.
3. AI Agent Guardrails
One of the more interesting applications for decision models is controlling what an AI agent is allowed to do. As AI agents gain access to infrastructure tools, cloud APIs, databases, and deployment systems, the question is no longer only whether an agent can perform an action. The more important question is whether that action should be allowed in the current context.
Consider an AI agent that wants to execute terraform apply. Before allowing the command to run, a decision model can evaluate the surrounding context and answer focused questions such as whether the requested action is permitted, whether production infrastructure is affected, whether human approval is required, and whether the proposed change is consistent with the organization’s policies.
The resulting workflow could look like this:
AI Agent
↓
Decision Model
↓
Policy Evaluation
↓
Allow / Deny / Escalate
If the change targets a development environment and falls within an approved scope, the workflow might allow the agent to continue. If the same operation affects production resources or introduces a high-risk infrastructure change, the workflow could pause execution and request human approval instead.
Cloudflare specifically identifies agent guardrails as a use case for Clef, where the model can evaluate whether an agent should be permitted to perform an action before the corresponding tool call takes place.
This pattern is particularly relevant for enterprise AI systems because it separates what an agent wants to do from what the system allows it to do. The agent can remain capable of planning and proposing actions, while a dedicated decision layer evaluates those actions against context and policy before execution.
For infrastructure automation, this could become an important safety pattern. An AI agent doesn’t necessarily need unrestricted access to every tool it can call. Instead, each sensitive action can pass through a decision and policy layer before reaching the execution environment.
Decision Models in DevOps and FinOps
The decision-model approach also maps well to infrastructure and cloud cost management. Consider a FinOps workflow that runs every morning and collects information from several sources, including AWS billing data, Kubernetes utilization, resource ownership, deployment history, environment, and business criticality. Looking at any one of these signals in isolation may not be enough to determine whether a resource is actually wasteful or what action should be taken.
The goal of the workflow isn’t necessarily to ask an AI model, “Write me a report about this resource.” Instead, the system can ask a series of focused questions that lead directly to an operational decision. For example: Is this resource likely to be wasteful? Which optimization category applies? Is automatic remediation safe? Should the resource owner be notified? Does the proposed action require approval?
A decision model can evaluate the available context and classify the situation, while deterministic workflow logic handles the operational steps that follow. This keeps the model focused on interpreting the state rather than giving it unrestricted control over cloud resources.
A typical workflow could look like this:
Cloud Cost Data
↓
Decision Model
↓
Waste Classification
↓
Policy Check
↓
Owner Notification
↓
Approval
↓
Automation
↓
Verification
For example, a Kubernetes workload with consistently low utilization might be classified as a potential optimization candidate. The workflow could then check whether the workload belongs to production, whether there is an active deployment, whether an owner has approved optimization, and whether the proposed change falls within the organization’s cost-management policy. Only after those checks pass would the automation perform the change.
This architecture is closer to how reliable enterprise automation should be designed. The decision model provides contextual judgment, policy controls determine what is permitted, and deterministic automation performs the approved action. The workflow can then verify the result and record what happened, creating an auditable chain from observation to decision to action.
Where GRiPO Fits Into This Architecture
This is also where decision models connect naturally with workflow automation platforms such as GRiPO.
GRiPO is built around the idea that automation shouldn’t stop at generating an answer. A workflow should be able to collect context, execute code, call APIs, involve AI agents, apply conditions, and perform actions.
GRiPO’s workflow architecture combines visual workflows, isolated code execution, AI agents, and plugins. Gripo
A decision-model pattern could fit into that architecture like this:

For example, a Kubernetes incident could trigger a workflow.
The workflow gathers cluster metrics and recent deployment information. A decision layer determines whether the situation looks like a capacity issue, deployment regression, or another category. A subsequent agent or workflow step can investigate further, while the sandbox provides an isolated environment for scripts and CLI operations.
GRiPO’s code sandbox can run Python, Bash, JavaScript, kubectl, Terraform, Helm, and AWS CLI commands in isolated containers, with runtime credentials and execution auditing. Gripo
The plugin layer then connects the workflow to systems such as Slack, Jira, PagerDuty, AWS, or other HTTP and gRPC APIs. Gripo
That creates an important separation:
Decision → policy → execution
rather than:
AI → unrestricted action
Clef Does Not Replace LLMs
This distinction deserves emphasis: a decision model isn’t intended to replace a powerful generative model. The two approaches solve different problems, and production AI systems can often benefit from using them together rather than treating them as competing technologies.
Consider a team investigating a production incident. An LLM may be the right tool for understanding the problem because it can read a runbook, summarize large volumes of logs, reason through several possible causes, generate diagnostic commands, explain what is happening, and propose a remediation plan. These are open-ended tasks where language generation and broader reasoning are useful.
Once the investigation reaches a bounded decision, however, a decision model can take over. Instead of asking the generative model to make every operational decision, the system can pass the relevant context to a decision layer and ask a focused question such as, “Which remediation path should we take?”
The complete workflow could look like this:
LLM “What might be happening?”
↓
Decision Model “Which remediation path should we take?”
↓
Application “Execute the approved workflow.”
↓
Sandbox “Run the remediation safely.”
↓
Verification “Did the service recover?”
Each component has a clearly defined responsibility. The LLM handles investigation and reasoning, the decision model evaluates a bounded operational choice, the application applies authorization and policy, and the sandbox provides an isolated environment for executing the remediation. The verification step then confirms whether the action actually produced the expected result.
This creates a more modular AI architecture than relying on a single model for every part of the workflow. It also provides clearer boundaries around automation. An AI agent can investigate and propose actions without automatically receiving unrestricted authority to execute them.
For enterprise environments, that separation can be valuable because it allows teams to introduce automation gradually. Open-ended reasoning can remain flexible, while sensitive operational decisions can pass through structured decision and policy layers before execution.
Clef vs Rules vs Generative LLM
A useful way to think about the three approaches is this:

The strongest production architecture often isn’t choosing one.
It’s combining them.
Is Clef Really Deterministic?
This distinction is important when evaluating decision models for production systems. Clef is bounded and schema-constrained, but that doesn’t mean it becomes a deterministic rules engine. The model is still making probabilistic judgments based on the context it receives. The fact that its possible outputs are constrained doesn’t remove uncertainty from the underlying model.
That’s why production systems shouldn’t treat a Clef decision as an automatic instruction to execute an action. Engineering teams still need to consider confidence, calibration, decision thresholds, testing, fallback behavior, and human escalation. A structured output makes the result easier for software to consume, but it doesn’t eliminate the need for safeguards around that result.
A more robust architecture separates the model’s assessment from the controls that determine whether an action is permitted:
Model Decision
↓
Confidence
↓
Business Policy
↓
Authorization
↓
Action
This is very different from allowing the model to directly control an operational system:
Model Decision
↓
Execute Anything
In the first architecture, the model contributes intelligence without becoming the final authority. Business policies can define acceptable thresholds, authorization systems can verify whether the requested action is permitted, and deterministic application logic can control the actual execution.
That separation becomes especially important for infrastructure, security, financial systems, and production operations. A decision that merely categorizes a support ticket can tolerate a different level of uncertainty than one that modifies a Kubernetes cluster, changes a firewall rule, terminates a cloud resource, or deploys infrastructure to production.
The practical lesson is straightforward: a decision model should be treated as an intelligent input to a controlled system, not as a replacement for the controls around that system.
What Does Open Source Mean for Clef?
What Open Source Means for Enterprise Adoption
Cloudflare has released the weights for Clef and Clef-flash on Hugging Face under the Apache 2.0 license. The published model information identifies Clef as a 27-billion-parameter model based on Qwen3.8-27B.
For engineering teams, this provides more flexibility in how they evaluate the technology. Teams can experiment with Clef through Cloudflare’s hosted Workers AI service, while organizations with the required infrastructure and operational expertise can also investigate self-hosting the released model weights.
For enterprises, however, open source shouldn’t automatically be interpreted as production-ready for unrestricted use. The availability of model weights removes some vendor dependency, but it doesn’t remove the engineering work required to validate a model for a particular workload.
Before deploying a decision model into a production workflow, teams should evaluate its performance against representative internal data and consider factors such as hardware requirements, inference cost, latency, probability calibration, privacy, data residency, observability, model versioning, security controls, licensing requirements, and failure handling.
The evaluation should also reflect the consequences of an incorrect decision. A model that occasionally misclassifies a low-priority support ticket may be acceptable, while the same error rate could be unacceptable if the model is deciding whether to terminate a production resource or approve a security-sensitive operation.
The right question therefore isn’t simply:
“Is the model open source?”
A more useful enterprise question is:
“Does this decision model perform reliably enough for the decisions we want to automate?”
That shift in perspective keeps the evaluation focused on the actual production requirement rather than the model’s licensing or availability alone.
When Should You Use Clef AI?
Clef makes the most sense when a system has messy or context-heavy input but needs a bounded output. That is the sweet spot for decision models. The application may need to interpret logs, events, user behavior, infrastructure state, or other complex information, but the final result still needs to fit within a defined set of possible outcomes.
A decision model is a good fit when the input requires contextual interpretation and the application has a clearly defined decision space. It becomes particularly useful when latency matters, when the application benefits from probability or confidence information, or when the model’s output feeds directly into application logic. Decision models can also provide an additional intelligence layer before an automated action, allowing the workflow to evaluate context before deterministic policies and execution steps take over.
Another advantage is reducing the complexity of parsing free-form LLM responses. If the application only needs values such as allow, deny, review, or escalate, generating a paragraph and then extracting one of those values adds unnecessary processing and another potential failure point.
Typical situations where a decision model can be a strong fit include:
- The input requires contextual interpretation.
- The output has a defined set of valid choices.
- Low latency is important.
- The application needs probability or confidence information.
- The result feeds directly into application logic.
- An AI assessment is needed before an automated action.
- Parsing free-form LLM output would introduce unnecessary complexity.
There are also situations where a decision model isn’t the right tool. If the primary requirement is to write long-form content, brainstorm ideas, maintain an open-ended conversation, generate code, produce detailed explanations, or perform unrestricted reasoning, a generative model is generally better suited to the task.
The distinction can be summarized simply: use a decision model when the system needs to interpret complex context and choose from a defined set of outcomes; use a generative model when the system needs open-ended reasoning or generation. In many production architectures, the two can work together rather than requiring teams to choose only one.
A Practical Architecture for Enterprise AI Automation
For teams building production AI systems, a useful architecture is:

An LLM can sit alongside this architecture when deeper reasoning or explanation is required.
A workflow platform can orchestrate the entire process.
And a sandbox can isolate actions that involve code, CLI tools, infrastructure changes, or potentially risky operations.
This is the direction AI automation is moving toward: not one giant model doing everything, but specialized models and deterministic systems working together.
What Makes Clef Important?
The biggest significance of Clef isn’t its parameter count. The more important development is the shift in how developers can think about AI inference. For years, a common pattern for integrating an LLM into an application looked something like this:
Prompt
↓
LLM
↓
Generated Text
↓
Parse / Interpret
↓
Application
This architecture works well when the application actually needs natural-language output. But when the software needs a specific decision, generating text first and interpreting that text afterward introduces an additional layer between the model and the application.
Decision models suggest a different abstraction:
Application State
↓
Question Schema
↓
Decision Model
↓
Structured Decision
↓
Application
The difference may look small, but it changes how developers can design AI-powered systems. Instead of treating every AI request as a text-generation problem, developers can identify the specific decisions an application needs and use a model designed to answer those questions directly.
Cloudflare’s release also demonstrates that decision models are moving beyond a purely experimental concept. Clef is available through Workers AI, its model weights have been released publicly, it supports multimodal inputs, and Cloudflare is positioning the technology for applications such as agent guardrails and other decision-oriented workflows. Cloudflare also documents a decision-model API approach that is designed around structured questions and answers.
The ecosystem is still relatively young, however. Engineering teams shouldn’t assume that published benchmark results will automatically translate into the same performance, calibration, latency, or reliability in their own production environments. A decision model should be evaluated against representative workloads and tested against the consequences of incorrect decisions before it is given authority over sensitive operations.
The underlying idea is nevertheless compelling. AI doesn’t always need to generate a response. Sometimes an application simply needs a well-defined decision. Clef represents an interesting step toward treating that decision as a first-class AI workload rather than forcing every problem through a general-purpose text-generation interface.
AI doesn’t always need to generate. Sometimes it just needs to decide.
FAQ Section
What is Clef AI?
Clef AI is Cloudflare’s open-source decision model designed to turn application state and typed questions into structured decisions with probabilities. It is designed for classification and bounded decision-making rather than free-form text generation. Cloudflare Docs
What is Cloudflare Clef used for?
Cloudflare positions Clef for use cases including support triage, threat intelligence, trust and safety classification, agent guardrails, and multimodal classification. Cloudflare Docs
Is Clef AI an LLM?
Clef uses a language-model backbone, but it is specialized for decision-model workloads. Instead of generating a sequence of text tokens as its primary output, it scores valid choices within a defined schema. Cloudflare Blog
What is a decision model?
A decision model evaluates an input state against defined questions and returns bounded, structured answers. This makes its output easier for application code and automation workflows to consume directly.
What is the difference between Clef and an LLM?
A generative LLM is designed for open-ended generation and reasoning. Clef is designed for bounded decisions where the application defines the valid answer space.
Is Cloudflare Clef open source?
Cloudflare has released Clef and Clef-flash weights on Hugging Face under the Apache 2.0 license. Hugging Face
What is Clef-flash?
Clef-flash is Cloudflare’s smaller 9B decision model, designed for faster, latency-sensitive workloads. Cloudflare also offers the larger 27B Clef model. Cloudflare Docs
Can Clef process images?
Yes. Clef is multimodal and can process text, JSON, images, and video. The current Workers AI documentation supports up to four embedded images per request within its documented limits. Cloudflare Docs
Can Clef be used for AI agent guardrails?
Yes. Cloudflare specifically identifies agent guardrails as a use case, where Clef can evaluate whether an agent should take an action before the agent invokes a tool. Cloudflare Docs
Is Clef faster than a generative LLM?
Cloudflare’s benchmarks show low decision latency compared with several other decision models and reports a 2.2-second end-to-end threat-intelligence workflow for Clef versus 4.7 seconds for gpt-oss-120b in its cited comparison. These are vendor-reported results and shouldn’t be treated as universal production performance guarantees. Cloudflare Blog
Can Clef replace generative AI?
No. Clef is better viewed as a specialized decision layer. Generative models remain useful for reasoning, content generation, explanations, code generation, and other open-ended tasks.
How can decision models help DevOps automation?
They can classify incidents, determine escalation paths, evaluate remediation options, assess whether an automated action is appropriate, and route infrastructure events before deterministic workflow steps execute.
Key Takeaways
- Clef AI is a decision model, not simply another chatbot or general-purpose LLM.
- Its core job is to turn application state + typed questions → structured decisions.
- Clef supports bounded question types such as yes/no, choice, and score. Cloudflare Docs
- Cloudflare released both Clef 27B and Clef-flash 9B. Cloudflare Docs
- The models are available through Workers AI and their weights have been released under Apache 2.0. Hugging Face
- Clef uses a Qwen backbone and a specialized decision architecture that scores valid schema choices without conventional token-by-token output generation. Cloudflare Blog
- Decision models are particularly useful when inputs are complex but outputs are bounded.
- They can work alongside LLMs rather than replacing them.
- For enterprise automation, the strongest pattern is often AI decision → policy → controlled execution.
- Workflow platforms such as GRiPO can provide the surrounding orchestration, isolated execution, integrations, and automation needed to turn those decisions into production workflows.
