A field guide to decision AI

Decision models.And the toolsaround them.

Not every AI task needs a paragraph. Explore models that return typed decisions, and the runtimes and harnesses that bring them into your software.

From context to decisionSYSTEM ONE
INPUTContext + a question + possible answers
Train for the taskDecision model

Weights trained for typed decisions.

Use a base modelDecision runtime

Inference techniques over existing weights.

Structured output, not a conversation

choicescoreboolean
A conceptual map. Supported types and probability quality vary by implementation.
Curated models & toolsTechnical contextPractical guides
Find the right approach

What is a decision model?

Decision models are AI models trained to return structured decisions—such as a selected option, a score or a boolean—rather than free-form prose. Often called System One models, they can classify, route and score inputs within a larger software workflow.

Typed output constrains the answer's format. It does not guarantee a correct decision, calibrated confidence or better performance on your task.

Understand System One

Trained to make a call.

Browse the model catalog

One category. Different implementations.

Explore hosted and open approaches. Compare the model, access requirements and evidence—not just the headline benchmark.

KevExplore an open decision model
LayaRead the model profile and sources

Runtimes& harnesses.

There is more than one way to get a typed decision. Runtimes add inference-time techniques to existing models. Wrappers make model APIs easier to use. Specialists address narrower tasks.

Explore runtimes & tools

Previously seen as "(NOT)Models". This collection separates supporting tools and specialist approaches from the general-purpose decision model catalog.

Same goal. Different layers.

A model, a runtime and a wrapper solve different parts of the problem. Start with what you need to control.

How three approaches to typed decisions differ
Compare
Decision modelThe trained weights
RuntimeThe inference layer
WrapperThe integration layer
What it addsModelWeights trained for decision tasks.RuntimeDecision extraction or constrained decoding over a base model.WrapperAn interface or workflow around an existing model API.
Consider it whenModelYou want to evaluate a purpose-trained decision system.RuntimeYou want control over the base model and how it runs.WrapperYou need simpler integration or orchestration.
Check firstModelTask accuracy, access, evaluation method and calibration evidence.RuntimeHardware, supported types and probability interpretation.WrapperUpstream dependencies, data handling and fallback behavior.
Explore an exampleModelJevRuntimeSemIfWrapperclassifier.dev

Methods and domain specialists also appear in the runtimes collection. Their scope may differ from these three categories; read each entry's inclusion rationale.

Explore the model quadrant

Compare models at a glance.

See how decision models stack up on maturity and capability. The quadrant visualizes our evaluation rubric—bubble size reflects adoption signals.

Open the interactive quadrant

Good questions.
Clearer decisions.

The distinctions that matter when evaluating decision models and their tooling.

How is a decision model different from a chat model?

A decision model is trained to return structured answers for decision tasks. A chat model primarily generates text, although tools can constrain that output or derive decisions from it. A shared JSON format does not make their training or reliability equivalent. Read the System One introduction.

Is a runtime a trained decision model?

Not by itself. A runtime controls inference: it can restrict decoding or read scores from an existing model without training new decision weights. The broader runtimes collection also includes wrappers and specialists, so each entry explains why it is listed there. Explore the collection.

Are probabilities always calibrated?

No. A probability or score is not automatically a trustworthy estimate of correctness. Calibration must be evaluated on a relevant dataset and can change across tasks. Look for evidence about calibration, thresholds and failure cases rather than relying on the output format alone.

When should I use a model, a runtime or a wrapper?

Evaluate a decision model when you need purpose-trained capabilities. Explore a runtime when base-model choice and execution control matter. Consider a wrapper when integration and orchestration are the main concern. In every case, test representative inputs, constraints and failure paths. Compare the approaches.

Can decision models replace generative models?

They often serve different roles. A decision layer can classify an input or select a route, while a generative model produces text when the workflow requires it. This can be combined in a dual-process system, but any cost or latency benefit depends on the actual workload. Learn about dual-process architecture.

A catalog, with context.

Evaluate the evidence, not just the label. Product descriptions here summarize the linked catalog entries; they are not independent benchmark results. Follow each profile for sources, reported capabilities and limitations.

Have a correction? Contribute to the catalog.