Jev
by TypeSafe AIA hosted decision model for software that needs a typed answer. Explore its interface, reported capabilities and limitations before integrating.
- Decision types
- choice · score · boolean
- Access
- Hosted API
A field guide to decision AI
Not every AI task needs a paragraph. Explore models that return typed decisions, and the runtimes and harnesses that bring them into your software.
Weights trained for typed decisions.
Inference techniques over existing weights.
Structured output, not a conversation
Decision models are AI models trained to return structured decisions—such as a selected option, a score or a boolean—rather than free-form prose. Often called System One models, they can classify, route and score inputs within a larger software workflow.
Typed output constrains the answer's format. It does not guarantee a correct decision, calibrated confidence or better performance on your task.
Understand System One01 / The models
A hosted decision model for software that needs a typed answer. Explore its interface, reported capabilities and limitations before integrating.
Explore hosted and open approaches. Compare the model, access requirements and evidence—not just the headline benchmark.
02 / Beyond the weights
There is more than one way to get a typed decision. Runtimes add inference-time techniques to existing models. Wrappers make model APIs easier to use. Specialists address narrower tasks.
Explore runtimes & toolsPreviously seen as "(NOT)Models". This collection separates supporting tools and specialist approaches from the general-purpose decision model catalog.
Browser-first runtime that extracts typed decisions from frozen open LLMs via WebGPU/wllama.
SGLang-based runtime for type-safe generation on autoregressive LLMs, with thinking mode and image input. Open source, and since 29 Sep 2026 also a hosted playground and API.
SGLang-based runtime providing TypeSafe-compatible /v1/systemone endpoint over Qwen3.6-35B-A3B.
HTTP/CLI/MCP wrapper over hosted Jev with smart escalation to reasoning LLMs for uncertain cases.
03 / Know the difference
A model, a runtime and a wrapper solve different parts of the problem. Start with what you need to control.
| Compare | Decision modelThe trained weights | RuntimeThe inference layer | WrapperThe integration layer |
|---|---|---|---|
| What it adds | ModelWeights trained for decision tasks. | RuntimeDecision extraction or constrained decoding over a base model. | WrapperAn interface or workflow around an existing model API. |
| Consider it when | ModelYou want to evaluate a purpose-trained decision system. | RuntimeYou want control over the base model and how it runs. | WrapperYou need simpler integration or orchestration. |
| Check first | ModelTask accuracy, access, evaluation method and calibration evidence. | RuntimeHardware, supported types and probability interpretation. | WrapperUpstream dependencies, data handling and fallback behavior. |
| Explore an example | ModelJev | RuntimeSemIf | Wrapperclassifier.dev |
Methods and domain specialists also appear in the runtimes collection. Their scope may differ from these three categories; read each entry's inclusion rationale.
Explore the model quadrantThe evaluation quadrant
See how decision models stack up on maturity and capability. The quadrant visualizes our evaluation rubric—bubble size reflects adoption signals.
Open the interactive quadrant04 / From understanding to implementation
Understand the trade-offs before choosing a model for classification, routing or scoring.
Read the selection guideMake your first callMove from the concept to an API integration with a practical introduction.
Open the getting-started guideDesign the workflowCombine a decision layer with a generative model in a dual-process architecture.
Explore the dual-process stackA little more context
The distinctions that matter when evaluating decision models and their tooling.
A decision model is trained to return structured answers for decision tasks. A chat model primarily generates text, although tools can constrain that output or derive decisions from it. A shared JSON format does not make their training or reliability equivalent. Read the System One introduction.
Not by itself. A runtime controls inference: it can restrict decoding or read scores from an existing model without training new decision weights. The broader runtimes collection also includes wrappers and specialists, so each entry explains why it is listed there. Explore the collection.
No. A probability or score is not automatically a trustworthy estimate of correctness. Calibration must be evaluated on a relevant dataset and can change across tasks. Look for evidence about calibration, thresholds and failure cases rather than relying on the output format alone.
Evaluate a decision model when you need purpose-trained capabilities. Explore a runtime when base-model choice and execution control matter. Consider a wrapper when integration and orchestration are the main concern. In every case, test representative inputs, constraints and failure paths. Compare the approaches.
They often serve different roles. A decision layer can classify an input or select a route, while a generative model produces text when the workflow requires it. This can be combined in a dual-process system, but any cost or latency benefit depends on the actual workload. Learn about dual-process architecture.
A catalog, with context.
Evaluate the evidence, not just the label. Product descriptions here summarize the linked catalog entries; they are not independent benchmark results. Follow each profile for sources, reported capabilities and limitations.
Have a correction? Contribute to the catalog.