When to Migrate
Migrate when the LLM’s output is a bounded decision, meaning the set of possible answers is known before the call.
Keep your LLM when the task requires generating new content:
- Text generation: drafting emails, summaries, reports, code
- Open-ended Q&A: answering user questions in natural language
- Multi-turn conversation: chatbots and copilots
- Complex reasoning: multi-step logic, math, planning
Setup
Install the relevant SDK and configure your client. The LLM examples use the OpenAI SDK, which works with any OpenAI-compatible endpoint. The decision model examples use the TypeSafe SDK to call the d1 decision model through the Liquid AI API. To get an API key, see Decision Models: Setup.- LLM
- Decision Model
Migration Examples
Each example shows the same task solved with an LLM and with a decision model.Classification
The standard LLM pattern for classification uses a JSON schema or enum constraint. The model generates tokens that conform to the schema, which you then parse.- LLM
- Decision Model
- Options are defined inline with descriptions. No Pydantic model or JSON schema needed.
- You get a probability distribution over all options alongside the top pick, plus a
confidencevalue that summarizes how clear-cut the answer is. - Zero output tokens generated. The model evaluates all options in a single call.
- Each option’s description lives next to its label in
criteriainstead of in a separate prompt.
Routing
A common pattern uses a small, cheap LLM to classify the complexity of an incoming request and route it to the appropriate model.- LLM
- Decision Model
- The router call is typically faster and cheaper than an LLM-based router.
- You get a
confidencevalue, so you can fall back to a stronger model when the router is uncertain instead of trusting a binary label. - Adding or removing a tier is a change to
criteria, not a prompt rewrite.
Binary Decision
The typical LLM approach asks whether content is safe, either with structured output or by parsing a yes/no from the response.- LLM
- Decision Model
- The probability (0.0 to 1.0) lets you set thresholds for different actions instead of committing to a binary label.
- The result is always a float. No string parsing or schema validation needed.
- Decision models produce more consistent results on repeated evaluations of the same input, reducing verdict flips.
Scoring
The LLM approach for scoring asks the model to rate something on a scale, typically by generating a number or picking from a rubric with structured output.- LLM
- Decision Model
- Levels are indexed from 0 in the order you list them, so a 4-level rubric scores from 0 to 3. A score of 2.999 means nearly all weight is on the top level (Critical).
- The score is continuous, not a forced integer. It captures how strongly the model leans toward a level.
- You get the full probability distribution across all levels, so you can detect ambiguous cases.
- No JSON parsing, no schema validation, no retries on malformed output.
Reranking
In retrieval-augmented generation (RAG) pipelines, a common pattern uses an LLM to score the relevance of each retrieved chunk before passing results to the generation step.- LLM
- Decision Model
- Each relevance score is a calibrated probability, not an arbitrary integer on a 1-5 scale.
- You can threshold directly (keep everything above 0.5) instead of guessing where to draw the line on integer scores.
- Each call is faster and cheaper, which compounds across large result sets.
- If you need graded relevance instead of relevant/not relevant, use a Score with defined relevance levels.
Multi-Question
LLM pipelines often chain multiple classification calls sequentially, one per decision. Decision models evaluate all questions against the same state in a single call.- LLM
- Decision Model
- Three network round-trips become one.
- All three decisions are evaluated against the exact same snapshot of state.
- Each question keeps its own type and probabilities. With an LLM you could merge the three schemas into one call, but you would still get only labels, with no probabilities.
- Questions are evaluated in parallel, so adding a question adds little latency compared to adding another LLM call.
Next Steps
- Decision Models for the API reference, primitives, and full code examples
- Model Library for all available Liquid AI models