Skip to main content
If you use an LLM to classify text, route requests, score content, or make yes/no decisions, you are paying for text generation to get a structured answer. Decision models return typed decisions directly without generating tokens, parsing JSON, or validating schemas. This guide shows when and how to migrate these calls from LLMs to the d1 decision model.

When to Migrate

Migrate when the LLM’s output is a bounded decision, meaning the set of possible answers is known before the call. Keep your LLM when the task requires generating new content:
  • Text generation: drafting emails, summaries, reports, code
  • Open-ended Q&A: answering user questions in natural language
  • Multi-turn conversation: chatbots and copilots
  • Complex reasoning: multi-step logic, math, planning
If the answer is one of N known options, use a decision model. If the answer is a new string the model must compose, use an LLM.

Setup

Install the relevant SDK and configure your client. The LLM examples use the OpenAI SDK, which works with any OpenAI-compatible endpoint. The decision model examples use the TypeSafe SDK to call the d1 decision model through the Liquid AI API. To get an API key, see Decision Models: Setup.

Migration Examples

Each example shows the same task solved with an LLM and with a decision model.

Classification

The standard LLM pattern for classification uses a JSON schema or enum constraint. The model generates tokens that conform to the schema, which you then parse.
What changes:
  • Options are defined inline with descriptions. No Pydantic model or JSON schema needed.
  • You get a probability distribution over all options alongside the top pick, plus a confidence value that summarizes how clear-cut the answer is.
  • Zero output tokens generated. The model evaluates all options in a single call.
  • Each option’s description lives next to its label in criteria instead of in a separate prompt.

Routing

A common pattern uses a small, cheap LLM to classify the complexity of an incoming request and route it to the appropriate model.
What changes:
  • The router call is typically faster and cheaper than an LLM-based router.
  • You get a confidence value, so you can fall back to a stronger model when the router is uncertain instead of trusting a binary label.
  • Adding or removing a tier is a change to criteria, not a prompt rewrite.

Binary Decision

The typical LLM approach asks whether content is safe, either with structured output or by parsing a yes/no from the response.
What changes:
  • The probability (0.0 to 1.0) lets you set thresholds for different actions instead of committing to a binary label.
  • The result is always a float. No string parsing or schema validation needed.
  • Decision models produce more consistent results on repeated evaluations of the same input, reducing verdict flips.

Scoring

The LLM approach for scoring asks the model to rate something on a scale, typically by generating a number or picking from a rubric with structured output.
What changes:
  • Levels are indexed from 0 in the order you list them, so a 4-level rubric scores from 0 to 3. A score of 2.999 means nearly all weight is on the top level (Critical).
  • The score is continuous, not a forced integer. It captures how strongly the model leans toward a level.
  • You get the full probability distribution across all levels, so you can detect ambiguous cases.
  • No JSON parsing, no schema validation, no retries on malformed output.

Reranking

In retrieval-augmented generation (RAG) pipelines, a common pattern uses an LLM to score the relevance of each retrieved chunk before passing results to the generation step.
What changes:
  • Each relevance score is a calibrated probability, not an arbitrary integer on a 1-5 scale.
  • You can threshold directly (keep everything above 0.5) instead of guessing where to draw the line on integer scores.
  • Each call is faster and cheaper, which compounds across large result sets.
  • If you need graded relevance instead of relevant/not relevant, use a Score with defined relevance levels.

Multi-Question

LLM pipelines often chain multiple classification calls sequentially, one per decision. Decision models evaluate all questions against the same state in a single call.
What changes:
  • Three network round-trips become one.
  • All three decisions are evaluated against the exact same snapshot of state.
  • Each question keeps its own type and probabilities. With an LLM you could merge the three schemas into one call, but you would still get only labels, with no probabilities.
  • Questions are evaluated in parallel, so adding a question adds little latency compared to adding another LLM call.

Next Steps