Decision models: Gemma 4 as a calibrated decision model, at scale
A decision model is an AI model used to hand software a decision rather than a paragraph. You give it a state (a message, a document, a record from your application) and the options you allow, and it returns a choice or, better, a probability for each option. There is no text to generate, parse or validate: the output is one of your options, with a number your code can act on.
The term has older meanings too. In business process management, Decision Model and Notation (DMN) is a standard for writing decision rules as tables and diagrams, and in research and management a decision model is often a decision tree that maps choices, outcomes and their probabilities. This guide is about the AI sense: a language model used to decide, not to write.
A leading open-weights model, used as a decision model
spinf is a specialized inference provider. It runs Gemma 4, Google DeepMind's open-weights model family, optimized to be used as a decision model at very large scale:
- You define the decision in plain language. A question and the options you allow, written as a template. No training run, no labelled dataset to start, and no new model to adopt for each new decision.
- You get a probability for every option. Normalised over your options, so they sum to 1. The answer is always one of yours.
- The probabilities come with a calibration check. Each question is also scored with empty content, so you can see how much the state itself moved each option.
- One read, many decisions. The state is read once and every question is answered from that read, so a second or a hundredth decision about the same state costs only its own tokens.
Because the model is a general one, the same call can take a routing decision, a policy check and a risk flag about
the same input, and a new decision is a new question, not a new model. spinf-12b is Gemma 4 12B; the larger
spinf-31b (Gemma 4 31B) is available on request.
An example: a decision on application state
The state does not have to be prose. Here it is an order record, and the decision is what to do next:
{
"model": "spinf-12b",
"messages": [{ "role": "user", "content": "Order:\n{\"order_id\": \"A-10482\", \"amount_usd\": 1840, \"customer_since\": \"2026-09-27\", \"previous_orders\": 0, \"shipping_country\": \"US\", \"billing_country\": \"BR\", \"items\": [\"gift card x4\"], \"payment\": \"new card, 3 declined attempts before success\"}" }],
"scoring": { "queries": [{
"id": "action",
"template": "\n\nNext action (ship, hold for review, or cancel):{?}",
"options": [" ship", " hold for review", " cancel"]
}]}
}
A live call to spinf-12b, with no examples, returned:
| Option | p | floor_p (no order) | Lift |
|---|---|---|---|
| ship | 0.59 | 0.75 | −0.16 |
| hold for review | 0.30 | 0.04 | +0.26 |
| cancel | 0.11 | 0.22 | −0.11 |
The largest raw probability is still "ship", because the question alone leans heavily that way. The floor shows what the order itself says: it pushes "hold for review" up from 0.04 to 0.30 and pushes the other two down. This is why a decision model should return calibrated probabilities rather than a single label: the label would have been "ship". Read the lift, or rank each score against your own history (Calibration), and set thresholds on it (probabilistic if statements). A few labelled orders from your history in the content sharpen the scores further (few-shot classification).
Measured on common decisions
On the test sets of our public benchmark, spinf-12b with a few labelled examples in the prompt:
- Is this ticket urgent? 92% correct (8 examples, 130 synthetic tickets).
- Can a bot resolve it? 93% correct (8 examples, same tickets).
- Does this post break the policy? 99% correct (4 examples, 130 synthetic posts).
- Is this message a prompt injection? 92% correct (4 examples, calibrated, on the public deepset/prompt-injections test split).
Synthetic sets are cleaner than real data: measure on a labelled sample of your own decisions before you rely on the thresholds. The sets, questions and code are on GitHub.
Decisions at massive scale
Only input is billed, at $0.09 per million tokens on spinf-12b; the scores are free. A decision on a 100-token record
with one 20-token question is about 120 tokens, so a million such decisions cost about $11, sent with many records
per call (each call bills at least 1,000 tokens). The first 50M tokens are free: about 400,000 of those decisions.
Scoring without output tokens shows the arithmetic.
This makes it practical to decide on everything rather than a sample: every order, every ticket, every post, every record in an archive, and to re-run the whole history when a policy or a threshold changes.
Bulk and real time
The on-demand API is built for throughput: backfills, backlogs, nightly sweeps, and measuring a new decision on your history before it goes live. Decisions inside a live flow (routing a ticket as it arrives, holding an order at checkout) run on a reserved endpoint, sized to your traffic, with the same questions and thresholds.
Try it
The classification API page shows a live call with several decisions about one email, and the use cases apply the same idea to tickets and content moderation. Next: probabilities from a language model or the quickstart.
Gemma is a trademark of Google LLC.
Raw Markdown for agents: /guides/decision-models.md · /llms.txt