# Decision models: Gemma 4 as a calibrated decision model, at scale

A **decision model** is an AI model used to hand software a decision rather than a paragraph. You give it a state (a
message, a document, a record from your application) and the options you allow, and it returns a choice or, better, a
probability for each option. There is no text to generate, parse or validate: the output is one of your options, with
a number your code can act on.

The term has older meanings too. In business process management, **Decision Model and Notation (DMN)** is a standard for
writing decision rules as tables and diagrams, and in research and management a decision model is often a **decision
tree** that maps choices, outcomes and their probabilities. This guide is about the AI sense: a language model used to
decide, not to write.

## A leading open-weights model, used as a decision model

spinf is a specialized inference provider. It runs Gemma 4, Google DeepMind's open-weights model family, optimized to
be used as a decision model at very large scale:

- **You define the decision in plain language.** A question and the options you allow, written as a template. No
  training run, no labelled dataset to start, and no new model to adopt for each new decision.
- **You get a probability for every option.** Normalised over your options, so they sum to 1. The answer is always one
  of yours.
- **The probabilities come with a calibration check.** Each question is also scored with empty content, so you can see
  how much the state itself moved each option.
- **One read, many decisions.** The state is read once and every question is answered from that read, so a second or a
  hundredth decision about the same state costs only its own tokens.

Because the model is a general one, the same call can take a routing decision, a policy check and a risk flag about
the same input, and a new decision is a new question, not a new model. `spinf-12b` is Gemma 4 12B; the larger
`spinf-31b` (Gemma 4 31B) is available on request.

## An example: a decision on application state

The state does not have to be prose. Here it is an order record, and the decision is what to do next:

```json
{
  "model": "spinf-12b",
  "messages": [{ "role": "user", "content": "Order:\n{\"order_id\": \"A-10482\", \"amount_usd\": 1840, \"customer_since\": \"2026-09-27\", \"previous_orders\": 0, \"shipping_country\": \"US\", \"billing_country\": \"BR\", \"items\": [\"gift card x4\"], \"payment\": \"new card, 3 declined attempts before success\"}" }],
  "scoring": { "queries": [{
    "id": "action",
    "template": "\n\nNext action (ship, hold for review, or cancel):{?}",
    "options": [" ship", " hold for review", " cancel"]
  }]}
}
```

A live call to `spinf-12b`, with no examples, returned:

| Option | `p` | `floor_p` (no order) | Lift |
|---|---|---|---|
| ship | 0.59 | 0.75 | −0.16 |
| hold for review | 0.30 | 0.04 | +0.26 |
| cancel | 0.11 | 0.22 | −0.11 |

The largest raw probability is still "ship", because the question alone leans heavily that way. The floor shows what
the order itself says: it pushes "hold for review" up from 0.04 to 0.30 and pushes the other two down. This is why a
decision model should return calibrated probabilities rather than a single label: the label would have been "ship".
Read the lift, or rank each score against your own history ([Calibration](/docs/calibration)), and set thresholds on
it ([probabilistic if statements](/guides/probabilistic-if)). A few labelled orders from your history in the content
sharpen the scores further ([few-shot classification](/guides/zero-shot-few-shot-classification)).

## Measured on common decisions

On the test sets of our public benchmark, `spinf-12b` with a few labelled examples in the prompt:

- **Is this ticket urgent?** 92% correct (8 examples, 130 synthetic tickets).
- **Can a bot resolve it?** 93% correct (8 examples, same tickets).
- **Does this post break the policy?** 99% correct (4 examples, 130 synthetic posts).
- **Is this message a prompt injection?** 92% correct (4 examples, calibrated, on the public deepset/prompt-injections test split).

Synthetic sets are cleaner than real data: measure on a labelled sample of your own decisions before you rely on the
thresholds. The sets, questions and code are on [GitHub](https://github.com/spinfinc/spinf-benchmarks).

## Decisions at massive scale

Only input is billed, at $0.09 per million tokens on `spinf-12b`; the scores are free. A decision on a 100-token record
with one 20-token question is about 120 tokens, so **a million such decisions cost about $11**, sent with many records
per call (each call bills at least 1,000 tokens). The first 50M tokens are free: about 400,000 of those decisions.
[Scoring without output tokens](/guides/no-output-tokens) shows the arithmetic.

This makes it practical to decide on everything rather than a sample: every order, every ticket, every post, every
record in an archive, and to re-run the whole history when a policy or a threshold changes.

## Bulk and real time

The on-demand API is built for throughput: backfills, backlogs, nightly sweeps, and measuring a new decision on your
history before it goes live. Decisions inside a live flow (routing a ticket as it arrives, holding an order at
checkout) run on a **reserved endpoint**, sized to your traffic, with the same questions and thresholds.

## Try it

The [classification API](/classification-api) page shows a live call with several decisions about one email, and the
use cases apply the same idea to [tickets](/use-cases/ticket-triage) and [content moderation](/use-cases/content-moderation).
Next: [probabilities from a language model](/guides/llm-probabilities) or the [quickstart](/docs/quickstart).

Gemma is a trademark of Google LLC.
