# Scoring without output tokens: the cost of classifying millions of documents

When a language model classifies by generating an answer, you pay for the text it writes, wait while it writes it,
and then parse it. For a label, a yes/no or a score, the generated text is overhead. spinf classifies by scoring
instead: it returns the probability of each answer you list and generates nothing. There are no output tokens, and
only the input is billed.

## What generation costs in a classification job

A generated answer adds three costs that grow with every document:

- **Output tokens.** Even a short reply ("Category: billing. Confidence: high.") is several tokens, and explanations
  run to hundreds. Output tokens are billed on every item.
- **Parsing.** Free text has to be mapped back to your labels, with rules for replies that do not match any of them.
- **Waiting.** Text is produced one token at a time after the input has been read.

A scored answer has none of these. The response lists your answers with their probabilities; the `choices` of a chat
completion are empty because nothing is written.

## What is billed

Only processed input tokens, at $0.09 per million on spinf-12b (our optimized Gemma 4 12B):

- **the content, once per call**, however many questions you ask about it;
- **each question's own tokens**: the filled template and its answers;
- **one empty floor per question per call**: the same question scored with no content, which gives the question's
  bias ([Calibration](/docs/calibration)). You can skip it with `"empty_floor": false` once you have it.

The scores themselves are free. There is a minimum of 1,000 tokens per call, counted on the whole call, so batch
work sends many documents per call with `inputs`: the minimum and the floors are then shared. The details are in
[Limits & pricing](/docs/limits).

## The cost of a million documents

Take documents of about 300 tokens, each asked 10 questions of about 20 tokens:

| | |
|---|---|
| content | 300 tokens |
| + 10 questions × 20 tokens | 200 tokens |
| = tokens per document | about 500 |
| × 1,000,000 documents | 500M tokens |
| × $0.09 per million | **about $45** |

With many documents per call, the floors add almost nothing per document. The same arithmetic on the use-case pages
gives about $117 for a million 800-token news articles with 20 questions each, and about $32 for a million
120-token survey answers with 8 questions.

## Many questions cost little

Because the content is read once per call, the cost per question is only its own tokens. Asking 10 questions of
a document costs far less than 10 separate classifications of it: the 300 tokens of content are paid once, not ten
times. This changes how jobs are designed:

- ask the questions you would otherwise skip, such as a policy check next to the topic and sentiment;
- ask the same question for every company, product or aspect with a placeholder, up to 1,000 values per question;
- re-score an archive when the codebook or the questions change, instead of living with the first version.

## Labelled examples are input too

Few-shot examples placed in the prompt raise accuracy the most ([zero-shot and few-shot
classification](/guides/zero-shot-few-shot-classification)), and they are part of the content, so they are billed
with every document. In our ticket benchmark, 8 short labelled examples added about 1,100 tokens per ticket: about
$0.12 per 1,000 tickets on spinf-12b. For content moderation, the use-case figures are about $0.01 per 1,000 posts
with no examples and about $0.08 with 8 examples.

## Batch LLM processing at scale

The on-demand API is built for throughput: backfills, archives and periodic batches, self-serve. The Thinking Text
news indexes were built this way on 15 million articles and about a trillion tokens across research iterations. For
projects over a trillion tokens we also offer dedicated deployments and file-based batch; for decisions in live flows,
a reserved endpoint.

New accounts start with 50M free tokens: enough to score about 100,000 documents of the size above, or to test
several question wordings on a few hundred of your own items many times over.

## Try it

The [classification API](/classification-api) page runs the cost example and a live call; the
[news analysis](/use-cases/news-analysis) and [survey & review analysis](/use-cases/survey-review-analysis) pages show
it per use case. Related: [Limits & pricing](/docs/limits) and
[probabilities from a language model](/guides/llm-probabilities).
