Decision engine at massive scale

Millions of decisions, one read of each content

Decision models at scale. Each content is read once and every prompt about it is answered in parallel, with a calibrated probability for each answer. For hundreds of thousands to millions of contents, on our API or on a private deployment.

Private deployments

The same engine on GPUs dedicated to you, sized to your volume and deadline, including in your own environment.

Leading open-weights models

Gemma 4 12B self-serve; Gemma 4 31B and Qwen3.8 27B on request. Your own fine-tune of a supported model on a private deployment.

Published research →

750,000 news articles read against 257 Polymarket questions, in a public paper.

Content not stored

Processed in memory by default, and never used to train models.

Cost at scale

1M documents for about $54

On spinf-12b, our optimized Gemma 4 12B, only input is billed, at $0.09 per million tokens: the content once per call, each question's own tokens, and one empty floor per question. The scores are free.

Combine use cases in the same call

A decision is rarely alone: the outcome, the evidence behind it, the policy it touches. Ask them all of the same content in one call: it is read once.

documents1,000,000× tokens each (about)600= tokens600M× $0.09 / 1M$54.00

Assumes 300-token documents and 20 prompts of about 15 tokens each. The first 50M tokens are free. Gemma is a trademark of Google LLC.

Pricing details
On demand · self-serve

$0.09 per million input tokens · the scores are free

Why spinf

Calibrated decisions, at any scale

One content, many prompts

Decision APIs that take one prompt per call send the same document again for every decision. In parallel processing mode the content is read once and every prompt is answered from that read: extra decisions cost only their own tokens, in money and in time.

Throughput for millions of contents

Gemma 4 12B processes over 20,000 tokens per second on a single GPU, and Qwen3.8 27B about 10,000, in parallel processing mode. A backlog of a million documents is a matter of hours, not weeks.

A probability, not a paragraph

Every decision is one of the answers you allow, with its probability and the floor of the same prompt with no content: thresholds you can set, audit and keep stable over millions of reads.

Live example

One article, five decisions

Four forecasts from one template, and a check on the evidence, all answered from one read of the article. The bars show p; the tick shows the floor, the same prompt with no article.

Content
Inflation cools for a third month as hiring slows Consumer prices rose 0.1% in September, below the 0.3% economists expected, and employers added 61,000 jobs, the weakest gain this year. Two members of the central bank's rate committee said on Tuesday that the case for cutting rates at the next meeting "has clearly strengthened".

Question: Does this article make it more likely that {outcome}?

the central bank cuts rates at its next meeting

Yes

0.85

No

0.15

the central bank raises rates at its next meeting

Yes

0.17

No

0.83

unemployment rises over the next three months

Yes

0.63

No

0.38

inflation picks up again next month

Yes

0.37

No

0.63

Does the article report new economic figures?

Yes

0.94

No

0.06

import os
import requests

resp = requests.post(
    "https://api.spinf.com/v1/score",
    headers={"Authorization": f"Bearer {os.environ['SPINF_API_KEY']}"},
    json={
      "model": "spinf-12b",
      "messages": [
        {
          "role": "user",
          "content": "Inflation cools for a third month as hiring slows\n\nConsumer prices rose 0.1% in September, below the 0.3% economists expected, and employers added 61,000 jobs, the weakest gain this year. Two members of the central bank's rate committee said on Tuesday that the case for cutting rates at the next meeting \"has clearly strengthened\"."
        }
      ],
      "scoring": {
        "queries": [
          {
            "id": "outcome",
            "template": "\n\nQuestion: Does this article make it more likely that {outcome}?\nAnswer:{?}",
            "options": [
              " Yes",
              " No"
            ],
            "combinations": [
              {
                "outcome": "the central bank cuts rates at its next meeting"
              },
              {
                "outcome": "the central bank raises rates at its next meeting"
              },
              {
                "outcome": "unemployment rises over the next three months"
              },
              {
                "outcome": "inflation picks up again next month"
              }
            ]
          },
          {
            "id": "data",
            "template": "\n\nDoes the article report new economic figures?\nAnswer:{?}",
            "options": [
              " Yes",
              " No"
            ]
          }
        ]
      }
    },
)
for query in resp.json()["results"][0]["queries"]:
    for combo in query["combinations"]:
        print(query["id"], combo["values"], [(o["text"], o["p"], o["floor_p"]) for o in combo["options"]])

A live call to spinf-12b, with no examples. The article moved “a rate cut at the next meeting” from 0.47 with no article to 0.85, and “a rate hike” from 0.44 to 0.17. A placeholder asks the same prompt for every outcome: add outcomes for the cost of their own tokens.

The API call behind this example, the docs and the benchmark code are easier to read on a larger screen. Start free now, and pick it up on your laptop.

Throughput · parallel processing mode

Built for hundreds of thousands to millions of contents

One content, many prompts: the engine reads each content once and scores every prompt about it in parallel, so adding decisions barely adds time.

20K+

tokens per second

Gemma 4 12B · a single GPU
10K

tokens per second

Qwen3.8 27B · a single GPU
~8 h

for 1M documents

20 prompts each · Gemma 4 12B · one GPU

Throughput in parallel processing mode on a single GPU, for 300-token documents with 20 prompts of about 15 tokens each. On the shared on-demand API it depends on the load; private deployments add GPUs to meet your deadline.

Research · Thinking Text

What the news says vs. what Polymarket prices

Thinking Text read every article about the 200 most traded Polymarket events against each of their outcomes, three ways, and compared the news with the market prices hour by hour: decisions at scale, in public.

The paper

What the news says vs. what Polymarket prices: 257 markets, nine months, 750k articles

Thinking Text, 2026-10-04. Read with spinf-12b through the spinf scoring API, with no examples.

750k

news articles

January to October 2026
849

outcomes

across 257 market questions
3

reads per outcome

for every article, with its floor

Figures from the Thinking Text Polymarket news index and its paper.

The questions matter: iterate on yours

As with any language model, the wording of the questions and a few labelled examples change the results a lot. With spinf that is cheap to get right: try a few wordings and example sets on a few hundred of your own items for cents, and keep the best. How to write questions

How it works

From questions to numbers

Self-serve on the on-demand API: sign up, create a key and send your first batch in minutes.

1

Write the decisions as prompts

One prompt per decision, with the answers you allow. Placeholders fill in every outcome, rule or criterion, so one template covers hundreds of decisions.

2

Calibrate on a sample

Score a few hundred contents you already know the answer for, with your free tokens. Keep the wording and the thresholds that match your decisions.

3

Run it at scale

Send millions of contents in batches on the on-demand API, or on a private deployment sized to your volume and deadline.

FAQ

Decision engine: questions, answered

More in the docs and the guides: decision models, probabilistic if statements and probabilities from a language model, or .

One content, many prompts. You send each content once with all the prompts you want answered about it; spinf reads the content a single time and answers every prompt from that read, in parallel. With hundreds of thousands to millions of contents, it is the fastest way to precise readouts.

In parallel processing mode Gemma 4 12B processes over 20,000 tokens per second on a single GPU, and Qwen3.8 27B about 10,000. A million 300-token documents with 20 prompts each take about 8 hours on one GPU with Gemma 4 12B. On the shared on-demand API throughput depends on the load; a private deployment adds GPUs to meet your deadline.

There, each call carries one content and one prompt, so a hundred decisions about the same document send it a hundred times and wait for each. Here the content is read once and every decision is answered from that read: the extra prompts cost only their own tokens.

spinf-12b (Gemma 4 12B) self-serve, and Gemma 4 31B and Qwen3.8 27B on request, all optimized by us for decisions at scale. On a private deployment we can also run your own fine-tune of a supported model.

Yes. Private deployments run the engine on GPUs dedicated to you: capacity we reserve for you, or an install in your own environment when your data must stay there. Tell us about the volume, the deadline and the constraints.

$0.09 per million input tokens on spinf-12b, and the scores are free: each call bills the content once, each prompt’s own tokens and one empty floor per prompt. The first 50M tokens are on us. Qwen3.8 27B, Gemma 4 31B and private deployments are priced per project.

No. Content sent for scoring is processed in memory and not stored by default, and it is never used to train models.
Try it on your data

The first 50M tokens are on us

Sign up, create a key and score your first batch in minutes.