LLM classification API

Classify and score millions of items with leading open-weights language models

spinf is a specialized inference provider: we run Gemma 4, optimized for classification and scoring at scale. Send your content with your questions and the answers you allow, and get the probability of each answer, for every question, from one read of the content. No training, no model to host, and only input is billed.

Read the quickstart
Cost at scale

1M documents for about $45

On spinf-12b, our optimized Gemma 4 12B, only input is billed, at $0.09 per million tokens: the content once per call, each question's own tokens, and one empty floor per question. The scores are free.

Combine use cases in the same call

Most jobs are several questions about the same content: tone, topic, urgency and policy for a message, sentiment and events per company for an article. Ask them together: the content is read once.

documents1,000,000× tokens each (about)500= tokens500M× $0.09 / 1M$45.00

Assumes 300-token documents, 10 questions of about 20 tokens each. The first 50M tokens are free.

Pricing details
On demand · self-serve

$0.09 per million input tokens · the scores are free

Why spinf

Calibrated answers for classification & scoring, at any scale

Your labels, with probabilities

The answer is always one of the answers you list, with its probability: nothing to parse, no invented labels, and thresholds you can set on your own history.

Hundreds of questions, one read

Sentiment, topic, urgency, policy, your own codes: ask them all of the same content in one call. The content is read once, and each extra question costs only its own tokens.

Taught by examples, not training

Write the questions in plain language, and put a few labelled examples in the content when the categories are subtle. No fine-tuning, no model to host or scale.

Live example

One email, four questions

Tone, topic, urgency and a yes/no check, answered together. The bars show p; the tick shows the floor, the same question asked with no email.

Content
Customer email: You charged my card twice for order #4471 and nobody has answered my two emails since Monday. If this is not refunded by Friday I will dispute it with my bank and close my account.
Is the customer's tone positive, neutral or negative?

positive

0.04

neutral

0.08

negative

0.88

Topic (billing, delivery, product, or account):

billing

0.85

delivery

0.02

product

0.07

account

0.05

Is this urgent (money at stake, a legal or bank threat, or a deadline within a few days)?

yes

0.81

no

0.19

Does the customer threaten to close the account or dispute the charge?

yes

0.88

no

0.12

import os
import requests

resp = requests.post(
    "https://api.spinf.com/v1/score",
    headers={"Authorization": f"Bearer {os.environ['SPINF_API_KEY']}"},
    json={
      "model": "spinf-12b",
      "messages": [
        {
          "role": "user",
          "content": "Customer email:\nYou charged my card twice for order #4471 and nobody has answered my two emails since Monday. If this is not refunded by Friday I will dispute it with my bank and close my account."
        }
      ],
      "scoring": {
        "queries": [
          {
            "id": "sentiment",
            "template": "\n\nIs the customer's tone positive, neutral or negative?\nAnswer:{?}",
            "options": [
              " positive",
              " neutral",
              " negative"
            ]
          },
          {
            "id": "topic",
            "template": "\n\nTopic (billing, delivery, product, or account):\nAnswer:{?}",
            "options": [
              " billing",
              " delivery",
              " product",
              " account"
            ]
          },
          {
            "id": "urgent",
            "template": "\n\nIs this urgent (money at stake, a legal or bank threat, or a deadline within a few days)?\nAnswer:{?}",
            "options": [
              " yes",
              " no"
            ]
          },
          {
            "id": "threat",
            "template": "\n\nDoes the customer threaten to close the account or dispute the charge?\nAnswer:{?}",
            "options": [
              " yes",
              " no"
            ]
          }
        ]
      }
    },
)
for query in resp.json()["results"][0]["queries"]:
    for combo in query["combinations"]:
        print(query["id"], combo["values"], [(o["text"], o["p"], o["floor_p"]) for o in combo["options"]])

A live call to spinf-12b, with no examples. Read the lift over the floor: each leading answer sits clearly above what the question alone would give. Add a few labelled items of your own to match your team’s decisions more closely.

In production · Thinking Text

Hundreds of questions, millions of articles

Thinking Text asks about five hundred fixed questions of every oil article and aggregates the answers into daily indexes of what the press says.

15M
news articles
five years, 2022–2026
~1T
tokens processed
across research iterations
~500
questions per article
one read of each article

Figures from the Thinking Text oil index, published daily with its score card.

Benchmark · five jobs

Measured on public and synthetic test sets

spinf-12b on the test sets of our public benchmark. The synthetic sets (108–130 items each) were written for it, labelled by a language model and checked blind by a second model; the public sets are named.

97%
correct polarity · public product reviews
300 reviews · no examples needed (calibrated)
92%
correct · prompt injection?
public deepset/prompt-injections test split · 4 examples in the prompt
95%
correct sentiment per company
255 mentions in synthetic news · 8 examples in the prompt
93%
correct ticket topic (5)
130 synthetic tickets · 8 examples in the prompt

Synthetic sets are cleaner than real data, and the wording of the questions matters: measure on a labelled sample of your own content before relying on the numbers.

Test sets, questions and code on GitHub: reproduce it with your own key
The questions matter: iterate on yours

As with any language model, the wording of the questions and a few labelled examples change the results a lot. With spinf that is cheap to get right: try a few wordings and example sets on a few hundred of your own items for cents, and keep the best. How to write questions

How it works

From questions to numbers

Self-serve on the on-demand API: sign up, create a key and send your first batch in minutes.

1

Write the questions

Plain-language questions with the answers you allow: yes/no, a list of categories, a scale. Placeholders ask the same question for every company, product or aspect.

2

Test on a sample

Score a few hundred items you have already labelled, with your free tokens. Try a few wordings and example sets for cents, and keep the best.

3

Run at scale

Send the rest in batches with many items per call on the on-demand API, or talk to us about a reserved endpoint for real time.

FAQ

Classification & scoring: questions, answered

More in the docs, or .

A specialized inference provider. We run leading open-weights language models on our own inference engine, optimized for specific jobs at very large scale, and offer them through a simple API: you get the performance and the price without hosting, tuning or scaling models yourself. Our first job is analysis: classification and scoring.

No. spinf scores: for each question you get the probability of each answer you listed, and nothing is generated. That is why the scores are free and only the input is billed.

spinf-12b is Gemma 4 12B and spinf-31b (on request) is Gemma 4 31B, both optimized by us for scoring at scale. The API accepts either name, for example "google/gemma-4-12B".

No. You write the questions and the answers in plain language. When the categories are subtle, put a few labelled examples in the content: in our benchmark, 8 examples often took accuracy from the 70s or 80s into the 90s.

$0.09 per million input tokens on spinf-12b, and the scores are free. Each call bills the content once, each question’s own tokens and one empty floor per question. The first 50M tokens are on us.

Yes, on a reserved endpoint: the on-demand API is built for throughput rather than real time. Archives, backlogs and periodic batches run on demand, self-serve.

No. Content sent for scoring is processed in memory and not stored by default, and it is never used to train models.
Try it on your data

The first 50M tokens are on us

Sign up, create a key and score your first batch in minutes.