Open source · decision inference

The open-source inference engine for decision models

Read once, decide hundreds of times. spinf serves decision models, and turns standard open models into decision models: each input is read once, every question about it is answered in parallel, and each answer comes back as a calibrated probability instead of generated text.

Open-source release planned for October 20, 2026
Latest news
  • 2026-10

    The spinf engine goes open source: release planned for October 20, 2026.

  • 2026-10

    The spinf/moderation prompt pack for content moderation, with its paper and a public benchmark. Read the paper →

  • 2026-10

    Decision mode at massive scale: 750,000 news articles read against 257 prediction-market questions. Decision engine →

About

A decision engine, not a chat server

spinf is a fast, easy-to-use engine for decision inference and decision model serving. Where general-purpose LLM servers generate text one prompt at a time, spinf answers many questions about each input in one pass and returns scores you can threshold: classification, extraction checks, routing, forecasting and policy decisions over millions of inputs.

spinf is fast

  • Decision mode: each input is read once and every question about it is answered from that read, in parallel.
  • Scores, not generated text: no token-by-token output loop for a decision.
  • Scheduling built for many short inputs with hundreds of questions each, to keep the GPU full.
  • Optimized for leading open-weights models.

spinf is built for decisions

  • A probability for each answer you allow: yes / no, a label, an outcome, a level.
  • The floor of each question (the same prompt with no input) to calibrate against.
  • Reproducible scores: the same input gets the same score, whatever else is in the batch.
  • Thresholds you can set, audit and keep stable over millions of decisions.

spinf is easy to use

  • One request: the input, your questions and their answers. Placeholders expand one template into hundreds of questions.
  • Ready-made prompt packs for common jobs, scored in the same call as your own questions.
  • The same request format as the hosted spinf API, so code moves between the two.
  • Runs on your own NVIDIA GPUs, in your own environment.

Decision mode

One input, many questions, one read

Most real jobs ask many questions about the same content: intent, urgency and policy for a message; category and attributes for a product; sentiment and events for an article. Decision mode reads the content once and answers them all from that read, so each extra decision costs only its own tokens, in money and in time.

General-purpose LLM servingspinf decision mode
Unit of workOne prompt, one generated answerOne input, hundreds of questions
OutputGenerated text, parsed afterwardsA probability for each allowed answer
200 questions about one documentThe document is sent and read 200 timesThe document is read once
Cost of one more questionThe whole prompt againOnly the question’s own tokens
Thresholds and auditsDepend on sampling and on parsing the textCalibrated, reproducible scores

Output: a probability for each allowed answer.

200 questions about one document: the document is read once.

Cost of one more question: only the question’s own tokens.

Models

Leading open-weights models, as decision models

Standard instruction models become decision models in decision mode, and dedicated decision models run as they are. Open weights: Gemma 4 by Google DeepMind and Qwen3.8 by the Qwen team.

Gemma 4 12B

google/gemma-4-12B

Gemma 4 31B

google/gemma-4-31B

Qwen3.8 27B

Qwen3.8 family

Your decision model

a fine-tune of a supported model

Gemma is a trademark of Google LLC. The list at release may grow.

Prompt packs

Ready-made decision sets to start from

Each pack is a set of questions for a common business job, with the few decisions they add up to. They ship with the release as quickstart examples, and run in the playground of the hosted API.

Content moderation

Harm categories with thresholds you set, for posts, comments and messages.

Available today on the hosted API

News analysis

Sentiment, impact and events per article, for markets and companies.

With the release

Inbound triage

Intent, urgency, sentiment, risk flags and phishing signals for emails, chats and tickets.

With the release

Product listings

Category, attributes and listing compliance for product pages and listings.

With the release

Entertainment and media

Genre, mood, themes, content advisories and audience, from a synopsis.

With the release

Reviews and feedback

Sentiment per aspect, feature requests and competitor mentions.

With the release

Brand safety

Topics for contextual advertising and brand-safety risk, for web pages.

With the release
Getting started

Try decision mode today

Until the release, decision mode runs on the hosted spinf API. One request carries the message, a question with five intents, and one template asked three ways: eight answers scored from a single read of the message.

POST /v1/score · hosted API
curl https://api.spinf.com/v1/score \
  -H "Authorization: Bearer $SPINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spinf-12b",
    "messages": [ { "role": "user",
      "content": "Charged twice for order 48213. Fix it today or I cancel." } ],
    "scoring": {
      "queries": [
        {
          "id": "intent",
          "template": "\n\nThe customer is writing about{?}",
          "options": [" billing", " delivery", " a technical problem",
                      " their account", " something else"]
        },
        {
          "id": "flag",
          "template": "\n\nDoes the message {signal}? Answer:{?}",
          "options": [" yes", " no"],
          "combinations": [
            { "signal": "threaten to cancel" },
            { "signal": "ask for a refund" },
            { "signal": "need a reply today" }
          ]
        }
      ]
    }
  }'
Run it your way

Open source, hosted or private

The same engine three ways: on your own GPUs, on demand through the spinf API, or on a private deployment we run for you.

Open source

Run the engine on your own GPUs, free, under an open-source license. Available from the release.

Hosted API

Decision mode on demand today, paid per million input tokens, with free credit to start.

Private deployment

The engine on GPUs dedicated to you, sized to your volume and deadline, run by us.

FAQ

Decision inference, explained

A language model used to choose among a fixed set of answers (yes or no, a label, an outcome, a level) and to return a probability for each, instead of writing free text. Some models are trained for it; any open instruction model can work as a decision model in decision mode.

Running a model to get decisions rather than generated text: the input is read, the answers you allow are scored, and a probability comes back for each. Nothing is generated, so a decision costs only the tokens it reads.

Serving decision models in production: many inputs, many questions per input, high throughput per GPU, and scores that stay stable and reproducible so thresholds and audits hold. General-purpose LLM servers are built for chat and text generation; spinf is built for decisions.

A chat server generates text one token at a time and treats each prompt as its own request: 200 questions about one document mean sending the document 200 times, or one long prompt whose answer you have to parse. In decision mode spinf reads the document once, answers every question from that read in parallel, and returns a probability for each allowed answer.

The release is planned for October 20, 2026. Leave your email on this page to hear when it is out. Until then, decision mode is available on the hosted spinf API.

An open-source license that lets anyone use it for free, announced with the release.

Gemma 4 12B and 31B and Qwen3.8 27B at release, and your own decision model if it is a fine-tune of a supported model.

NVIDIA GPUs, on your own servers or in your cloud account. The supported GPU generations are listed with the release.

Ready-made sets of questions for a common job (content moderation, news analysis, inbound triage, product listings, entertainment metadata, reviews, brand safety) that turn many scores into a few decisions. They run in the same call as your own questions.

Yes. The spinf API runs the same engine on demand, paid per million input tokens with free credit to start, and private deployments run it on GPUs dedicated to you.
Release planned for October 20, 2026

Be the first to run it

Leave your email and we will write once, when the repository is public.

Decision engine at scale