
The open-source inference engine for decision models
Read once, decide hundreds of times. spinf serves decision models, and turns standard open models into decision models: each input is read once, every question about it is answered in parallel, and each answer comes back as a calibrated probability instead of generated text.
2026-10
The spinf engine goes open source: release planned for October 20, 2026.
2026-10
The spinf/moderation prompt pack for content moderation, with its paper and a public benchmark. Read the paper →
2026-10
Decision mode at massive scale: 750,000 news articles read against 257 prediction-market questions. Decision engine →
A decision engine, not a chat server
spinf is a fast, easy-to-use engine for decision inference and decision model serving. Where general-purpose LLM servers generate text one prompt at a time, spinf answers many questions about each input in one pass and returns scores you can threshold: classification, extraction checks, routing, forecasting and policy decisions over millions of inputs.
spinf is fast
- Decision mode: each input is read once and every question about it is answered from that read, in parallel.
- Scores, not generated text: no token-by-token output loop for a decision.
- Scheduling built for many short inputs with hundreds of questions each, to keep the GPU full.
- Optimized for leading open-weights models.
spinf is built for decisions
- A probability for each answer you allow: yes / no, a label, an outcome, a level.
- The floor of each question (the same prompt with no input) to calibrate against.
- Reproducible scores: the same input gets the same score, whatever else is in the batch.
- Thresholds you can set, audit and keep stable over millions of decisions.
spinf is easy to use
- One request: the input, your questions and their answers. Placeholders expand one template into hundreds of questions.
- Ready-made prompt packs for common jobs, scored in the same call as your own questions.
- The same request format as the hosted spinf API, so code moves between the two.
- Runs on your own NVIDIA GPUs, in your own environment.
One input, many questions, one read
Most real jobs ask many questions about the same content: intent, urgency and policy for a message; category and attributes for a product; sentiment and events for an article. Decision mode reads the content once and answers them all from that read, so each extra decision costs only its own tokens, in money and in time.
| General-purpose LLM serving | spinf decision mode | |
|---|---|---|
| Unit of work | One prompt, one generated answer | One input, hundreds of questions |
| Output | Generated text, parsed afterwards | A probability for each allowed answer |
| 200 questions about one document | The document is sent and read 200 times | The document is read once |
| Cost of one more question | The whole prompt again | Only the question’s own tokens |
| Thresholds and audits | Depend on sampling and on parsing the text | Calibrated, reproducible scores |
Output: a probability for each allowed answer.
200 questions about one document: the document is read once.
Cost of one more question: only the question’s own tokens.
Leading open-weights models, as decision models
Standard instruction models become decision models in decision mode, and dedicated decision models run as they are. Open weights: Gemma 4 by Google DeepMind and Qwen3.8 by the Qwen team.
Gemma 4 12B
google/gemma-4-12B
Gemma 4 31B
google/gemma-4-31B
Qwen3.8 27B
Qwen3.8 family
Your decision model
a fine-tune of a supported model
Gemma is a trademark of Google LLC. The list at release may grow.
Ready-made decision sets to start from
Each pack is a set of questions for a common business job, with the few decisions they add up to. They ship with the release as quickstart examples, and run in the playground of the hosted API.
Content moderation
Harm categories with thresholds you set, for posts, comments and messages.
News analysis
Sentiment, impact and events per article, for markets and companies.
Inbound triage
Intent, urgency, sentiment, risk flags and phishing signals for emails, chats and tickets.
Product listings
Category, attributes and listing compliance for product pages and listings.
Entertainment and media
Genre, mood, themes, content advisories and audience, from a synopsis.
Reviews and feedback
Sentiment per aspect, feature requests and competitor mentions.
Brand safety
Topics for contextual advertising and brand-safety risk, for web pages.
Try decision mode today
Until the release, decision mode runs on the hosted spinf API. One request carries the message, a question with five intents, and one template asked three ways: eight answers scored from a single read of the message.
curl https://api.spinf.com/v1/score \
-H "Authorization: Bearer $SPINF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "spinf-12b",
"messages": [ { "role": "user",
"content": "Charged twice for order 48213. Fix it today or I cancel." } ],
"scoring": {
"queries": [
{
"id": "intent",
"template": "\n\nThe customer is writing about{?}",
"options": [" billing", " delivery", " a technical problem",
" their account", " something else"]
},
{
"id": "flag",
"template": "\n\nDoes the message {signal}? Answer:{?}",
"options": [" yes", " no"],
"combinations": [
{ "signal": "threaten to cancel" },
{ "signal": "ask for a refund" },
{ "signal": "need a reply today" }
]
}
]
}
}'Open source, hosted or private
The same engine three ways: on your own GPUs, on demand through the spinf API, or on a private deployment we run for you.
Open source
Run the engine on your own GPUs, free, under an open-source license. Available from the release.
Hosted API
Decision mode on demand today, paid per million input tokens, with free credit to start.
Private deployment
The engine on GPUs dedicated to you, sized to your volume and deadline, run by us.
Decision inference, explained
Be the first to run it
Leave your email and we will write once, when the repository is public.