Put every product, conversation and ad in the right category
Catalogues, marketplace listings, chat and call transcripts, ad creatives: ask plain-language questions about each item (the category, the policy checks, the tags your team uses) and get the probability of each answer. We run Gemma 4 12B, optimized to read the content once and answer every question from that read: a million items cost about $23 to classify.
1M items for about $23
On spinf-12b, our optimized Gemma 4 12B, only input is billed, at $0.09 per million tokens: the content once per call, each question's own tokens, and one empty floor per question. The scores are free.
Combine use cases in the same call
Check the same listing, message or ad against your content policy, or route the conversation to the right queue: it is read once, and each extra question adds only its own tokens.
Assumes 150-token listings or messages, 4 questions, batched calls. The first 50M tokens are free. Gemma is a trademark of Google LLC.
$0.09 per million input tokens · the scores are free
Calibrated answers for content classification, at any scale
Your taxonomy, no training
Product categories, conversation reasons, ad verticals, the tags your team uses: write them as questions with the answers you allow. No labelled dataset and no model to train.
Category and checks from one read
The item is read once and every question is answered from that read: the category, the policy checks and each tag add only their own tokens, not another copy of the content.
Cheap enough for the whole catalogue
Classify every listing, conversation and ad, including the backlog, not a sample: on spinf-12b, about $0.02 per 1,000 items with four questions each. Re-run it whenever the taxonomy changes.
One listing, a category and two checks
A category among seven and two yes/no marketplace checks, answered together. The bars show p; the tick shows the floor, the same question with no listing.
Marketplace listing: Stainless steel insulated water bottle, 32 oz, keeps drinks cold 24h / hot 12h. Leak-proof lid, fits most car cup holders. Opened but never used, comes in original box. Great for hiking and the gym. Local pickup or shipping.
Category (electronics, home and kitchen, fashion, beauty, toys, sports and outdoors, or automotive):
electronics
0.03
home and kitchen
0.24
fashion
0.05
beauty
0.01
toys
0.01
sports and outdoors
0.66
automotive
0.00
Does the listing make a health or medical claim?
yes
0.23
no
0.78
Is this item prohibited on a general marketplace (weapons, drugs, counterfeit goods or adult products)?
yes
0.04
no
0.95
import os
import requests
resp = requests.post(
"https://api.spinf.com/v1/score",
headers={"Authorization": f"Bearer {os.environ['SPINF_API_KEY']}"},
json={
"model": "spinf-12b",
"messages": [
{
"role": "user",
"content": "Marketplace listing:\nStainless steel insulated water bottle, 32 oz, keeps drinks cold 24h / hot 12h. Leak-proof lid, fits most car cup holders. Opened but never used, comes in original box. Great for hiking and the gym. Local pickup or shipping."
}
],
"scoring": {
"queries": [
{
"id": "category",
"template": "\n\nCategory (electronics, home and kitchen, fashion, beauty, toys, sports and outdoors, or automotive):{?}",
"options": [
" electronics",
" home and kitchen",
" fashion",
" beauty",
" toys",
" sports and outdoors",
" automotive"
]
},
{
"id": "health_claim",
"template": "\n\nDoes the listing make a health or medical claim?\nAnswer:{?}",
"options": [
" yes",
" no"
]
},
{
"id": "prohibited",
"template": "\n\nIs this item prohibited on a general marketplace (weapons, drugs, counterfeit goods or adult products)?\nAnswer:{?}",
"options": [
" yes",
" no"
]
}
]
}
},
)
for query in resp.json()["results"][0]["queries"]:
for combo in query["combinations"]:
print(query["id"], combo["values"], [(o["text"], o["p"], o["floor_p"]) for o in combo["options"]])A live call to spinf-12b, with no examples. The probabilities show the doubt too: a water bottle for hiking and the gym is mostly sports and outdoors, partly home and kitchen. Set a threshold per question, and send the uncertain items to a person or to a second question.
The API call behind this example, the docs and the benchmark code are easier to read on a larger screen. Start free now, and pick it up on your laptop.
What Gemma 4 12B gets right, with no examples
Examples of results with spinf-12b, our optimized Gemma 4 12B, from plain-language questions and no labelled examples in the prompt, on synthetic test sets written for the job and checked blind by a second model. The reading is Gemma 4; what spinf adds is the scale: every item, every question, one read, at a fraction of the usual cost.
correct topic of a customer message (5)
correct policy category (7)
correct main topic (5)
Synthetic sets are cleaner than real content. A handful of well-defined categories per question works best: for a large taxonomy, ask in two levels (the department, then the categories within it). Labelled examples in the prompt raise the numbers further; measure on a sample of your own items before relying on them.
Test sets, questions and code on GitHub: reproduce it with your own keyThe same questions over product photos and video ads
In preview, on request: images, and video sent as frames, are read like text, and every question shares the one read. On 207 Super Bowl TV commercials (2000–2020), Gemma 4 12B named the brand among ten from frames of the ad, with no examples.
correct brand (10) · TV commercials
Images and video are in preview and enabled per account on request; formats and limits may change. Talk to us to try them on your catalogue or your ads.
The questions matter: iterate on yours
As with any language model, the wording of the questions and a few labelled examples change the results a lot. With spinf that is cheap to get right: try a few wordings and example sets on a few hundred of your own items for cents, and keep the best. How to write questions
From questions to numbers
Self-serve on the on-demand API: sign up, create a key and send your first batch in minutes.
1
Write the taxonomy as questions
One closed question per dimension (category, reason, vertical) with the answers you allow, and a yes/no question per tag or policy check.
2
Test on a labelled sample
Score a few hundred items you have already classified, with your free tokens. Try a few wordings, add examples where a category is subtle, and keep the best.
3
Classify the whole catalogue
Send the backlog in batches on the on-demand API, then new items as they arrive. Keep the probabilities: they tell you which items to check.
Content classification: questions, answered
More in the docs and the guides: zero-shot and few-shot classification and decision models, or .
Same engine, other questions
The first 50M tokens are on us
Sign up, create a key and score your first batch in minutes.