Guides to classification and scoring with language models
spinf is a specialized inference provider: an API to leading open-weights language models, optimized for analysis at scale. It returns the probability of each answer you allow, for every question you ask about your content, and generates no text. These guides explain the ideas behind that, and how to apply them to common jobs.
Concepts
- Probabilities from a language model: confidence scores over the answers you allow, the empty-content floor, and calibrated scores you can compare.
- Probabilistic if statements: decisions in code on confidence thresholds, with a band for human review, chosen on your own history.
- Zero-shot and few-shot classification: classification with no fine-tuning, and what 4 or 8 labelled examples in the prompt change, measured.
- Scoring without output tokens: why scoring costs less than generating, and what a million documents cost when only input is billed.
Jobs
- Entity-level sentiment for news: sentiment toward each company an article names, aggregated into daily indicators.
- Aspect-based sentiment analysis: sentiment per aspect for reviews and survey answers, measured on SemEval-2014.
- Coding open-ended survey responses: your codebook as questions, every verbatim coded.
- Custom content moderation policies: your rules as questions, a few examples, thresholds per category.
- Intent classification and routing: intents as answer options, thresholds for automation, and many intents in two steps.
Reference
The documentation covers the API itself: the quickstart, the scoring API, writing questions and calibration. The benchmark behind the figures quoted in these guides is public: spinf-benchmarks.
Raw Markdown for agents: /guides/index.md · /llms.txt