Intent classification and ticket routing with confidence thresholds
Support teams sort every incoming ticket, chat and email: which queue it belongs to, whether it is urgent, whether a bot or a help article can resolve it. Intent classification (or intent detection) does this sorting automatically, and support ticket triage works when the automation knows when it is unsure.
spinf, a specialized inference provider, runs leading open-weights language models optimized for classification at scale. For each message it returns the probability of each answer to your routing questions, so you can automate the clear cases and send the uncertain ones to a person, at a confidence threshold you choose.
Intents as answer options
Each routing decision is a closed question, and the intents, queues or categories are its answers:
{
"id": "intent",
"template": "\n\nIntent (login problem, outage, billing, cancellation, or feature request):{?}",
"options": [" login problem", " outage", " billing", " cancellation", " feature request"]
}
Add the other decisions as questions on the same message: "Is this urgent (business blocked, security, legal threat, data loss, or a deadline within 24 hours)?", "Can a bot resolve this?", "Does it need a human agent?". The message is read once, and each extra question costs only its own tokens.
The answer is always one of your labels, with its probability: no free text to parse and no invented intents. Start
each message with a short label, such as Customer support message: or Customer email:, so the model knows what it
is reading. The same approach works for email classification and for chat messages.
Accuracy, measured
We measured ticket classification on 130 synthetic support tickets, labelled by a language model and checked blind by a second model, with no examples and with 8 labelled example tickets placed in the prompt (never the tickets being scored):
- Topic (5 topics): 93% correct with 8 examples; 87% with no examples (calibrated). The larger spinf-31b, on request, gets 95% with no examples.
- Can a bot resolve it? 93% correct with 8 examples; 75% with no examples (calibrated).
- Urgent? 92% correct with 8 examples; 85% with no examples. An urgent ticket outranks a non-urgent one 98% of the time (AUC 0.98).
- Needs a human? 87% correct with 8 examples, and 87% with no examples.
The tickets, questions, example sets and code are public in spinf-benchmarks. Synthetic tickets are cleaner than real ones: expect lower numbers on your data, and measure them on your own history before going live.
The examples are the biggest lever. Put 4 to 8 of your own labelled tickets in the content, covering every intent and the cases your team gets wrong, and the scores move closer to your team's decisions. No training run is needed. See zero-shot and few-shot classification.
Choose a confidence threshold
A probability per intent turns routing into a threshold decision:
- Score a labelled sample of past tickets (a few hundred is usually enough).
- For each question, find the probability above which the automatic decision matches your team closely enough.
- Above the threshold, route automatically; below it, send the ticket to a person.
You can use different thresholds per queue: a strict one for tickets a bot will answer on its own, a looser one for tickets that only change their queue. The probabilistic if statements guide covers thresholds, the review band and how to measure the trade-off.
Many intents: two steps
Intent classification works best with a manageable list: a few to a few dozen intents, each named clearly. For a long list, use hierarchical intents and ask two questions: first the broad category ("billing, technical, account, shipping, or cancellation"), then the intent within that category, with a few labelled examples for each. Measure both steps on a labelled sample of your own messages.
Connect it to your helpdesk
spinf is an API. Call it from your helpdesk's webhooks or automation rules (for example in Zendesk, Intercom or Freshdesk) with the ticket text, and write the answer back as a tag, a priority or a queue.
Calibrate first, then route live
- Calibrate on your history. Score last year's tickets on the on-demand API, with your free tokens, compare with what your team decided, and pick your thresholds.
- Go live on a reserved endpoint. Routing tickets as they arrive runs on capacity reserved for you: the on-demand API is built for throughput rather than real time.
- Keep checking. Re-score a sample of new tickets in bulk every week to catch drift.
Cost
Only input is billed, at $0.09 per million input tokens on spinf-12b, and the scores are free. With 8 labelled examples in the prompt, scoring costs about $0.12 per 1,000 tickets. The first 50M tokens are on us. Ticket content is processed in memory and not stored by default, and it is never used to train models.
Try it
See the ticket & message routing use case for a live example, or the classification API for other jobs. Related: probabilistic if statements and probabilities from a language model.
Raw Markdown for agents: /guides/intent-classification.md · /llms.txt