Coding open-ended survey responses with a language model
Open-ended questions are where respondents say what the closed questions missed: why they gave a low NPS score, what went wrong with a delivery. They are also the most expensive part of a survey to analyse: coding the verbatims by hand means reading every answer against a codebook, so most teams code a sample.
spinf, a specialized inference provider, runs leading open-weights language models optimized for this kind of job: the same questions asked of every response, with a probability for each answer you allow. This guide shows how to turn a codebook into questions, check the result against your own coders, and re-code an archive when the coding frame changes.
From codebook to questions
A codebook (or coding frame) is a list of codes with a definition for each. With spinf, each code becomes a closed question, and the answers are the words you want back:
- One yes/no question per code when a response can carry several codes: "Does the customer mention a late or missed delivery?", "Does the customer mention the price?", "Does the answer suggest a new feature?".
- One closed question for a single-choice code, such as the main theme or the overall tone: "Overall, the answer is (positive, mixed, or negative):".
- A template with a placeholder for codes that repeat across topics. One template, "How does the customer feel about the {aspect} (positive, negative, or not mentioned)?", asks the same question for delivery, packaging, support, quality and price in one call. See the aspect-based sentiment guide.
Put the definition of the code in the question, the way you would write it in the codebook. Start each response with a
short label, such as Survey answer: or NPS comment:, so the model knows what it is reading.
{
"id": "delivery_issue",
"template": "\n\nDoes the customer mention a late, missed or damaged delivery?\nAnswer:{?}",
"options": [" yes", " no"]
}
For each response and each code, spinf returns the probability of each answer you listed; nothing is generated.
Teach the subtle codes with examples
Some codes are obvious ("mentions the price"); others depend on how your team draws the line ("complaint vs suggestion"). For those, put a few coded responses in the content before the one to score: 4 to 8 examples, covering each answer and the cases your coders disagree on. No training run is needed.
In our public benchmark, examples made the largest difference on per-aspect sentiment: 94% correct with 8 examples, against 84% with no examples (calibrated), on 625 labels from 125 synthetic reviews. Overall sentiment (positive, mixed or negative) was 96% correct with 8 examples on the same reviews. The test sets and the code are on GitHub. Synthetic answers are cleaner than real ones, so measure on your own data, as described below.
Check agreement with your coders
Before coding a whole survey, measure spinf the way you would measure a new human coder:
- Take a sample of responses that your team has already coded: 100 to 300 is usually enough.
- Score them with your questions, with and without examples, and with two or three wordings of the subtle codes.
- For each code, compare with your coders: agreement, and the codes where the two disagree. Read the disagreements; they often show a code whose definition was ambiguous for people too.
- Keep the version that agrees best, and choose a threshold per code (for example, count a code when its probability is above 0.5, and send answers between 0.4 and 0.6 to a person).
A round on 300 responses costs cents, and fits well within the free tokens of a new account. Once you are happy, keep the questions, answers and examples fixed, so that results compare from one wave to the next.
Count, segment and track
Because every code is a probability over your own answers, the results add up cleanly: the share of responses that mention delivery, per segment, per wave, per product. You can count with a threshold, or sum the probabilities to get an expected count that does not depend on where you put the cut. Each answer also comes with its floor, the same question asked with no response, so different codes can be compared; see Calibration.
Re-code the archive when the codebook changes
Codebooks change: a new theme appears, two codes merge, a definition is tightened. With manual coding, the old waves keep the old frame. With spinf, re-scoring the archive is a batch job: send past responses with the new questions, many answers per call, and the whole history is coded the same way.
The cost stays small because only input is billed, and each extra question costs only its own tokens. With 120-token answers and 8 questions, a response is about 350 billed tokens in batched calls: 1 million responses is about 350 million tokens, or about $31.50 at $0.09 per million tokens. The scores are free, and the first 50M tokens are on us.
Any text works, from survey verbatims to NPS comments and app store reviews. Content sent for scoring is processed in memory and not stored by default, and it is never used to train models.
Try it
See the survey & review analysis use case for a live example and the cost at scale, and Writing questions for wording and examples. Related: aspect-based sentiment analysis and zero-shot and few-shot classification.
Raw Markdown for agents: /guides/open-ended-survey-coding.md · /llms.txt