Aspect-based sentiment analysis (ABSA) for reviews and surveys

"Delivery was two days late and the box arrived dented, but support sent a replacement right away and the blender itself is excellent." One overall label for this review, positive, negative or mixed, loses what matters. Aspect-based sentiment analysis (ABSA) asks how the customer feels about each aspect separately: negative on delivery and packaging, positive on support and on the product.

This guide shows how to run aspect sentiment with spinf, how to turn it into numbers you can count, and what it measured on public and synthetic reviews.

One template, every aspect

spinf asks questions of the content with templates. A placeholder such as {aspect} asks the same question for every aspect you name, all over one read of the review:

{
  "id": "aspect",
  "template": "\n\nHow does the customer feel about the {aspect} (positive, negative, or not mentioned)?\nAnswer:{?}",
  "options": [" positive", " negative", " not mentioned"],
  "combinations": [
    { "aspect": "delivery" }, { "aspect": "packaging" }, { "aspect": "customer support" },
    { "aspect": "product quality" }, { "aspect": "price" }
  ]
}

For each aspect you get the probability of each answer, and they sum to 1. The option "not mentioned" matters: most reviews talk about two or three aspects, and a forced positive or negative on the others would add noise to every count. Add an overall question in the same call ("Overall, the review is (positive, mixed, or negative):") and any other code you need. The review is read once, and each extra question costs only its own tokens.

The aspects are yours: delivery and price for a shop, onboarding and billing for a software product, staff and cleanliness for a hotel. There is no fixed list and no training run.

From scores to a feedback dashboard

Probabilities over your own answers add up cleanly:

  • Shares. The mean of p(negative) for "delivery" over a month of reviews estimates the share of reviews negative about delivery, without choosing a cut-off for each review.
  • Segments and trends. Group by product, region, channel or survey wave, and follow each aspect over time.
  • The reviews behind a number. Sort by the probability to read the reviews that drive a change.

Each answer also comes with its floor, the same question asked with empty content. With no examples in the prompt, reading the lift over the floor corrects a question that leans towards one answer. See Calibration.

Measured on public and synthetic reviews

These figures come from our public benchmark, on spinf-12b (Gemma 4 12B, optimized for scoring):

TaskTest setResult
Overall sentiment (3)125 synthetic reviews96% correct with 8 examples
Sentiment per aspect625 synthetic aspect labels94% correct with 8 examples; 84% with none (calibrated)
Polarity300 public product reviews97% correct with no examples (calibrated)
Aspect sentimentSemEval-2014 restaurants, 1,166 public labels89% correct with 4 examples; 83% with none (calibrated)

On SemEval-2014 the larger spinf-31b (Gemma 4 31B, on request) reached 90% with 8 examples (one example set). "Examples" are labelled reviews placed in the prompt, never the reviews being scored; the spinf-12b figures are a mean over three different example sets. The synthetic reviews were labelled by a language model and checked blind by a second model.

The public datasets are the better guide: synthetic reviews are cleaner than real feedback. Measure on a coded sample of your own responses before relying on the numbers. The test sets, questions and code are public: spinf-benchmarks on GitHub.

When to add examples

Overall polarity of a clear review needs no examples. Aspects are harder: the model has to find where the review talks about each aspect, and decide what counts as, say, "support" in your business. There, a few labelled reviews in the prompt help most (84% to 94% on the synthetic set). Write them in the same format as the question, cover every answer at least once, and include reviews that mention an aspect only in passing. See Zero-shot and few-shot classification.

With examples in the prompt, read p as it is: the floor is computed without your examples, so it no longer describes the question the model sees.

Cost for every response

Only input is billed, at $0.09 per million tokens on spinf-12b. A 120-token answer with 8 questions and aspects, sent in batches, is about 350 billed tokens, so a million survey answers or reviews cost about $32. That is cheap enough to score every response instead of a hand-read sample, and to re-score the archive when you add an aspect. The first 50M tokens are free.

Try it

See the use cases

Raw Markdown for agents: /guides/aspect-based-sentiment-analysis.md · /llms.txt