Custom content moderation policies with a language model

Every platform has rules that a generic moderation list does not cover: counterfeit goods on a marketplace, off-topic posts in a professional community, medical claims in product reviews, recruiting scams in a job board. Trust and safety teams end up with a policy document that only people can apply, and a queue that grows faster than the team.

spinf, a specialized inference provider, runs leading open-weights language models optimized for scoring content at scale. For moderation, that means a policy classifier you write yourself: your rules as questions, a few examples per category, and a probability for each rule on every post. This guide covers how to write the policy, set thresholds per category, clear a moderation backlog and moderate live content.

Write the policy as questions

A custom content moderation policy becomes two kinds of questions:

  • A yes/no question for "does it break the policy?", with the policy summarized in the question: "Does this post break the platform's content policy (harassment, hate, sexual content, violent threats, self-harm, spam or scams)?".
  • A closed question for the category, with your moderation categories as the answers: "Policy category (safe, harassment, hate, sexual, violence, self-harm, or scam):".

You can add a question for any rule of your own ("Does this listing sell counterfeit goods?", "Does this review make a medical claim?"), and ask all of them of the same post in one call. The post is read once, and each extra question costs only its own tokens.

Start each post with a short label, such as User post: or Listing:, so the model knows what it is reading. The answers always come back as one of the options you listed, with its probability: there is no generated text to parse.

Teach the edge cases with examples

A policy is mostly its edge cases: the joke that is not harassment, the polite message that is a scam. Put a few labelled posts in the content before the post to score, covering each category and the cases your moderators argue about. No training run and no model to host.

On 130 synthetic posts across six policy categories and safe content, including edgy-but-safe and polite-but-harmful posts, spinf-12b got:

QuestionNo examplesWith examples
Breaks policy? (yes/no)92% correct99% correct, with 4 examples
Policy category (7)79% correct98% correct, with 8 examples

The posts, questions, example sets and code are public in spinf-benchmarks. Real posts are harder than synthetic ones: teach your policy with examples from your own platform, and measure on a labelled sample of your own content before relying on the numbers.

Set thresholds per category

A probability per rule lets you act differently on different risks:

  • Remove what is clearly over the line: a high threshold, set so that almost nothing removed would be reversed on appeal.
  • Queue for review the borderline band below it, where a moderator decides.
  • Allow the rest.

Set each threshold on your own history: score a few hundred posts your team has already decided, and pick the cut-off per category where the automatic decisions match your moderators closely enough. A stricter category (threats, self-harm) can have a lower threshold for review than a milder one (off-topic). The probabilistic if statements guide covers the mechanics.

Clear the moderation backlog

A backlog, or a periodic sweep of everything published last week, runs on the on-demand API, self-serve. Send posts in batches with many posts per call, compare the scores with past decisions, and work through the queue from the highest probability down.

Because only input is billed, scoring every post is affordable, not just a sample: on spinf-12b, about $0.01 per 1,000 posts with no examples, and about $0.08 with 8 examples in the prompt, at $0.09 per million input tokens. The scores are free, and the first 50M tokens are on us.

Moderate live posts

To check posts before they are published, run the same questions and thresholds on a reserved endpoint: capacity reserved for you and sized to your traffic. The on-demand API is built for throughput rather than real time. A good path is to calibrate on your history with the on-demand API first, then move to the reserved endpoint once the thresholds are set.

Good to know

  • UGC moderation of other formats: text is available to everyone today. Images and audio are in preview, on request.
  • Data: content sent for scoring is processed in memory and not stored by default, and it is never used to train models.
  • Other checks in the same call: classify the intent of a message, or check it for prompt injection, while you moderate it. See the content moderation use case.

Try it

See the content moderation use case for a live example and the cost at scale. Related: probabilistic if statements and zero-shot and few-shot classification.

See the use cases

Raw Markdown for agents: /guides/custom-content-moderation.md · /llms.txt