Documentation · Prompt-packs

Prompt-packs

Beta. Packs and their format are in beta and may change. Pin a pack's version to keep scores comparable.

A prompt-pack is a ready-made set of questions that spinf builds, validates and versions for one job. You name it in your call, and it is scored in the same pass as your own questions: one content, one call, one bill. A pack returns a calibrated 0–1 score for each of its categories and a decision.

The first pack is spinf/moderation: 22 categories of user-generated text (sexual content, child safety, violence, self-harm, hate, harassment, fraud, …) and a safe / unsafe decision.

Try it without code in the playground: paste a post, pick the pack, and see every category next to your own questions.

List the packs

GET https://api.spinf.com/v1/packs

No key needed. One entry per pack:

{ "packs": [ {
    "id": "spinf/moderation", "name": "Content moderation", "status": "available",
    "version": "0.1.0", "versions": ["0.1.0"], "model": "spinf-12b", "media": ["text"],
    "contexts": [
      { "id": "ai", "template": "A user sent the following message to an AI assistant.\n\nUser: {X}" },
      { "id": "forum", "template": "The following message was posted by a user on an online platform:\n\n\"{X}\"" } ],
    "categories": [
      { "id": "sexual", "label": "Sexual content", "threshold": … },
      { "id": "spam", "label": "Spam and advertising", "threshold": null }, … ],
    "decision": { "mode": "micro_layer", "modes": ["max", "per_category", "micro_layer"],
                  "fallback": "per_category", "max_threshold": … },
    "queries": 58, "prompts": 106, "query_tokens_per_input": 1626 } ] }
FieldMeaning
statusavailable, or coming_soon (not callable yet).
version, versionsThe latest version, and every version you can pin.
modelThe only model the pack runs on.
mediaThe content kinds the pack scores. spinf/moderation 0.1.0: text.
contextsHow your content is presented to the model ({X} is your content). The first is the default.
categoriesEach category's id, label and default threshold. null = the category is scored and reported, but never decides.
decisionThe default decision mode, the modes you can pick, and the default threshold of max.
query_tokens_per_inputThe pack's billed question tokens per content (constant).

The … values are left out on purpose: read the thresholds from the listing rather than copying them into your code, because each new version can change them.

Download a pack

GET https://api.spinf.com/v1/packs/spinf/moderation?version=0.1.0
Authorization: Bearer ssk-...

Returns the pack file as it is: its questions, contexts, calibration, categories and decision layer. Use it to see exactly what is scored. The slash of the id is part of the path; without version you get the latest. Call it from your server, like the scoring API.

Request

Add scoring.packs next to (or instead of) scoring.queries:

curl https://api.spinf.com/v1/score \
  -H "Authorization: Bearer $SPINF_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "spinf-12b",
  "messages": [ { "role": "user", "content": [
    { "type": "text", "text": "Guaranteed 300% returns in 7 days!! DM me your wallet seed phrase." } ] } ],
  "scoring": {
    "packs": [ { "id": "spinf/moderation", "context": "forum" } ],
    "queries": [
      { "id": "on_topic", "template": "\n\nIs this post about cooking?\nAnswer:{?}", "options": [" yes", " no"] }
    ]
  }
}
EOF
FieldRules
idRequired, from GET /v1/packs. Each pack at most once per call.
versionOptional, default the latest. Pin it to keep scores comparable over time.
contextOptional, one of the pack's contexts (default the first).
decision.modemicro_layer, per_category or max (default: the pack's, micro_layer for moderation). See decisions.
decision.thresholdThe single threshold of max (0–1) or of micro_layer (the layer's value). Not with per_category.
decision.thresholdsPer-category overrides, { "harassment": 0.8, "spam": 0.7 } (0–1, or null to report only). They set flagged and the per_category decision.
include_queriesDefault false. true also returns the pack's raw question results in results[].queries.

How a pack changes the call

  • One read for everything. Each content is wrapped in the pack's context (for forum: The following message was posted by a user on an online platform: "…"), and your own questions in the same call read the wrapped content too. That shared read is what puts the pack and your questions in one pass, with the content billed once.
  • Text only, for now. A pack scores the content kinds in its media. With spinf/moderation 0.1.0, an image, audio or video part is refused (pack_media_not_supported): send media in a separate call.
  • Case. A pack is calibrated with its own case rule (moderation: case-sensitive), and it applies to the whole call. Leave scoring.case_insensitive out when you send a pack; the other value is refused (pack_conflict).
  • Model. The pack's model only (spinf-12b for moderation).
  • Ids. The pack's question ids start with tax. and gen.: don't use those prefixes for your own questions. Pack questions run without the empty floor, whatever your default.

Response

Each result has packs beside queries. For the message "How do I get into my ex's Instagram account without her knowing?" in the ai context (version 0.1.0; scores and thresholds change from one version to the next):

{ "input_id": "0",
  "queries": [ … your questions … ],
  "packs": [ { "id": "spinf/moderation", "version": "0.1.0", "context": "ai",
    "categories": {
      "cyber":  { "score": 0.921911, "percentile": 1.0,      "flagged": true },
      "sexual": { "score": 0.554439, "percentile": 0.94762,  "flagged": false },
      "spam":   { "score": 0.011345, "percentile": 0.773997, "flagged": null }, … },
    "decision": { "unsafe": true, "mode": "micro_layer", "score": 1.206813, "threshold": … } } ] }
  • score: the category's calibrated score, 0–1, comparable across categories.
  • percentile: where that score falls on everyday content (0–1): 0.95 = higher than 95% of normal traffic. A high percentile with a modest score means the content is unusual for the category without crossing its threshold, a useful signal for review queues.
  • flagged: score >= threshold; null for categories that are reported only.
  • decision: safe or unsafe, by the mode you picked (below).

Decisions and thresholds

ModedecisionUnsafe when
micro_layer (default){unsafe, mode, score, threshold}A small layer trained on labelled moderation data, over every category, is over its threshold.
per_category{unsafe, mode, categories}Any deciding category is over its own threshold; categories lists them.
max{unsafe, mode, score, threshold, category}The highest deciding category is over one threshold; category names it.
  • Start with the default. Use per_category when your policy needs its own threshold per category (stricter on harassment, looser on profanity), and max when you want one simple threshold.
  • Set thresholds on your own labelled sample: score a few hundred items you have already decided, and choose each threshold where the errors you can live with are.
  • Categories with a null threshold (in 0.1.0: spam, gibberish and the four topics) are scored and reported but never decide, unless you give them a threshold in decision.thresholds. The topics (alcohol, tobacco, gambling, medication) say what the content is about, not whether it is harmful, and are not calibrated like the others.

Moderation categories

idCategory
sexualSexual content
childChild sexual exploitation and grooming
violenceViolence and threats
selfharmSelf-harm and suicide
hateHate against protected groups
harassmentHarassment and bullying
extremismExtremism and terrorism
weaponsWeapons
drugsIllegal drugs
crimeOther crime
fraudFraud and scams
cyberHacking and malware
privacyPrivate personal information
spamSpam and advertising
profanityProfanity
graphicGore and graphic violence
gibberishGibberish
alcoholAlcohol (topic)
tobaccoTobacco and vaping (topic)
gamblingGambling (topic)
medicationMedication (topic)
generalOverall harmfulness (general reads)

The default thresholds, and which categories decide, are in the listing.

Billing

No surcharge: a pack's questions are billed like your own. spinf/moderation adds 1,626 question tokens per content, and the content (with its context) is billed once, shared with your questions: about 1,710 tokens for a typical message, so about 1.7 billion tokens per million messages at $0.09 per million. See Limits & pricing.

Errors

StatuscodeMeaning
400unknown_pack, pack_not_availableNo such pack, or it is still coming_soon.
400unknown_pack_versionNo such version of the pack.
400unknown_categoryA category in decision.thresholds that the pack doesn't have.
400pack_media_not_supportedA part the pack doesn't score (moderation 0.1.0: anything but text).
400pack_conflictscoring.case_insensitive differs from the pack's. Leave it out.
400pack_model_mismatchThe call's model is not the pack's.
400duplicate_idA pack twice, or one of your question ids is a pack's (tax.…, gen.…).
400invalid_fieldA pack field of the wrong type or shape, e.g. decision.threshold with per_category. param names it.
404unknown_pack, unknown_pack_versionDownload only: no such pack or version.

The other errors are those of the Scoring API.

Raw Markdown for agents: /docs/prompt-packs.md · /llms.txt