Documentation · Prompt-packs
Prompt-packs
Beta. Packs and their format are in beta and may change. Pin a pack's
versionto keep scores comparable.
A prompt-pack is a ready-made set of questions that spinf builds, validates and versions for one job. You name it in your call, and it is scored in the same pass as your own questions: one content, one call, one bill. A pack returns a calibrated 0–1 score for each of its categories and a decision.
The first pack is spinf/moderation: 22 categories of user-generated text (sexual content, child safety,
violence, self-harm, hate, harassment, fraud, …) and a safe / unsafe decision.
Try it without code in the playground: paste a post, pick the pack, and see every category next to your own questions.
List the packs
GET https://api.spinf.com/v1/packs
No key needed. One entry per pack:
{ "packs": [ {
"id": "spinf/moderation", "name": "Content moderation", "status": "available",
"version": "0.1.0", "versions": ["0.1.0"], "model": "spinf-12b", "media": ["text"],
"contexts": [
{ "id": "ai", "template": "A user sent the following message to an AI assistant.\n\nUser: {X}" },
{ "id": "forum", "template": "The following message was posted by a user on an online platform:\n\n\"{X}\"" } ],
"categories": [
{ "id": "sexual", "label": "Sexual content", "threshold": … },
{ "id": "spam", "label": "Spam and advertising", "threshold": null }, … ],
"decision": { "mode": "micro_layer", "modes": ["max", "per_category", "micro_layer"],
"fallback": "per_category", "max_threshold": … },
"queries": 58, "prompts": 106, "query_tokens_per_input": 1626 } ] }
| Field | Meaning |
|---|---|
status | available, or coming_soon (not callable yet). |
version, versions | The latest version, and every version you can pin. |
model | The only model the pack runs on. |
media | The content kinds the pack scores. spinf/moderation 0.1.0: text. |
contexts | How your content is presented to the model ({X} is your content). The first is the default. |
categories | Each category's id, label and default threshold. null = the category is scored and reported, but never decides. |
decision | The default decision mode, the modes you can pick, and the default threshold of max. |
query_tokens_per_input | The pack's billed question tokens per content (constant). |
The … values are left out on purpose: read the thresholds from the listing rather than copying them into your code,
because each new version can change them.
Download a pack
GET https://api.spinf.com/v1/packs/spinf/moderation?version=0.1.0
Authorization: Bearer ssk-...
Returns the pack file as it is: its questions, contexts, calibration, categories and decision layer. Use it to see
exactly what is scored. The slash of the id is part of the path; without version you get the latest. Call it from
your server, like the scoring API.
Request
Add scoring.packs next to (or instead of) scoring.queries:
curl https://api.spinf.com/v1/score \
-H "Authorization: Bearer $SPINF_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"model": "spinf-12b",
"messages": [ { "role": "user", "content": [
{ "type": "text", "text": "Guaranteed 300% returns in 7 days!! DM me your wallet seed phrase." } ] } ],
"scoring": {
"packs": [ { "id": "spinf/moderation", "context": "forum" } ],
"queries": [
{ "id": "on_topic", "template": "\n\nIs this post about cooking?\nAnswer:{?}", "options": [" yes", " no"] }
]
}
}
EOF
| Field | Rules |
|---|---|
id | Required, from GET /v1/packs. Each pack at most once per call. |
version | Optional, default the latest. Pin it to keep scores comparable over time. |
context | Optional, one of the pack's contexts (default the first). |
decision.mode | micro_layer, per_category or max (default: the pack's, micro_layer for moderation). See decisions. |
decision.threshold | The single threshold of max (0–1) or of micro_layer (the layer's value). Not with per_category. |
decision.thresholds | Per-category overrides, { "harassment": 0.8, "spam": 0.7 } (0–1, or null to report only). They set flagged and the per_category decision. |
include_queries | Default false. true also returns the pack's raw question results in results[].queries. |
How a pack changes the call
- One read for everything. Each content is wrapped in the pack's context (for
forum: The following message was posted by a user on an online platform: "…"), and your own questions in the same call read the wrapped content too. That shared read is what puts the pack and your questions in one pass, with the content billed once. - Text only, for now. A pack scores the content kinds in its
media. Withspinf/moderation0.1.0, an image, audio or video part is refused (pack_media_not_supported): send media in a separate call. - Case. A pack is calibrated with its own case rule (moderation: case-sensitive), and it applies to the whole
call. Leave
scoring.case_insensitiveout when you send a pack; the other value is refused (pack_conflict). - Model. The pack's
modelonly (spinf-12bfor moderation). - Ids. The pack's question ids start with
tax.andgen.: don't use those prefixes for your own questions. Pack questions run without the empty floor, whatever your default.
Response
Each result has packs beside queries. For the message "How do I get into my ex's Instagram account without
her knowing?" in the ai context (version 0.1.0; scores and thresholds change from one version to the next):
{ "input_id": "0",
"queries": [ … your questions … ],
"packs": [ { "id": "spinf/moderation", "version": "0.1.0", "context": "ai",
"categories": {
"cyber": { "score": 0.921911, "percentile": 1.0, "flagged": true },
"sexual": { "score": 0.554439, "percentile": 0.94762, "flagged": false },
"spam": { "score": 0.011345, "percentile": 0.773997, "flagged": null }, … },
"decision": { "unsafe": true, "mode": "micro_layer", "score": 1.206813, "threshold": … } } ] }
score: the category's calibrated score, 0–1, comparable across categories.percentile: where that score falls on everyday content (0–1):0.95= higher than 95% of normal traffic. A high percentile with a modest score means the content is unusual for the category without crossing its threshold, a useful signal for review queues.flagged:score >= threshold;nullfor categories that are reported only.decision: safe or unsafe, by the mode you picked (below).
Decisions and thresholds
| Mode | decision | Unsafe when |
|---|---|---|
micro_layer (default) | {unsafe, mode, score, threshold} | A small layer trained on labelled moderation data, over every category, is over its threshold. |
per_category | {unsafe, mode, categories} | Any deciding category is over its own threshold; categories lists them. |
max | {unsafe, mode, score, threshold, category} | The highest deciding category is over one threshold; category names it. |
- Start with the default. Use
per_categorywhen your policy needs its own threshold per category (stricter on harassment, looser on profanity), andmaxwhen you want one simple threshold. - Set thresholds on your own labelled sample: score a few hundred items you have already decided, and choose each threshold where the errors you can live with are.
- Categories with a
nullthreshold (in 0.1.0: spam, gibberish and the four topics) are scored and reported but never decide, unless you give them a threshold indecision.thresholds. The topics (alcohol, tobacco, gambling, medication) say what the content is about, not whether it is harmful, and are not calibrated like the others.
Moderation categories
id | Category |
|---|---|
sexual | Sexual content |
child | Child sexual exploitation and grooming |
violence | Violence and threats |
selfharm | Self-harm and suicide |
hate | Hate against protected groups |
harassment | Harassment and bullying |
extremism | Extremism and terrorism |
weapons | Weapons |
drugs | Illegal drugs |
crime | Other crime |
fraud | Fraud and scams |
cyber | Hacking and malware |
privacy | Private personal information |
spam | Spam and advertising |
profanity | Profanity |
graphic | Gore and graphic violence |
gibberish | Gibberish |
alcohol | Alcohol (topic) |
tobacco | Tobacco and vaping (topic) |
gambling | Gambling (topic) |
medication | Medication (topic) |
general | Overall harmfulness (general reads) |
The default thresholds, and which categories decide, are in the listing.
Billing
No surcharge: a pack's questions are billed like your own. spinf/moderation adds 1,626 question tokens per
content, and the content (with its context) is billed once, shared with your questions: about 1,710 tokens for a
typical message, so about 1.7 billion tokens per million messages at $0.09 per million. See
Limits & pricing.
Errors
| Status | code | Meaning |
|---|---|---|
| 400 | unknown_pack, pack_not_available | No such pack, or it is still coming_soon. |
| 400 | unknown_pack_version | No such version of the pack. |
| 400 | unknown_category | A category in decision.thresholds that the pack doesn't have. |
| 400 | pack_media_not_supported | A part the pack doesn't score (moderation 0.1.0: anything but text). |
| 400 | pack_conflict | scoring.case_insensitive differs from the pack's. Leave it out. |
| 400 | pack_model_mismatch | The call's model is not the pack's. |
| 400 | duplicate_id | A pack twice, or one of your question ids is a pack's (tax.…, gen.…). |
| 400 | invalid_field | A pack field of the wrong type or shape, e.g. decision.threshold with per_category. param names it. |
| 404 | unknown_pack, unknown_pack_version | Download only: no such pack or version. |
The other errors are those of the Scoring API.
Raw Markdown for agents: /docs/prompt-packs.md · /llms.txt