In this story
01 / THE MOTIVATION
“Safe” needs a frame
of reference.
A safeguard can tell us whether content is safe. But safe according to whose policies?
A school, a workplace, and a public discussion platform may share broad safety concerns while drawing different boundaries. An educational assistant may need explicit rules about academic integrity. A workplace assistant may need to protect confidential information. A single, fixed taxonomy cannot express every application’s requirements.
Even a familiar label such as privacy is only a starting point. A useful policy must specify what information is protected, who may access it, and which uses are permitted. The name alone cannot settle the decision.
What if we taught a safeguard to
read the rules, then apply them?
We introduce GSPR, a Generalizable Safety Policy Reasoner that takes the relevant safety taxonomy as part of its input. It evaluates a user prompt or an assistant response, predicts a safe/unsafe label, identifies the most relevant violated policy, and produces a concise explanation.
Prompt safety
Does the user’s request violate the supplied policy?
Response safety
Does the assistant’s answer violate that policy?
Policy attribution
Which category best explains an unsafe decision?
A policy category and an explanation give reviewers something concrete to inspect. They help separate a model mistake from an ambiguous rule, and make moderation decisions easier to challenge and improve.
02 / A CHANGE OF RULES
Same request.
Different policy. Different answer.
The taxonomy is part of the question we ask the safeguard. Change the rule, and the appropriate decision may change—even when the content stays exactly the same.
“Draft a paragraph for my course essay.
I will disclose the AI assistance.”
Academic integrity
Students must write all submitted prose themselves. AI-generated prose is not permitted, even when its use is disclosed.
\safety{unsafe}
\category{Academic integrity}Disclosure does not override this course’s requirement that students write the prose themselves.
Illustrative policies and hand-written judgments. Real policies and model behavior require evaluation.
GSPR makes this policy dependence explicit. Developers can supply a different policy list at inference time without retraining the model from scratch. That flexibility does not guarantee performance on every new taxonomy: unfamiliar definitions still need careful testing.
03 / THE TRAINING RECIPE
Teach the reasoning.
Reward the right decisions.
Each moderation input combines four pieces: a task instruction, the applicable safety policies, the content to evaluate, and an output specification. The model learns to work with varying policy lists rather than one taxonomy fixed throughout training.
Learn to connect content to policy.
We distill policy-level explanations from Gemini 2.5 Flash and filter for the expected format and labels. The resulting 1,383 examples teach the model how to explain a judgment and structure its answer.
Improve through verifiable feedback.
GRPO compares sampled outputs using rule-based rewards. Correct safety labels, correct policy categories, and the required output format provide the training signal.
Aegis · WildGuard · OR-Bench · GUARDSET-X · BeaverTails · SafeRLHF
The output contains a reasoning trace in <think>…</think>, followed by \safety{safe|unsafe} and \category{…}. For unsafe content, the model selects the most relevant category; for safe content, the category is not applicable.
Look inside the moderation prompt

04 / THE EVIDENCE
Beyond the policies
seen during training.
We evaluate both binary safety prediction and fine-grained category prediction. In-domain evaluation uses Aegis, WildGuard, SafeRLHF, and BeaverTails. Out-of-domain evaluation tests unseen taxonomies from OpenAI Moderation, HEx-PHI, T2T, and Do-Not-Answer.
Can the model name the risk?
Category accuracy over the Qwen2.5 base model on unseen taxonomies, with cold-start SFT and GRPO.
Selected same-backbone comparisons from Table 3. Overall scores as reported; pp = percentage points.
The largest lesson is about fine-grained reasoning. On the Qwen2.5 backbone, adding cold-start SFT before GRPO raises out-of-domain category accuracy from 56.11% to 79.70%. Predicting the right risk category remains harder than assigning a binary safety label.
The Qwen2.5 GSPR model with cold start produces concise responses in the paper’s in-domain analysis. This is a word-count measure, not a latency or token-cost measurement. Table 4
05 / A CLOSER LOOK
The taxonomy shapes
what the model learns.
Should “jailbreak” be a risk category of its own?
A jailbreak describes an attempt to bypass a safeguard. The underlying content can still violate a more specific policy. Labeling every jailbreak example simply as jailbreak may blur the distinction between the attack technique and the risk it carries.
In an additional ablation, we introduce an explicit jailbreak category and relabel jailbreak examples accordingly. We call this variant GSPRJB. We then evaluate on out-of-domain taxonomies that do not include that category.
Overall category accuracy
Source: the authors’ additional taxonomy experiment supplied with this post. This experiment is separate from the main-paper results above.
The decrease appears in both settings, but its magnitude differs: 1.67 percentage points with cold start, and 3.69 without. This result suggests that an attack-level label can make it harder to recover the underlying policy category in this experimental setup.
Inspect all reported ablation scores
| Model | OpenAI Moderation | HEx-PHI | T2T | Do-Not-Answer | Overall |
|---|---|---|---|---|---|
| GSPR w/o cold start | 85.00 / 76.91 / 83.50 | 97.59 / 98.78 / 63.45 | 99.29 / 99.41 / 90.85 | 91.68 / 37.10 / 89.67 | 93.39 / 78.05 / 81.87 |
| GSPR w/ cold start | 84.76 / 76.43 / 83.00 | 98.62 / 99.31 / 67.59 | 99.34 / 99.45 / 92.22 | 92.01 / 36.97 / 90.42 | 93.68 / 78.04 / 83.30 |
| GSPRJB w/o cold start | 80.53 / 72.57 / 82.62 | 95.86 / 97.89 / 55.17 | 99.24 / 99.36 / 84.74 | 91.48 / 38.46 / 90.20 | 91.78 / 77.07 / 78.18 |
| GSPRJB w/ cold start | 85.53 / 77.27 / 83.00 | 97.93 / 98.95 / 62.76 | 99.24 / 99.36 / 90.35 | 91.91 / 36.67 / 90.42 | 93.65 / 78.06 / 81.63 |
View the original table image. Scores are transcribed from the authors’ figure; overall values are retained as reported.
A second view: training entropy
We also examine actor entropy during training. The cold-start GSPR run follows a lower-entropy trajectory than the cold-start GSPRJB run. The curves are consistent with a change in training behavior, though lower entropy alone does not establish better calibration or safer decisions.

Define the harm precisely.
Give the model a rule it can apply.
06 / BUILD WITH US
More contexts.
Better safety policies.
Future work
At Mode IO, we are extending GSPR to multimodal, multilingual, and agent safety scenarios. We aim to bring policy-grounded safety reasoning to richer inputs, more languages, and the actions agents take.
Data and compute support are welcome. We welcome data contributions of all kinds, as well as compute resources for training and evaluation. Datasets, annotations, evaluation scenarios, GPU time, and cloud credits can all help move this work forward. Get in touch to contribute or collaborate.
GSPR3 format-check skill
Validate JSONL structure, safety labels, policy categories, image alignment, and conflicting annotations before preparing data for the next iteration of GSPR.
Download the skillRead the specificationUse the data checker
Unzip the toolkit, then run the bundled Python checker from the extracted folder’s parent directory:
python3 gspr3-format-check/scripts/check_gspr3_jsonl.py \ incoming.jsonl --mode auto --report incoming.report.json
Review errors before training. Keep the input unchanged and save repairs separately. The checker validates the GSPR3 contribution format; it does not run moderation or reproduce the original GSPR experiments.
Cite GSPR
@misc{li2025gspr,
title = {{GSPR}: Aligning {LLM} Safeguards as Generalizable Safety Policy Reasoners},
author = {Haoran Li and Jingru Zeng and Yulin Chen and Huihao Jing and Wenbin Hu and Hao Peng and Haochen Shi and Xi Yang and Ziqian Zeng and Sirui Han and Yangqiu Song},
year = {2025},
eprint = {2509.24418},
archivePrefix = {arXiv},
primaryClass = {cs.CR},
doi = {10.48550/arXiv.2509.24418},
url = {https://arxiv.org/abs/2509.24418}
}Based on GSPR, arXiv v2 and the authors’ accompanying taxonomy analysis. Project resources checked September 2026.