01 POLICY-AWARE AI SAFETY

Safe according
to whose
policies?

GSPR: Towards Generalizable
Safety Policy Reasoners

The rules change across applications.
A safeguard should be able to reason with them.

THE IDEA IN ONE PICTUREFig. 01
[ Flexible policies. Inspectable decisions. ]
ONE MODEL.
MANY POLICY WORLDS.
19training taxonomies
167safety policies
2moderation tasks
Inside the researchPaper · Section 3
In this story

01 / THE MOTIVATION

“Safe” needs a frame
of reference.

A safeguard can tell us whether content is safe. But safe according to whose policies?

A school, a workplace, and a public discussion platform may share broad safety concerns while drawing different boundaries. An educational assistant may need explicit rules about academic integrity. A workplace assistant may need to protect confidential information. A single, fixed taxonomy cannot express every application’s requirements.

Even a familiar label such as privacy is only a starting point. A useful policy must specify what information is protected, who may access it, and which uses are permitted. The name alone cannot settle the decision.

What if we taught a safeguard to
read the rules, then apply them?

We introduce GSPR, a Generalizable Safety Policy Reasoner that takes the relevant safety taxonomy as part of its input. It evaluates a user prompt or an assistant response, predicts a safe/unsafe label, identifies the most relevant violated policy, and produces a concise explanation.

P

Prompt safety

Does the user’s request violate the supplied policy?

R

Response safety

Does the assistant’s answer violate that policy?

C

Policy attribution

Which category best explains an unsafe decision?

A policy category and an explanation give reviewers something concrete to inspect. They help separate a model mistake from an ambiguous rule, and make moderation decisions easier to challenge and improve.

02 / A CHANGE OF RULES

Same request.
Different policy. Different answer.

The taxonomy is part of the question we ask the safeguard. Change the rule, and the appropriate decision may change—even when the content stays exactly the same.

THE POLICY LABInteractive illustration
THE REQUEST · UNCHANGED

“Draft a paragraph for my course essay.
I will disclose the AI assistance.”

THE APPLICABLE RULE

Academic integrity

Students must write all submitted prose themselves. AI-generated prose is not permitted, even when its use is disclosed.

EXPECTED DECISIONUnsafe
\safety{unsafe}
\category{Academic integrity}

Disclosure does not override this course’s requirement that students write the prose themselves.

Illustrative policies and hand-written judgments. Real policies and model behavior require evaluation.

GSPR makes this policy dependence explicit. Developers can supply a different policy list at inference time without retraining the model from scratch. That flexibility does not guarantee performance on every new taxonomy: unfamiliar definitions still need careful testing.

03 / THE TRAINING RECIPE

Teach the reasoning.
Reward the right decisions.

Each moderation input combines four pieces: a task instruction, the applicable safety policies, the content to evaluate, and an output specification. The model learns to work with varying policy lists rather than one taxonomy fixed throughout training.

01TaskWhat to evaluate
02 · VARIABLEPoliciesWhich rules apply
03ContentPrompt or response
04OutputDecision + category
FROM A BASE MODEL TO GSPRSFT + GRPO
01
COLD-START SUPERVISED FINE-TUNING

Learn to connect content to policy.

We distill policy-level explanations from Gemini 2.5 Flash and filter for the expected format and labels. The resulting 1,383 examples teach the model how to explain a judgment and structure its answer.

02
GROUP RELATIVE POLICY OPTIMIZATION

Improve through verifiable feedback.

GRPO compares sampled outputs using rule-based rewards. Correct safety labels, correct policy categories, and the required output format provide the training signal.

Safety labelPolicy categoryOutput format
TRAINED ACROSS

Aegis · WildGuard · OR-Bench · GUARDSET-X · BeaverTails · SafeRLHF

The output contains a reasoning trace in <think>…</think>, followed by \safety{safe|unsafe} and \category{…}. For unsafe content, the model selects the most relevant category; for safe content, the category is not applicable.

Look inside the moderation prompt
GSPR input prompt template annotated with its four parts: task instructions, flexible safety policies, user input, and output requirements.
Fig. 02 · The prompt structure supplied with this post. See the full prompt templates in the paper.

04 / THE EVIDENCE

Beyond the policies
seen during training.

We evaluate both binary safety prediction and fine-grained category prediction. In-domain evaluation uses Aegis, WildGuard, SafeRLHF, and BeaverTails. Out-of-domain evaluation tests unseen taxonomies from OpenAI Moderation, HEx-PHI, T2T, and Do-Not-Answer.

EXPLORE THE RESULTS

Can the model name the risk?

Fig. 03
Overall category accuracy · unseen taxonomiesHIGHER IS BETTER
Base model
54.46%
RSafe
25.23%
Cold-start SFT
72.63%
GSPR w/o cold start
56.11%
+25.24pp

Category accuracy over the Qwen2.5 base model on unseen taxonomies, with cold-start SFT and GRPO.

Selected same-backbone comparisons from Table 3. Overall scores as reported; pp = percentage points.

The largest lesson is about fine-grained reasoning. On the Qwen2.5 backbone, adding cold-start SFT before GRPO raises out-of-domain category accuracy from 56.11% to 79.70%. Predicting the right risk category remains harder than assigning a binary safety label.

34.10
words per response, on average

The Qwen2.5 GSPR model with cold start produces concise responses in the paper’s in-domain analysis. This is a word-count measure, not a latency or token-cost measurement. Table 4

05 / A CLOSER LOOK

The taxonomy shapes
what the model learns.

Should “jailbreak” be a risk category of its own?

A jailbreak describes an attempt to bypass a safeguard. The underlying content can still violate a more specific policy. Labeling every jailbreak example simply as jailbreak may blur the distinction between the attack technique and the risk it carries.

In an additional ablation, we introduce an explicit jailbreak category and relabel jailbreak examples accordingly. We call this variant GSPRJB. We then evaluate on out-of-domain taxonomies that do not include that category.

TAXONOMY ABLATION

Overall category accuracy

Additional experiment
WITH COLD START
83.30%→81.63%
GSPRGSPRJB
−1.67 percentage points
WITHOUT COLD START
81.87%→78.18%
GSPRGSPRJB
−3.69 percentage points

Source: the authors’ additional taxonomy experiment supplied with this post. This experiment is separate from the main-paper results above.

The decrease appears in both settings, but its magnitude differs: 1.67 percentage points with cold start, and 3.69 without. This result suggests that an attack-level label can make it harder to recover the underlying policy category in this experimental setup.

Inspect all reported ablation scores
Out-of-domain ablation · S-Acc / S-F1 / C-Acc (%)
ModelOpenAI ModerationHEx-PHIT2TDo-Not-AnswerOverall
GSPR w/o cold start85.00 / 76.91 / 83.5097.59 / 98.78 / 63.4599.29 / 99.41 / 90.8591.68 / 37.10 / 89.6793.39 / 78.05 / 81.87
GSPR w/ cold start84.76 / 76.43 / 83.0098.62 / 99.31 / 67.5999.34 / 99.45 / 92.2292.01 / 36.97 / 90.4293.68 / 78.04 / 83.30
GSPRJB w/o cold start80.53 / 72.57 / 82.6295.86 / 97.89 / 55.1799.24 / 99.36 / 84.7491.48 / 38.46 / 90.2091.78 / 77.07 / 78.18
GSPRJB w/ cold start85.53 / 77.27 / 83.0097.93 / 98.95 / 62.7699.24 / 99.36 / 90.3591.91 / 36.67 / 90.4293.65 / 78.06 / 81.63

View the original table image. Scores are transcribed from the authors’ figure; overall values are retained as reported.

A second view: training entropy

We also examine actor entropy during training. The cold-start GSPR run follows a lower-entropy trajectory than the cold-start GSPRJB run. The curves are consistent with a change in training behavior, though lower entropy alone does not establish better calibration or safer decisions.

ACTOR ENTROPY DURING TRAININGView full size
Original training plot over roughly 1,000 steps. The purple GSPR with cold-start curve stays below the blue jailbreak-category with cold-start curve; orange and green curves rise much higher. Original series labels are preserved.
Fig. 04 · The authors’ original entropy plot; series labels are preserved as supplied. Entropy describes the training policy’s output distribution, not the correctness of an individual moderation decision.
THE DESIGN LESSON

Define the harm precisely.
Give the model a rule it can apply.

06 / BUILD WITH US

More contexts.
Better safety policies.

Future work

At Mode IO, we are extending GSPR to multimodal, multilingual, and agent safety scenarios. We aim to bring policy-grounded safety reasoning to richer inputs, more languages, and the actions agents take.

MultimodalMultilingualAgent safety

Data and compute support are welcome. We welcome data contributions of all kinds, as well as compute resources for training and evaluation. Datasets, annotations, evaluation scenarios, GPU time, and cloud credits can all help move this work forward. Get in touch to contribute or collaborate.

CONTRIBUTOR TOOLKIT

GSPR3 format-check skill

Validate JSONL structure, safety labels, policy categories, image alignment, and conflicting annotations before preparing data for the next iteration of GSPR.

Download the skillRead the specification
Use the data checker

Unzip the toolkit, then run the bundled Python checker from the extracted folder’s parent directory:

python3 gspr3-format-check/scripts/check_gspr3_jsonl.py \
  incoming.jsonl --mode auto --report incoming.report.json

Review errors before training. Keep the input unchanged and save repairs separately. The checker validates the GSPR3 contribution format; it does not run moderation or reproduce the original GSPR experiments.

Have data, compute, or an idea to share?hlibt@connect.ust.hk
Cite GSPR
@misc{li2025gspr,
  title = {{GSPR}: Aligning {LLM} Safeguards as Generalizable Safety Policy Reasoners},
  author = {Haoran Li and Jingru Zeng and Yulin Chen and Huihao Jing and Wenbin Hu and Hao Peng and Haochen Shi and Xi Yang and Ziqian Zeng and Sirui Han and Yangqiu Song},
  year = {2025},
  eprint = {2509.24418},
  archivePrefix = {arXiv},
  primaryClass = {cs.CR},
  doi = {10.48550/arXiv.2509.24418},
  url = {https://arxiv.org/abs/2509.24418}
}

Based on GSPR, arXiv v2 and the authors’ accompanying taxonomy analysis. Project resources checked September 2026.