Before the assistant model
If a user request is blocked, the assistant model is not called.
Maintain the rules. Route each request and proposed action. Review safety and category at the leaf before the harness proceeds.
1,000 held-out routes · no confidence filtering
73,415 / 79,987 labeled safety records
147,643 records · excludes full routing + review
Content, cyber, privacy and compliance share one tree. Each leaf holds the applicable taxonomy and review rules. Change the tree to fit your agent.
Jev chooses a branch at each level using the request and its audit context. At the selected leaf, it predicts safety and the applicable categories. The harness combines these outputs with authorization rules.
What is gravity?
The request reaches the assistant model.
Illustrated workflow; this page does not call the reviewer.
If a user request is blocked, the assistant model is not called.
If a proposed action is blocked, the tool does not execute. The interception appears in the conversation.
English narration · English and Chinese captions. The opening animates an actual Jev action-review trace. The following Hermes + DeepSeek session explains each input and highlights Jev interceptions in red.
Demo and recording details ↗Qwen3.5-2B with rank-8 LoRA and classification heads. Routing and leaf review use structured probabilities rather than generated explanations.
Child branches for routing; leaf taxonomy for review.
Last non-padding hidden state
Two-class linear head
safe / unsafeCandidate embedding similarity
softmax / sigmoidScalar + ordered thresholds
Ordinal severity147,643 native task records across 109 benchmark views. Safety accuracy and category micro-F1 use records with the corresponding supervision.
| Domain | Task cases | Safety accuracy | Category micro-F1 |
|---|
| Benchmark | Cases | Safety accuracy | Category F1 | Median forward |
|---|
Selected v3.2 continuation checkpoint at update 1,508 after one epoch; a validation plateau was not established. Native tests retain original v3.1 candidates and gold labels. Full-path routing starts at root without gold intermediate nodes, with caller audit context supplied.
Controlled action-state replay caught 49/52 malicious proposals and allowed 87/97 benign proposals, including one API error among interruptions. No attacker action was executed; this is not end-to-end attack success or a production error estimate.
Quoted injection analysis can still be incorrectly blocked. Some expanded categories have no positive training examples. Ordinal severity exact accuracy is 15.28% on 1,466 cases. Editing the tree does not retrain the model. Keep native permissions and explicit scope checks.
Machine-readable evidence ↗Copy and run on macOS or Linux. This installer connects to our hosted Jev v3.2 reviewer over HTTPS; no local GPU is needed. It installs or updates the supplement, preserves your policy, and checks the connection.
curl -fsSL https://huggingface.co/hubin/jev-guard-v3.2-2b/resolve/main/downloads/install-hermes-jev.txt -o /tmp/install-hermes-jev.sh
bash /tmp/install-hermes-jev.sh --endpoint https://approaches-lemon-antique-garden.trycloudflare.com
# Then start the guarded conversation:
jev-hermesThe assistant still needs its provider credentials. Configure your assistant provider with hermes setup; existing credentials are preserved.
The same installer can prompt for your reviewer URL. No SSH host is assumed in this mode. The model is downloadable on Hugging Face; deploy its reviewer server before connecting Hermes.
bash /tmp/install-hermes-jev.sh