true | <two-digit risk category id> | <short reason>
false | <short reason>
Training Data
This checkpoint is trained with the PolicyShiftBench public data release:
- Dataset:
PolicyShiftBench/PolicyShiftBench
- Main evaluation splits: ID/adaptive branch and OOD/shift branch
- Training stages: randomized policy SFT followed by boundary-pair policy adaptation
Intended Use
Use this model for research on policy-conditioned multimodal safety, adaptive image moderation, and robustness under policy shifts. The model should be evaluated with explicit policy bundles rather than as a fixed universal safety classifier.
Limitations
This is a research checkpoint. It may fail under policies, languages, visual domains, or deployment settings not represented in the benchmark. Outputs should not be treated as legal or compliance advice.
Citation
If you use this model, please cite the paper:
@article{song2026policyshiftguard,
title = {PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails},
author = {Song, Mingyang and Xu, Luxin and Sun, Haoyu and Pan, Minzhou and Cheng, Yu and Li, Bo},
journal = {arXiv preprint arXiv:2607.05910},
year = {2026}
}