Model summary
Table with columns: Property, Value| Property | Value |
|---|
| Upstream model | Qwen/Qwen3.8-27B |
| Architecture | Qwen3_5ForConditionalGeneration |
| Weight format | BF16 Transformers checkpoint |
| Quantization | None |
| Maximum configured context | 262,144 tokens |
This card intentionally excludes internal transformation, training, data,
deployment, and reproducibility details.
Evaluation
R1-HARMFUL-BENCH-450
Table with columns: Dataset, Prompts, Substring refusal flags, Flag rate, Manually confirmed safety refusals| Dataset | Prompts | Substring refusal flags | Flag rate | Manually confirmed safety refusals |
|---|
| AdvBench | 150 | 3 | 2.0% | 0 |
| StrongREJECT | 150 | 0 | 0.0% | 0 |
| XSTest | 150 | 0 | 0.0% | 0 |
|
Harmful subset (AdvBench + StrongREJECT): 300 prompts, 3 substring
flags (1.0%), 0 manually confirmed safety refusals, and 1 strict task
non-completion (0.33%). Manual review classified the other 2 flags as
capability disclaimers followed by substantive responses.
Quality checks: 0 API errors, 0 empty responses, and 0 manually
confirmed incoherent responses. The automatic detector produced 17
incoherence flags, all caused by repeated Markdown or diagram separator
characters rather than degraded output.
Intended use and limitations
This checkpoint is intended for private, controlled research and evaluation.
It may produce inaccurate, biased, unsafe, or otherwise undesirable output.
The checkpoint name is not a safety guarantee or a claim of suitability for
deployment. Evaluate it for the intended domain and apply appropriate access
controls and safeguards before any broader use.
License and attribution
The upstream model is released under the Apache License 2.0. Review the
upstream Qwen3.8-27B model card for
its documentation, license information, and original limitations.