About Hirundo
Hirundo develops machine-unlearning technology for
removing unwanted data and reducing unwanted behavior in trained AI models.
Hirundo's behavioral-unlearning workflow identifies a target behavior,
modifies the model, and evaluates the resulting checkpoint on both
behavior-specific and general-capability benchmarks.
Learn more:
What Hirundo changed
The target behavior for this checkpoint was prompt injection: cases where
adversarial or untrusted instructions cause the model to disregard its
intended task or reveal protected context.
The checkpoint was produced in a Hirundo behavioral-unlearning run with the
following configuration:
Table with columns: Field, Value| Field | Value |
|---|
| Base model | Qwen/Qwen3.5-4B |
| Target behavior | Security (prompt injection) |
| Unlearning aggressiveness | 1.0 |
Evaluation summary
Hirundo compared the hardened checkpoint with the unmodified base checkpoint
under the same evaluation setup. For the security evaluations below, lower
scores indicate fewer successful attacks or leaks and therefore better
performance. Utility results are reported as benchmark scores, where higher
is better.
Headline result: PurpleLlama prompt injection
On Meta's
PurpleLlama textual prompt-injection benchmark,
attack success rate decreased from 28.29% to 9.16%.
Table with columns: Benchmark, Base model ASR, Hirundo model ASR, Relative reduction| Benchmark | Base model ASR | Hirundo model ASR | Relative reduction |
|---|
| PurpleLlama prompt injection | 28.29% | 9.16% | 67.61% |
General-capability evaluation
The hardened model was evaluated on 6 general-capability benchmarks.
Table with columns: Benchmark, Change vs. base model| Benchmark | Change vs. base model |
|---|
| AIME25 | +4.59 pp |
| GPQA | +3.35 pp |
| IFBENCH | −0.49 pp |
| LiveCodeBench | +2.70 pp |
| MMLU-Pro | +1.61 pp |
| SciCode | 0.00 pp |
Across the six tasks, the unweighted mean change was +1.96 percentage points. The largest measured decrease was 0.49 percentage points on IFBENCH.
The means are descriptive rather than standardized aggregate scores: the
benchmarks measure different capabilities and may use different scoring
procedures.
Intended use
This checkpoint is intended for:
- research on model-level prompt-injection mitigation;
- comparative security and robustness evaluation;
- development of applications that require a model with lower measured
prompt-injection susceptibility than the upstream checkpoint; and
- further evaluation or fine-tuning by teams that understand the base model
and the security requirements of their deployment.
The model should be evaluated in the application's real prompt structure,
tooling environment, retrieval pipeline, and threat model before deployment.
from transformers import AutoModelForMultimodalLM, AutoProcessor
model_id = "hirundo-io/Qwen3.5-4B-hardened"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
model_id, dtype="auto", device_map="auto"
)
messages = [{"role": "user", "content": "Explain the difference between encryption and hashing."}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Consult the
Qwen/Qwen3.5-4B model card
for upstream architecture, supported modalities and languages, context length,
generation parameters, and base-model considerations.
License
This derivative retains the license of the upstream model. See the repository's
LICENSE file and the base-model repository
for details.
Citation
If you use this checkpoint, cite the upstream model and identify the checkpoint
as Hirundo's prompt-injection-hardened derivative:
@misc{hirundo_qwen3_5_4b_hardened_2026,
title = {Hirundo Qwen3.5 4B Hardened},
author = {{Hirundo}},
year = {2026},
howpublished = {\url{https://huggingface.co/hirundo-io/Qwen3.5-4B-hardened}},
note = {Prompt-injection-hardened derivative of Qwen/Qwen3.5-4B}
}