Held-out case evaluation
We evaluated 60 tasks from 20 held-out vulnerability cases. Each case has guided whitebox, focused whitebox, and focused blackbox views. Targets ran in isolated environments, and verifiers checked the final target state. The table reports adjudicated first-draw pass@1.
Table with columns: Model, Tasks solved, Pass@1, Estimated cost for 60 tasks| Model | Tasks solved | Pass@1 | Estimated cost for 60 tasks |
|---|
| Claude Opus 5 High | 43/60 | 71.7% | $74.68 (provider pricing) |
| apex-flash-1 | 40/60 | 66.7% | $2.38 |
| GLM-5.3-Flash | 36/60 | 60.0% | $4.56 (provider pricing) |
Capabilities
The checkpoint retains the base model's image-text-to-text architecture. Our reported evaluation covers text-based security tasks; we have not evaluated image or video performance.
Training and use
Training used production-like software and protocol environments with the Codex agent harness. The model is intended as a focused worker under a larger agent's direction. We recommend the Codex harness for this checkpoint.
What comes next
We are building harder, multi-step investigations as the model improves, broadening the data mix, and training across agent harnesses. We plan to publish more held-out and public benchmark results as they are validated. Explore Apex to see how this work supports production security.