Models
Weight sizes are approximate; inference also requires memory for runtime allocations and the KV cache.
Task-held-out, localized scientific-code repair on familiar codebases, with original numerical verification. Qwen3.5-4B and Qwen3.5-9B are compared with PhAI-IDE-4B and PhAI-IDE-9B, respectively, on identical tasks. Pass rates are percentages; gains are percentage points.
Table with columns: Size, Environment, Tasks, Qwen3.5, PhAI-IDE, Gain (pp)| Size | Environment | Tasks | Qwen3.5 | PhAI-IDE | Gain (pp) |
|---|
| 4B | PLUTO-Particles-Dust | 3 | 0.00 | 33.33 | +33.33 |
| 9B | LAPS | 16 | 31.25 | 50.00 | +18.75 |
| 9B | MITgcm-biogeo |
Comparison with published models
Scores (%), grouped by benchmark and model size. Each reference entry gives its published score and the PhAI-IDE score difference in percentage points. Reference models are approximately the same size: 3–4B, 7–9B, and 67–72B, respectively.
Table with columns: PhAI-IDE, Benchmark, Score, Reference models: score (difference)| PhAI-IDE | Benchmark | Score | Reference models: score (difference) |
|---|
| 4B | BBH multistep-arithmetic-two | 97.60 | Llama-3.2-3B-Instruct (3.21B): 53.2 (+44.40); Phi-3.5-mini-8k-instruct (3.82B): 95.6 (+2.00) |
| 9B | BBH word-sorting | 60.40 | Llama-3.1-8B-Instruct (8.03B): 51.2 (); (7.62B): 15.6 () |
Reference scores come from the linked publications, model cards, and independent evaluation reports; evaluation settings and sample counts vary by source. Differences describe reported scores across evaluations, rather than matched-protocol head-to-head gains. BBH entries refer to the named tasks.
Quick start
Use Transformers 5.16.1, PyTorch and Accelerate. Set model_id to any model in the table above; the example selects the matching model class.
from transformers import AutoTokenizer, AutoModelForCausalLM, AutoModelForImageTextToText
model_id = "AItonomy/PhAI-IDE-4B"
loader = AutoModelForCausalLM if model_id.endswith("72B") else AutoModelForImageTextToText
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = loader.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": "Explain how to verify a numerical simulation."}],
add_generation_prompt=True, enable_thinking=False, return_dict=True, return_tensors="pt",
).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Training procedure
ScienceIDE demonstrations were collected with GPT-5.6-sol and filtered using a numerical-equivalence verifier. They capture code inspection, tool use, and responses to execution feedback.
All three models use ms-swift supervised fine-tuning with LoRA across trainable linear layers for three epochs. The release merges each final checkpoint's adapter into its base model. Retained assistant targets provide the next-token training signal, while conversation history and tool observations provide context. The trajectories retain the native exec / wait interaction format. Heuristic target masking selects assistant actions for supervision while preserving the surrounding interaction history.
Table with columns: Shared setting, Value| Shared setting | Value |
|---|
| Training dataset | Codex trajectories |
| Training examples / tasks | 4,567 segments / 564 tasks |
| Validation examples / tasks | 544 segments / 81 tasks |
| Train/validation task overlap | 0 |
| Training epochs | 3 |
| LoRA rank / alpha / dropout | 32 / 64 / 0.05 |
| Released weights | LoRA merged into BF16 Safetensors |
Long trajectories are organized into segments. Source partition assignments are preserved, with no task identifiers shared between training and validation.
Framework versions
The release was validated with the following environment.
Table with columns: Component, Version| Component | Version |
|---|
| Python | 3.11 |
| ms-swift | 4.5.3 |
| Transformers | 5.16.1 |
| PyTorch | 2.6.0+cu124 |
| PEFT | 0.20.0 |
| Datasets | 4.8.4 |
| Tokenizers | 0.23.2 |
| Accelerate | 1.14.0 |