Training
- Base:
openbmb/MiniCPM5-1B
- Base/tokenizer revision:
4e9de7a0778dc1c362e983e6858f0e77542cbdca
- Rows: 4,900; dataset SHA-256:
f3421604542d8f333576db814b471751900bc6aa2cfe109084c68d0f9ddf9c20
- Context: 2,048 tokens; no truncation; maximum rendered row 2,036 tokens
- Precision: BF16 LoRA, not QLoRA
- LoRA: rank 16, alpha 32, dropout 0.05; attention and MLP projections
- Epochs: 1; cosine learning-rate schedule; 3% warmup
- Peak learning rate: 2e-4
- Effective batch: 16 (1 x 16 gradient accumulation)
- Seed: 3407
- Loss: native assistant response only
- Trainer runtime: 1,020.55 seconds
- Adapter SHA-256:
8c0f24b5fce0063237b6f89ea22558b013b9e454bb3009895b0f81c1f8a65209
Frozen validation result
Evaluation used 1,050 private validation rows with normal five-tool retrieval,
no gold injection, 100% action-gold retrieval recall, thinking disabled, and no
decoding constraint. The validation dataset SHA-256 is
85094c96ca7fa2f96cbb0f7f85bd08510d56b9d4639646156d8806680bca9715.
Table with columns: Metric, Result| Metric | Result |
|---|
| End-to-end exact | 90.86% |
| Routing | 98.38% |
| Action exact | 69.55% |
| No-tool precision | 100% |
| No-tool recall | 99.86% |
| Listed-tool rate | 100% |
| Valid-call rate | 100% |
| Latency average | 0.272 s |
| Latency p50 | 0.141 s |
The untuned base on the identical CUDA validation condition scored 16.29%
end-to-end exact, 28.38% routing, 52.24% action exact, and 1.08% no-tool recall.
Limitations
This adapter is not yet promoted for autonomous execution. The frozen validation
set still contains 17 wrong-tool rows and 79 wrong-argument rows; action exact is
69.55%. Nested XML arguments are a recurring failure mode. Use strict listed-name
and schema validation or constrained decoding and reject invalid calls at runtime.
Constraints cannot repair semantically wrong listed tools or schema-valid wrong
arguments.
The private dataset and row-level evaluation repository is
turnercore/automaticity-v9.