Results
The SFT stage of the conditioned model. Numbers are reported for the GRPO stages, in the two companion repos.
Training
Table | |
|---|
| Stage | SFT (distillation), all levels in one corpus |
| Conditioning | the level is named in the prompt, not selected by adapter |
| Engine | HF transformers + peft |
| LoRA | r=16, alpha=32 |
| Hardware | 1x NVIDIA A100 80GB |
Usage
Solve this using Level {N} ({Verbose|Concise|Symbolic|Shorthand|Extreme}).
Problem: {your problem}
from transformers import AutoModelForCausalLM
from peft import PeftModel
model = AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-7B-Think", torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "ssurface/cot-dialect-olmo3-7b-think-conditioned-sft")
Limitations
- A design comparison, not a recommended model. The per-level adapters in this collection are the ones the results are reported on.
- Single seed.