This is the inference-weight export of the 50-iteration baseline Formal-R1
GRPO checkpoint used before the subsequent self-OPD experiment and RL
continuation.
- Base model:
Qwen/Qwen3-4B-Instruct-2507
- Training method: GRPO
- Training iterations: 50 (
iter_0000049)
- Rollout seed: 42
- Training data:
examples/formal-r1/data/real_ok_valid.jsonl
- Evaluation data used by the launcher:
examples/formal-r1/data/real_valid.jsonl
cnl_rl base revision: 6541bdf
- Prompt source update from
cnl_itp: ee4cde0
The uploaded repository contains Hugging Face inference weights only. Optimizer
and distributed-training state remain in the original torch-distributed
checkpoint and are not included.
The model is specialized for the project's controlled-natural-language theorem
proving format. Outputs should still be checked by the corresponding verifier.