Benchmarks (greedy pass@1, same Colab/evalplus harness for all three)
Table with columns: Metric, base, SFT (MoM-python-slm), this model (GRPO), GRPO vs SFT| Metric | base | SFT (MoM-python-slm) | this model (GRPO) | GRPO vs SFT |
|---|
| MBPP | 66.7 | 69.6 | 72.5 | +2.9 |
| MBPP+ | — | — | 62.7 | — |
domain problem_solving (exec) | 0.700 | 0.713 | 0.767 | +5.4 |
domain spec_to_code (exec) | 0.632 | 0.714 | 0.729 | +1.5 |
domain api_usage (application) | — | 0.855 | 0.900 | +4.5 |
| HumanEval | 68.9 | 70.7 | 67.7 | −3.0 |
| HumanEval+ | — | — | 62.2 | — |
domain api_signature (param-recall) | 0.217 | 0.299 | 0.301 | +0.0 |
What GRPO did (load-bearing read)
GRPO is a specialization trade, not a free lunch. Gains land on exactly the execution-rewarded,
spec-driven dimensions — MBPP +2.9 and domain problem_solving +5.4 over SFT — while the
un-reinforced HumanEval completion format gives back −3.0 (slightly under base). That's the textbook
RLVR signature: the model sharpens "write a correct function from a spec" (what the MoM node actually
does) at a small cost to "graft a body under a fixed signature" (a format it never saw a reward for).
- Use this model for the spec-driven node role — it's the strongest on MBPP and the held-out domain eval.
- Use the SFT sibling if HumanEval-completion is a
hard gate — it remains the HumanEval-strongest checkpoint (70.7).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizertok = AutoTokenizer.from_pretrained("srivarenya/MoM-python-slm-grpo")model = AutoModelForCausalLM.from_pretrained( "srivarenya/MoM-python-slm-grpo", dtype="bfloat16", device_map="auto")
Prompt with the training system prompt + a Python task; the model returns reasoning then code. Reward,
training recipe, and the self-contained GRPO Colab notebook are in the project repository.
Next cross-check: LiveCodeBench (contamination-resistant), before/after vs the SFT sibling.