Model details
Table | |
|---|
| Base model | openai/gpt-oss-120b (MoE, harmony format) |
| Adapter | LoRA rank 16 (Apache-2.0, same as base) |
| Renderer / format | gpt_oss_no_sysprompt (harmony); direct final channel, no chain-of-thought |
| Task | ASR transcript formatting (single-turn: system + raw transcript → formatted transcript) |
| Languages | English (multi-dialect: en-US/GB/AU/IN/CA/IE/ZA) |
| Training | Tinker API — 6-stage curriculum SFT (L0→L5) then RLVR alignment |
Training pipeline
A layered curriculum (each stage continued-SFT from the previous, with replay
of all prior categories to prevent forgetting), then RL alignment:
Table with columns: Stage, Adds, Categories| Stage | Adds | Categories |
|---|
| L0 | basics | sent_end_punct, sent_start_cap, pronoun_I_cap, mixed_punct_cap, passthrough |
| L1 | surface | filler_removal, homophone_fix, itn_fix, comma_insertion, day_month_cap, mixed_L1 |
| L2 | disfluency | stutter_removal, false_start_collapse, backtrack_phrase, mixed_L2 |
| L3 | artifacts | proper_noun_fix, url_email_fix, mixed_L3 |
| L4 | layout | email_block, list_content, mixed_L4 |
21 categories total. Full dataset: Akash-Sakala/transcript-formatter-curriculum.
Results
Held-out evaluation over 2,384 rows across all 21 categories, scored with a
two-round protocol: deterministic Round 1 (EM / ContentOK / CER), then an
independent Claude judge re-adjudicates the Round-1 fails to credit cases where
the model is correct but the gold label is unrealistic ("adjusted accuracy").
Table with columns: Metric, Value| Metric | Value |
|---|
| Judge-adjusted accuracy | 99.9% |
| Round-1 PASS | 98.4% |
| Exact Match | 94.8% |
| ContentOK (no hallucination/omission) | 98.5% |
| CER / WER | 0.0011 / 0.0033 |
| FalseEdit (passthrough safety) | 0.9% |
Per-layer adjusted accuracy: L0 100% · L1 99.9% · L2 99.8% · L3 100% · L4 100%.
Only 2 genuine residual errors in 2,384 rows (one ITN edge case, one stacked
disfluency). No forgetting across layers; no hallucination.
EM understates quality for the layout categories (many valid paragraph/list
renderings); read ContentOK / adjusted accuracy / CER, not EM.
Usage
⚠ The system prompt is required (the model was trained with it verbatim) — see
merge_and_infer.py for the exact text and a runnable example.
import torchfrom transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModel base = AutoModelForCausalLM.from_pretrained( "openai/gpt-oss-120b", torch_dtype=torch.bfloat16, device_map="auto")model = PeftModel.from_pretrained(base, "Akash-Sakala/gpt-oss-120b-transcript-formatter-lora")tok = AutoTokenizer.from_pretrained("openai/gpt-oss-120b") SYSTEM_PROMPT = open("merge_and_infer.py").read() # copy the verbatim prompt from theremessages = [{"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": "i spoke to aisha she confirmed the PR will be merged by friday"}]ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)out = model.generate(ids, max_new_tokens=4096, do_sample=False)print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=False))# → "I spoke to Aisha. She confirmed the PR will be merged by Friday."
Merge once for serving (vLLM/SGLang):
# Option A — PEFTmerged = PeftModel.from_pretrained(base, "Akash-Sakala/gpt-oss-120b-transcript-formatter-lora").merge_and_unload()merged.save_pretrained("./gpt-oss-120b-transcript-formatter-merged") # Option B — Tinker cookbook (handles sharded / quantised merges for vLLM)from tinker_cookbook import weightsweights.build_hf_model( base_model="openai/gpt-oss-120b", adapter_path="./adapter", # this repo, downloaded locally output_path="./merged", dtype="bfloat16",)
Hardware: gpt-oss-120b is a 120B MoE — use a multi-GPU / large-VRAM host (or an
MXFP4-quantised serving stack). For cheaper deployment, use the distilled 20B
student (direct-final, no reasoning).
Intended use & limitations
- Use: post-processing ASR/Whisper output into readable text (dictation,
meeting notes, voice memos, emails). Greedy decoding (temperature 0) for
deterministic output.
- Preserves words. It edits format, not meaning — it does not paraphrase,
summarise, answer, or invent content; ambiguous cases pass through unchanged.
- Limitations: English only; rare/novel proper-noun spellings can be missed;
paragraph segmentation of free-dictated emails is subjective; not a general
assistant (it only formats transcripts).
- License: Apache-2.0, inherited from the base
openai/gpt-oss-120b.
Citation / provenance
Trained with the Tinker API. Teacher checkpoint:
85613875 (RL-final). LoRA rank 16. See the dataset card for the full curriculum.