Training
Serve with vLLM
Ships the native Qwen3 40960 context (the yarn extension to 163840 was applied at post-training time); no overrides needed:
CUDA_VISIBLE_DEVICES=0 \
python -m vllm.entrypoints.openai.api_server \
--model TIGER-Lab/FIM-Mid-8B \
--served-model-name FIM-Mid-8B \
--host 127.0.0.1 \
--port 8400 \
--tensor-parallel-size 1 \
--max-model-len 40960 \
--gpu-memory-utilization 0.9 \
> vllm_fim_mid8b.log 2>&1 &
Post-training
To reproduce FIM-8B, run SWE-Lego trajectory SFT from this checkpoint — the exact config is posttraining/swe_lego/FIM_Posttrain_8B.yaml (LLaMA-Factory, full fine-tuning, lr 1.0e-4, 2 epochs — the official SWE-Lego recipe's 4 overfits this base — cutoff 131072 with yarn rope scaling, qwen3_nothink template, turn_mask enabled), which already points at this repo id. See posttraining/swe_lego/ for the walkthrough.
Citation
@article{wang2026fim,
title={Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models},
author={Wang, Yubo and Liang, Jiarong and Zhang, Yuxuan and Liu, Xuye and Wei, Cong and Zhang, Yuyu and Nie, Ping and Chen, Wenhu},
journal={arXiv preprint arXiv:2607.12463},
year={2026}
}