Training Objective
This adapter is trained to improve structured output accuracy
(JSON / YAML / XML / TOML / CSV).
Loss is applied only to the final assistant output,
while intermediate reasoning (Chain-of-Thought) is masked.
Training Configuration
- Base model: Qwen/Qwen3-4B-Instruct-2507
- Method: QLoRA (4-bit)
- Max sequence length: 1024
- Epochs: 1
- Learning rate: 3e-05
- LoRA: r=64, alpha=128
Usage
from transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModelimport torch base = "Qwen/Qwen3-4B-Instruct-2507"adapter = "your_id/your-repo" tokenizer = AutoTokenizer.from_pretrained(base)model = AutoModelForCausalLM.from_pretrained( base, torch_dtype=torch.float16, device_map="auto",)model = PeftModel.from_pretrained(model, adapter)
Sources & Terms (IMPORTANT)
Training data: daichira/structeval-t-sft-hq-yaml-cleaned
Dataset License: MIT License. This dataset is used and distributed under the terms of the MIT License.
Compliance: Users must comply with the MIT license (including copyright notice) and the base model's original terms of use.