Model summary
Intended use
Automate radiology coding workflows, integrate into RIS/PACS systems, or build a report-to-code pipeline for teleradiology platforms.
This model is not a substitute for a certified medical professional's judgment. Output should be reviewed by a qualified person before being used in a clinical or billing decision. The model can make mistakes, especially on rare or compound cases.
How to use
from transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModelimport torch base_model = "unsloth/Qwen2.5-3B-Instruct"adapter = "AmareshHebbar/radiology-coder-qwen25-3b" tokenizer = AutoTokenizer.from_pretrained(base_model)model = AutoModelForCausalLM.from_pretrained( base_model, torch_dtype=torch.bfloat16, device_map="auto",)model = PeftModel.from_pretrained(model, adapter) messages = [ {"role": "system", "content": "You are a radiology coding specialist. Given a radiology report impression, return the ICD-10-CM codes for all findings."}, {"role": "user", "content": "IMPRESSION: 1.8cm hypoechoic nodule right thyroid lobe, TIRADS 4. Recommend FNA."},]inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)outputs = model.generate(inputs, max_new_tokens=128, temperature=0.1, do_sample=True)print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
Expected output:
ICD-10 Codes:\n1. E04.1 — Non-toxic single thyroid nodule (primary finding)\nRecommend: FNA if TIRADS ≥4. Add malignancy code post-biopsy.
With Unsloth (faster inference, recommended)
from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name="AmareshHebbar/radiology-coder-qwen25-3b", max_seq_length=512, load_in_4bit=True,)FastLanguageModel.for_inference(model) messages = [ {"role": "system", "content": "You are a radiology coding specialist. Given a radiology report impression, return the ICD-10-CM codes for all findings."}, {"role": "user", "content": "IMPRESSION: Acute pulmonary embolism involving bilateral main pulmonary arteries. No right heart strain."},]prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)inputs = tokenizer(prompt, return_tensors="pt").to("cuda")outputs = model.generate(**inputs, max_new_tokens=128, temperature=0.1, do_sample=True)print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
With vLLM (production serving)
vllm serve unsloth/Qwen2.5-3B-Instruct \ --enable-lora \ --lora-modules radiology-coder-qwen25-3b=AmareshHebbar/radiology-coder-qwen25-3b \ --host 0.0.0.0 --port 8000 --dtype bfloat16
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")response = client.chat.completions.create( model="radiology-coder-qwen25-3b", messages=[ {"role": "system", "content": "You are a radiology coding specialist. Given a radiology report impression, return the ICD-10-CM codes for all findings."}, {"role": "user", "content": "IMPRESSION: 3.2cm hypodense lesion right hepatic lobe, arterial enhancement with washout, consistent with hepatocellular carcinoma."}, ], temperature=0.1,)print(response.choices[0].message.content)
Training details
Data
Trained on 25,090 examples extracted from 25k radiology-relevant clinical notes filtered by imaging keywords. No synthetic or LLM-generated training data — every example pairs real-world input with its authoritative output.
- Train: 20,072 examples
- Validation: 2,509 examples
- Test: 2,509 examples
See the dataset card for the full extraction pipeline.
Hyperparameters
Table with columns: Parameter, Value| Parameter | Value |
|---|
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Quantization | 4-bit NF4 (QLoRA) |
| Max sequence length | 512 |
| Optimizer | paged_adamw_8bit |
| Learning rate | 2e-4, cosine schedule |
Training infrastructure
Fine-tuned with Unsloth for 2x faster training and reduced VRAM, using TRL's SFTTrainer. Training run on a single NVIDIA A40 GPU. Experiment tracking via Weights & Biases.
Limitations and bias
- Training data reflects a specific snapshot in time; outputs may become outdated as source authorities issue updates.
- The model may occasionally produce a plausible-sounding but incorrect output for rare or highly compound cases — always have a qualified person verify before downstream use.
- English-language input only.
Full collection: [link your HF collection here]
Citation
@misc{medicalai2026, author = {Hebbar, Amaresh}, title = {Medical AI Fine-tuning Suite}, year = {2026}, publisher = {HuggingFace}, url = {https://huggingface.co/AmareshHebbar}}