Model summary
Intended use
Hospital reimbursement prediction, DRG validation tooling, and revenue cycle automation under the Medicare Inpatient Prospective Payment System (IPPS).
This model is not a substitute for a certified medical professional's judgment. Output should be reviewed by a qualified person before being used in a clinical or billing decision. The model can make mistakes, especially on rare or compound cases.
How to use
from transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModelimport torch base_model = "unsloth/Qwen2.5-1.5B-Instruct"adapter = "AmareshHebbar/icd10-to-drg-qwen25-1b" tokenizer = AutoTokenizer.from_pretrained(base_model)model = AutoModelForCausalLM.from_pretrained( base_model, torch_dtype=torch.bfloat16, device_map="auto",)model = PeftModel.from_pretrained(model, adapter) messages = [ {"role": "system", "content": "You are a DRG grouper. Given ICD-10-CM codes, return the MS-DRG code, relative weight, and geometric mean LOS."}, {"role": "user", "content": "ICD-10-CM code: I21.09"},]inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)outputs = model.generate(inputs, max_new_tokens=128, temperature=0.1, do_sample=True)print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
Expected output:
MS-DRG 280 — Acute Myocardial Infarction, Discharged Alive with MCC\nRelative Weight: 2.8613\nGeometric Mean LOS: 4.7 days
With Unsloth (faster inference, recommended)
from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name="AmareshHebbar/icd10-to-drg-qwen25-1b", max_seq_length=512, load_in_4bit=True,)FastLanguageModel.for_inference(model) messages = [ {"role": "system", "content": "You are a DRG grouper. Given ICD-10-CM codes, return the MS-DRG code, relative weight, and geometric mean LOS."}, {"role": "user", "content": "ICD-10-CM code: J18.9"},]prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)inputs = tokenizer(prompt, return_tensors="pt").to("cuda")outputs = model.generate(**inputs, max_new_tokens=128, temperature=0.1, do_sample=True)print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
With vLLM (production serving)
vllm serve unsloth/Qwen2.5-1.5B-Instruct \ --enable-lora \ --lora-modules icd10-to-drg-qwen25-1b=AmareshHebbar/icd10-to-drg-qwen25-1b \ --host 0.0.0.0 --port 8000 --dtype bfloat16
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")response = client.chat.completions.create( model="icd10-to-drg-qwen25-1b", messages=[ {"role": "system", "content": "You are a DRG grouper. Given ICD-10-CM codes, return the MS-DRG code, relative weight, and geometric mean LOS."}, {"role": "user", "content": "MS-DRG: 343"}, ], temperature=0.1,)print(response.choices[0].message.content)
Training details
Data
Trained on 5,385 examples extracted from real CMS MS-DRG v43.1 Definitions Manual + FY2026 Final Rule Table 5 weights. No synthetic or LLM-generated training data — every example pairs real-world input with its authoritative output.
- Train: 4,308 examples
- Validation: 538 examples
- Test: 539 examples
See the dataset card for the full extraction pipeline.
Hyperparameters
Table with columns: Parameter, Value| Parameter | Value |
|---|
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Quantization | 4-bit NF4 (QLoRA) |
| Max sequence length | 512 |
| Optimizer | paged_adamw_8bit |
| Learning rate | 2e-4, cosine schedule |
Training infrastructure
Fine-tuned with Unsloth for 2x faster training and reduced VRAM, using TRL's SFTTrainer. Training run on a single NVIDIA A40 GPU. Experiment tracking via Weights & Biases.
Limitations and bias
- Training data reflects a specific snapshot in time; outputs may become outdated as source authorities issue updates.
- The model may occasionally produce a plausible-sounding but incorrect output for rare or highly compound cases — always have a qualified person verify before downstream use.
- English-language input only.
Full collection: [link your HF collection here]
Citation
@misc{medicalai2026, author = {Hebbar, Amaresh}, title = {Medical AI Fine-tuning Suite}, year = {2026}, publisher = {HuggingFace}, url = {https://huggingface.co/AmareshHebbar}}