Model summary
Intended use
Support ABDM/ABHA health record integration for AYUSH practitioners and integrative medicine hospitals that need to document in standard codes.
This model is not a substitute for a certified medical professional's judgment. Output should be reviewed by a qualified person before being used in a clinical or billing decision. The model can make mistakes, especially on rare or compound cases.
How to use
from transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModelimport torch base_model = "unsloth/Qwen2.5-1.5B-Instruct"adapter = "AmareshHebbar/ayurveda-icd-qwen25-1b" tokenizer = AutoTokenizer.from_pretrained(base_model)model = AutoModelForCausalLM.from_pretrained( base_model, torch_dtype=torch.bfloat16, device_map="auto",)model = PeftModel.from_pretrained(model, adapter) messages = [ {"role": "system", "content": "You are an integrative medicine specialist. Map the traditional medicine condition to its ICD-10-CM equivalent with a clinical note."}, {"role": "user", "content": "Traditional medicine condition: Amavata"},]inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)outputs = model.generate(inputs, max_new_tokens=128, temperature=0.1, do_sample=True)print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
Expected output:
Nearest ICD-10: M06.9 — Rheumatoid arthritis, unspecified\nClinical note: Amavata is characterised by joint inflammation and constitutional symptoms closely mirroring RA. Check RF, anti-CCP antibodies.
With Unsloth (faster inference, recommended)
from unsloth import FastLanguageModel model, tokenizer = FastLanguageModel.from_pretrained( model_name="AmareshHebbar/ayurveda-icd-qwen25-1b", max_seq_length=512, load_in_4bit=True,)FastLanguageModel.for_inference(model) messages = [ {"role": "system", "content": "You are an integrative medicine specialist. Map the traditional medicine condition to its ICD-10-CM equivalent with a clinical note."}, {"role": "user", "content": "Traditional medicine condition: Madhumeha"},]prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)inputs = tokenizer(prompt, return_tensors="pt").to("cuda")outputs = model.generate(**inputs, max_new_tokens=128, temperature=0.1, do_sample=True)print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
With vLLM (production serving)
vllm serve unsloth/Qwen2.5-1.5B-Instruct \ --enable-lora \ --lora-modules ayurveda-icd-qwen25-1b=AmareshHebbar/ayurveda-icd-qwen25-1b \ --host 0.0.0.0 --port 8000 --dtype bfloat16
from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")response = client.chat.completions.create( model="ayurveda-icd-qwen25-1b", messages=[ {"role": "system", "content": "You are an integrative medicine specialist. Map the traditional medicine condition to its ICD-10-CM equivalent with a clinical note."}, {"role": "user", "content": "Traditional medicine condition: Pandu"}, ], temperature=0.1,)print(response.choices[0].message.content)
Training details
Data
Trained on 3,002 examples extracted from real CMS ICD-10-CM Z and Y codes plus traditional medicine terminology crosswalk. No synthetic or LLM-generated training data — every example pairs real-world input with its authoritative output.
- Train: 2,401 examples
- Validation: 300 examples
- Test: 301 examples
See the dataset card for the full extraction pipeline.
Hyperparameters
Table with columns: Parameter, Value| Parameter | Value |
|---|
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0 |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Quantization | 4-bit NF4 (QLoRA) |
| Max sequence length | 512 |
| Optimizer | paged_adamw_8bit |
| Learning rate | 2e-4, cosine schedule |
Training infrastructure
Fine-tuned with Unsloth for 2x faster training and reduced VRAM, using TRL's SFTTrainer. Training run on a single NVIDIA A40 GPU. Experiment tracking via Weights & Biases.
Limitations and bias
- Training data reflects a specific snapshot in time; outputs may become outdated as source authorities issue updates.
- The model may occasionally produce a plausible-sounding but incorrect output for rare or highly compound cases — always have a qualified person verify before downstream use.
- English-language input only.
Full collection: [link your HF collection here]
Citation
@misc{medicalai2026, author = {Hebbar, Amaresh}, title = {Medical AI Fine-tuning Suite}, year = {2026}, publisher = {HuggingFace}, url = {https://huggingface.co/AmareshHebbar}}