🎓 Academic Context
This study was conducted to address the challenge of global VLM models failing to recognize local cultural elements, specifically Turkish cuisine.
- University: Fırat University, Faculty of Technology
- Department: Software Engineering
- Course: Senior Design Project (Bitirme Projesi)
- Supervisor: Assoc. Prof. Dr. Özal Yıldırım
- Student: Turhan Göksu
Model Details
Model Description
General-purpose VLMs often misclassify specific cultural dishes (e.g., confusing Mantı with Ravioli or Pasta). This model was developed to identify Turkish meals and answer nutrition questions about them.
It utilizes QLoRA (LoRA adapters trained on a 4-bit quantized base model) for efficient fine-tuning, preserving the capabilities of the base Qwen2-VL model while injecting domain-specific knowledge about Turkish cuisine.
Model Sources
Supported Turkish Foods (47 Categories)
The model has been trained to recognize and provide nutritional information for the following Turkish dishes:
Main Courses & Kebabs
Adana Kebap, Patlıcan Kebabı, İskender, Taş Kebabı, Et Döner, Hünkar Beğendi, Mantı, Hamsi Tava
Stuffed Dishes (Dolma & Sarma)
Beyaz Lahana Sarması, Biber Dolması, Midye Dolma, Mumbar Dolması, Yaprak Sarma
Pastries & Börek
Gözleme, Su Böreği, Lahmacun, Kıymalı Pide, Simit
Appetizers & Street Food
Çiğ Köfte, İçli Köfte, Çoban Salatası, Kısır, Midye Tava, Mücver, Kabak Mücver, Tantuni, Sucuklu Yumurta, Cacık, Menemen
Soups (Çorba)
Mercimek Çorbası, Yayla Çorbası, Domates Çorbası, Şehriye Çorbası, Tarhana Çorbası
Vegetable Dishes
Kuru Fasulye, Taze Fasulye
Köfte Varieties
Mercimek Köftesi
Desserts (Tatlı)
Baklava, Sütlaç, Kazandibi, Tulumba Tatlısı, Lokma, Lokum, Kemal Paşa Tatlısı, Kalburabastı
Beverages
Türk Kahvesi, Sahlep
Note: The model may attempt to recognize other Turkish foods, but it was only trained on the 47 dishes listed above, and answers for other dishes are unreliable.
Uses
Direct Use
The model is intended for:
- Food Tracking Applications: Automating food logging for Turkish users.
- Diet and Nutrition Assistance: Providing instant feedback on meals (e.g., "Is this suitable for a diet?").
- Cultural Gastronomy Education: Helping foreigners identify Turkish foods.
Out-of-Scope Use
- Medical Diagnosis: This model is not a doctor. Suggestions should not be taken as medical prescriptions.
- Non-Food Images: The model is optimized for food images; it may hallucinate or provide irrelevant answers on unrelated images.
- Unsupported Foods: For Turkish dishes not in the training set, the model may provide incorrect classifications or nutritional estimates.
- Portion-Specific Nutrition: The model returns typical per-portion values for the recognized dish; it does not measure the portion in the photo.
How to Get Started with the Model
You can use this model with transformers and peft.
Installation
pip install -U transformers peft accelerate pillow
Inference Code
import torchfrom PIL import Imagefrom transformers import Qwen2VLForConditionalGeneration, Qwen2VLProcessorfrom peft import PeftModel # 1. Load Base Model and Processorbase_model_id = "Qwen/Qwen2-VL-2B-Instruct"adapter_id = "Turhan123/turkish-cuisine-vlm" # Load Base Model (fp16 works on T4 and newer GPUs)model = Qwen2VLForConditionalGeneration.from_pretrained( base_model_id, torch_dtype=torch.float16, device_map="auto",)processor = Qwen2VLProcessor.from_pretrained(base_model_id) # 2. Load the Fine-Tuned Adaptermodel = PeftModel.from_pretrained(model, adapter_id) # System prompt used during trainingSYSTEM_PROMPT = "Sen Türk mutfağı konusunda uzman bir diyetisyen ve VLM asistanısın. Yemekleri tanı ve besin değerlerini söyle." # 3. Define Helper Functiondef ask_dietitian(image_path, question="Bu yemek nedir ve besin değerleri nasıldır?"): image = Image.open(image_path) conversation = [ {"role": "system", "content": [{"type": "text", "text": SYSTEM_PROMPT}]}, { "role": "user", "content": [ {"type": "image", "image": image}, {"type": "text", "text": question}, ], }, ] text_prompt = processor.apply_chat_template(conversation, add_generation_prompt=True) inputs = processor( text=[text_prompt], images=[image], padding=True, return_tensors="pt" ).to(model.device) output_ids = model.generate(**inputs, max_new_tokens=150, do_sample=False) generated_ids = [ output_ids[len(input_ids):] for input_ids, output_ids in zip(inputs.input_ids, output_ids) ] return processor.batch_decode(generated_ids, skip_special_tokens=True)[0] # Example Usage# response = ask_dietitian("kebap.jpg")# print(response)
Training Details
Training Data
The model was trained on a Custom Turkish Food Dataset curated specifically for this project.
- Categories: 47 types of foods including Kebabs, Desserts, Soups, and Street Foods.
- Size: 17,631 image–question–answer rows over 957 images (about 20 images per category, 13–20 Q&A pairs per image).
- Images: Collected from the Kaggle datasets listed below; most are 256×256 pixels.
- Structure: Image + Question (User) + Answer (Assistant/Dietitian), all in Turkish.
- Answers: Hand-written per category (identification, calories, macros, ingredients, diet suitability). All images of a category share the same Q&A set.
Example (baklava):
Q: Kilo aldırır mı?
A: Evet, şerbet (şeker) ve yağ içeriği çok yüksek olduğu için porsiyon kontrolü yapılmazsa hızla kilo aldırır.
Training Procedure
- Training Regime: Supervised Fine-Tuning (SFT) with QLoRA, using
transformers, peft, bitsandbytes and trl (SFTTrainer).
- Quantization: 4-bit NF4 base model with double quantization, bfloat16 compute dtype.
- Precision: BF16 mixed precision training on NVIDIA A100 GPU (Google Colab).
- Data Split: Random 95% / 5% train/eval split.
Optimization Hyperparameters:
- Epochs: 10
- Learning Rate: 2e-4
- Batch Size: 8 per device, gradient accumulation 2 (effective 16)
- LoRA Rank (r): 64
- LoRA Alpha: 128
- Dropout: 0.05
- Target Modules:
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Evaluation
Testing Results
No quantitative benchmark was run. The model was checked informally with a "blind test" on random images collected from the internet (not seen during training).
Observations:
- Dish Recognition: The model identified dishes from the 47 supported categories and answered in the style of the training data.
- Dietary Explanations: Besides numbers, the model reproduces the explanations it was trained on (e.g., why a dessert is high in calories).
Environmental Impact
- Hardware Type: NVIDIA A100 GPU (Google Colab)
- Compute Region: Cloud-based
- Carbon Emitted: Low (Fine-tuning was efficient via QLoRA, taking <3 hours).
Acknowledgements & Special Thanks
I would like to express my special thanks to the creators of the following resources, which made this project possible:
Special Thanks for Tutorial & Inspiration:
Dataset Sources:
Technical References:
- Qwen2-VL Paper: Wang et al., "Qwen2-VL: To See the World More Clearly", arXiv:2409.12191
- LoRA Technique: Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", arXiv:2106.09685
- QLoRA Technique: Dettmers et al., "QLoRA: Efficient Finetuning of Quantized LLMs", arXiv:2305.14314
Bias, Risks, and Limitations
- Calorie Estimates: The nutritional values are hand-written general estimates based on standard recipes, not taken from a verified nutrition database. Actual values may vary significantly based on portion size and specific cooking methods.
- Category-Level Answers: Because every image of a dish shares the same answers, the model effectively recognizes the dish and returns its standard information; it cannot reflect portion size or variations visible in the photo.
- Limited Food Coverage: The model is trained on 47 specific Turkish dishes. Performance on other Turkish foods or international cuisines may be suboptimal.
- Small, Low-Resolution Training Set: About 20 images per dish, mostly 256×256 pixels, which limits robustness to unusual photos.
- No Rigorous Evaluation: The train/eval split was made per Q&A row, so the same images appear in both sets and the eval loss does not measure generalization. No held-out accuracy is reported.
- Hallucinations: Like all Large Language Models, the model might generate incorrect information for ambiguous or very low-quality images.
- Not Medical Advice: This model is not a substitute for professional medical or dietary consultation.