الوصف (Arabic)
نموذج Qwen2.5-7B-Instruct معدل باستخدام QLoRA على بيانات اللجوء الأوروبية والألمانية.
الهدف: حفظ القوانين واستظهارها — وليس تعليم التفكير أو الاستدلال.
النموذج يميز بين القوانين السابقة (ما قبل 2026) والقوانين المحدّثة (حتى أبريل 2026).
Description (German)
Qwen2.5-7B-Instruct fine-tuned with QLoRA on European and German asylum law data.
Goal: Memorization and recall of laws — not reasoning.
The model distinguishes between previous (pre-2026) and updated (April 2026) laws.
Data
DAPT (Domain-Adaptive Pre-Training)
- Raw legal texts with packing
- Sources: AsylG, EU-Regulation 2024/1348, Dublin III, other EU/German laws
SFT (Supervised Fine-Tuning)
- Priority (updated 2026 laws): 5x weighted
- Historical (pre-2026 laws): capped at 58140
- Synthetic temporal comparisons: 48 examples
Temporal Awareness Strategy
- Dual system prompts (UPDATED / PREVIOUS)
- Time tags before each question
- Input field with amendment date (23. April 2026)
- Synthetic temporal comparison examples
- Priority x5 weighting + Historical capped
Training
- Base model: Qwen/Qwen2.5-7B-Instruct
- Method: QLoRA (rank 16, alpha 32, dropout 0)
- Hardware: 2x Tesla T4 (Kaggle)
- DAPT: 6 epochs, lr=2e-4, packing=True
- SFT: 5 epochs, lr=1e-4, early stopping
Evaluation
- Overall keyword score: Run evaluation cell first
- Questions: Run evaluation cell first
- Categories: temporal comparisons, basic recall, cross-temporal reasoning
Usage
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
"Qwen/Qwen2.5-7B-Instruct",
load_in_4bit=True,
device_map="sequential",
)
model.load_adapter("delta34/Qwen2.5-7B-AsylRef-LoRA/lora_adapter")
# OR with llama.cpp
./llama-cli -m Qwen2.5-7B.gguf --lora lora_qwen_asylref.gguf
Limitations
- المخرجات: lora.gguf فقط — يتطلب وجود النموذج الأساسي Qwen2.5-7B-Instruct
- النموذج يحفظ القوانين ولكن قد لا يقدم تفسيرات قانونية عميقة
- البيانات محدودة بالقوانين الأوروبية والألمانية فقط
- Temporal awareness تعتمد على جودة الأمثلة التركيبية
Known Issues
None at this time
- SFT formatting issue: قد يظهر الخطأ "'str' object has no attribute 'items'" بسبب إرجاع formatting_func قيمة str بدلاً من list. تم الإصلاح في الإصدار v2.0.0.
MLflow
- Experiment: qwen-asylref-training
- Parent run ID: not tracked (see individual DAPT/SFT runs in MLflow)
- Git hash: no-git