Table of Contents
Model Overview
Table with columns: Attribute, Value| Attribute | Value |
|---|
| Base Model | Qwen/Qwen2.5-1.5B-Instruct |
| Task | Text Segmentation / Multilingual Semantic Chunking |
| Finetuning Method | Supervised Fine-Tuning (SFT) with LoRA |
| Framework Stack | Unsloth + Hugging Face TRL (SFTTrainer) |
| Supported Languages | Arabic, English, and Code-Switched (Mixed) Text |
| License | Apache 2.0 |
This model acts as a dedicated LoRA Adapter layered over Qwen2.5-1.5B-Instruct. When fed a text corpus, it splits it into logically self-contained, semantically coherent chunks. It preserves 100% of the input text—no paraphrasing, translation, summarization, or omission occurs.
Intended Use
This model is engineered to serve as a core component in enterprise-grade Arabic/multilingual NLP pipelines, specifically:
- Retrieval-Augmented Generation (RAG): Chunking unstructured documents into highly cohesive text segments prior to vector embedding generation.
- Advanced Preprocessing: A smart, context-aware alternative to naive rule-based text splitters (e.g., regex, token count limits) which frequently break up semantic contexts in complex Arabic phrasing.
- Technical Document Analysis: Flawlessly handling technical Arabic material heavily interspersed with English terminology.
- Information Extraction (IE): Isolating independent factual claims or narrative transitions for downstream analytical tasks.
Caution: This model is strictly a sequence segmenter. It is not designed for abstractive generation, translation, open-ended QA, or text summarization.
Training Details
LoRA Configuration
Table with columns: Parameter, Value| Parameter | Value |
|---|
| LoRA Rank (r) | 64 |
| LoRA Alpha | 64 |
| LoRA Dropout | 0 |
| Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Bias |
Supervised Fine-Tuning (SFT) Hyperparameters
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Training Epochs | 3 |
| Per-Device Train Batch Size | 1 |
| Gradient Accumulation Steps | 4 |
| Effective Batch Size | 4 |
| Learning Rate | 1e-4 |
| LR Scheduler Type | Cosine |
| Warmup Ratio | 0.1 |
| Optimizer | adamw_8bit |
The final adapter utilized in production corresponds to checkpoint-1000, chosen for its optimal evaluation convergence and stable validation loss profile without indicating signs of overfitting.
Dataset Architecture
The dataset was constructed via a rigorous Knowledge Distillation pipeline utilizing gemini-3.1-flash-lite as the Teacher Model to generate high-quality ground-truth mappings (Raw Text → Syntactically Valid JSON Chunks).
Data Sources & Distribution
Three distinct text styles were synthesized to ensure cross-domain generalization:
Table with columns: Source, Description, Count| Source | Description | Count |
|---|
| Wikipedia (Modern Standard Arabic) | Formal, encyclopedic prose | 568 |
| XLSum (BBC Arabic News) | Direct, objective journalistic style | 1476 |
| Technical Wikipedia (Mixed) | Code-switched documents (5%–40% English token density) | 239 |
| Total Corpus Size | | 2,257 Articles |
Data Splits
Table with columns: Split, Sample Count| Split | Sample Count |
|---|
| Train | 2000 |
| Validation | 257 |
Each training sequence wraps the original raw layout into an instruction text prompt, with a strictly deterministic, character-exact JSON schema mapping as the targeted output.
Infrastructure and Hardware
Table with columns: Component, Details| Component | Details |
|---|
| Graphics Processing Unit (GPU) | NVIDIA Tesla T4 |
| Environment Environment | Google Colab (Free Tier) |
| Training Optimization Stack | Unsloth AI + Hugging Face TRL |
| Experiment Tracking | Weights & Biases (WandB) |
Integrating Unsloth was vital to successfully execute this training pipeline on a commodity T4 GPU's restricted VRAM profile (15GB), providing nearly a 2x acceleration advantage alongside minimal hardware activation overhead.
Following convergence, the model was deployed via the high-concurrency vLLM engine and stress-tested using Locust to simulate realistic concurrent production workloads.
Serving Configuration (vLLM)
Table with columns: Metric / Parameter, Configured Value| Metric / Parameter | Configured Value |
|---|
| Inference Engine | vLLM Engine Process |
| GPU Memory Utilization | 0.8 (80% of VRAM allocated to KV Cache) |
Max Model Length (max_model_len) | 2048 Tokens |
| Max LoRA Rank | 64 |
| Inference Precision | half (Float16 Execution) |
A sustained load test simulating simultaneous clients generating continuous validation payloads yields the following verified metrics:
Table with columns: Performance Metric, Evaluation Result| Performance Metric | Evaluation Result |
|---|
| Inference Throughput | 437.20 Tokens / Second |
| GPU KV Cache Usage: | ~3.2%(Significant headroom remaining) |
| Request Failure Rate | 0.00% (Zero Errors Over Duration) |
| Sustained Duration | 60 Seconds |
| Concurrency Load Tool | Locust (Simulating 20 Concurrent Users) |
Architectural Note: These benchmarks were successfully conducted under a concurrent load of 20 users. vLLM's Continuous Batching and PagedAttention maximized serving stability, keeping the GPU KV Cache usage at a highly efficient ~3.2% with zero out-of-memory (OOM) or timeout constraints.
Usage Guide
Loading via Unsloth (Recommended for Inference)
from unsloth import FastLanguageModel # Load the model and tokenizer natively using Unslothmodel, tokenizer = FastLanguageModel.from_pretrained( model_name = "Mo-Abdelfattah/arabic-semantic-chunker-qwen1.5b", max_seq_length = 3500, dtype = None, load_in_4bit = False,)FastLanguageModel.for_inference(model)
from transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import PeftModel # Load base model firstbase_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") # Wrap with the LoRA adapter weightsmodel = PeftModel.from_pretrained(base_model, "Mo-Abdelfattah/arabic-semantic-chunker-qwen1.5b")tokenizer = AutoTokenizer.from_pretrained("Mo-Abdelfattah/arabic-semantic-chunker-qwen1.5b")
Production Serving via vLLM CLI
vllm serve "Qwen/Qwen2.5-1.5B-Instruct" \ --dtype=half \ --gpu-memory-utilization 0.8 \ --max-model-len 2048 \ --max-lora-rank 64 \ --enable-lora \ --lora-modules arabic-chunker="Mo-Abdelfattah/arabic-semantic-chunker-qwen1.5b"
The model strictly adheres to outputting valid, parsable JSON matching the following structural schema:
{ "original_text_length": 2345, "semantic_chunks": [ "First coherent semantic segment extracted from the text...", "Second logical chunk containing continuing semantic context..." ]}
Limitations & Boundary Conditions
- Strict Task Bound: The model is exclusively trained for text sequence segmentation. Abstractive reasoning, creative generation, or open-domain synthesis tasks are explicitly unsupported.
- Context Length Restrictions: The training pipeline sliced input sources at a maximum boundary of 1,500 characters per document sample; performance on exceedingly long inputs (e.g., >10k tokens) may degrade.
- Distillation Dependency: The training data utilizes synthesized extractions from
gemini-3.1-flash-lite, meaning that any localized logical parsing bias present in the teacher model may occasionally be inherited.
Author
Mohamed Abdelfattah
AI & Data Platform Engineer specializing in high-throughput Arabic NLP architectures, robust LLM alignment pipelines, and optimal distributed serving infrastructures.