Why it exists — the NGARi model pipeline
NGARi's architecture pairs a large teacher model with small, deployable edge models:
Teacher (27B-class, e.g. qwen3.8-27B)
│ generates reasoning traces, synthetic data, eval judgments
▼
Edge models (1.5B–2B: ngari-ft-distilled, ngari-tool)
│ distilled / fine-tuned on teacher outputs
▼
Deployment: air-gapped edge hardware (Jetson AGX Orin, 8GB RAM)
This repo is the distilled student in that pipeline — capabilities that normally need a much larger model, compressed into a 1.5B footprint that runs entirely on owned hardware.
Provenance (verified Aug 3, 2026)
Table with columns: Attribute, Value| Attribute | Value |
|---|
| Base model | Qwen/Qwen2.5-1.5B-Instruct (Apache 2.0) — pinned in adapter_config.json |
| LoRA | rank 32, alpha 64, dropout 0.05, all linear projections |
| Synthetic data teacher | qwen3:8b (v1; 27B-class teacher planned for v2) |
| License | Apache 2.0 (NGARi-authored artifacts) |
| Hardware validated | aarch64 / NVIDIA Jetson AGX Orin, 8GB RAM, air-gap verified |
Google Gemma models were served only on NGARi hardware and were never used in NGARi training. All training used the Apache-2.0 Qwen2.5 lineage.
Evaluation
Chat quality — ngari-ft-distilled_chat_eval.json
{
"model": "ngari-ft-distilled",
"num_examples": 200,
"avg_score": 0.3766,
"avg_latency_sec": 2.81,
"tokens_per_sec": 39.78,
"total_time_sec": 561.98
}
{
"model": "ngari-ft-distilled:stable",
"num_examples": 20,
"tool_detection_rate": 0.6,
"tool_name_accuracy": 0.55,
"params_validity_rate": 0.6,
"avg_latency_sec": 2.38
}
For high-accuracy tool calling, use NGARiAI/ngari-tool (100% on all three tool metrics). This model's role is QA + safety judging.
Files
Table with columns: File, Purpose| File | Purpose |
|---|
model-*.safetensors (+ config) | Merged full model — use with Transformers |
adapter_model.safetensors | PEFT LoRA adapter — apply on the base |
ngari-ft-distilled-q4_K_M.gguf / -f16.gguf | GGUF — use with Ollama / llama.cpp |
Usage
# Ollama (GGUF)
ollama create ngari-ft-distilled -f Modelfile
ollama run ngari-ft-distilled "your prompt"
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("NGARiAI/ngari-ft-distilled")
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
adapter = PeftModel.from_pretrained(base, "NGARiAI/ngari-ft-distilled")
Companion repos
Sovereign AI
Trained and verified on user-owned edge hardware with zero cloud dependency. Verified air-gap (monitored via /proc/net/dev). "AI You Own. Completely."