Deployment Profile
- Artifact size: 6.31 GiB
- Recommended starting envelope: 10.5 GiB free VRAM
- Estimate covers model weights plus practical single-image runtime headroom; larger images, batching, long generations, and server overhead need more.
Use this variant when CUDA memory matters more than maximum fidelity. The
deployment profile reports measured artifact size, a recommended free-VRAM
envelope for one-image inference, and full same-case quality deltas against
the merged model.
Run It
Download the strict runtime from the primary repository:
hf download MirilAI/Miril-DroneVLM-2B-2 inference.py router_contract.py requirements.txt --local-dir miril-drone-runtime
python -m pip install -r miril-drone-runtime/requirements.txt
python miril-drone-runtime/inference.py \
--model-id MirilAI/Miril-DroneVLM-2B-2-bnb4 \
--image drone_frame.jpg \
--prompt "Track the white car."
The exported checkpoint carries its quantization configuration. Do not add --load-4bit when loading this already-quantized repository.
Use transformers>=5.12.1. This artifact retains Gemma 4 E2B's shared-KV layout: language layers 15 through 34 reuse key/value states and intentionally have no separate k_proj, v_proj, or k_norm tensors. The artifact audit records the expected and observed tensor owners.
The helper rejects invalid JSON, schema violations, contradictory status/coordinate combinations, and non-target responses with coordinates. Only a valid target_found + precise_point may become a precise marker. A coarse_grid_direction remains a broad location cue and is never a landing, delivery, pointing, or tracking target.
Complete Benchmark

Table with columns: Metric, Merged BF16, CUDA bnb4| Metric | Merged BF16 | CUDA bnb4 |
|---|
| Valid JSON | 100.0% | 96.7% |
| Schema valid | 96.1% | 92.1% |
| Route accuracy | 94.8% | 91.2% |
| Caption / answer F1 | 38.2% | 37.2% |
| Spatial status | 79.2% | 69.7% |
| Precise target retained | 50.2% |
Complete held-out validation

Table with columns: Metric, Merged BF16, CUDA bnb4| Metric | Merged BF16 | CUDA bnb4 |
|---|
| Valid JSON | 99.9% | 99.0% |
| Schema valid | 99.9% | 98.7% |
| Route accuracy | 99.9% | 98.6% |
| Caption / answer F1 | 38.6% | 37.2% |
| Spatial status | 86.1% | 79.8% |
| Precise target retained | 71.7% |
Cleaned held-out deployment audit

After training, a stricter held-out audit removed pointing rows whose targets fall below the model-visible size threshold, then ran every release artifact on the complete revised validation and test splits. Strict-cleaned rows use only accepted evidence. Coverage-matched rows add evidence-preserving questions on the same held-out images to restore the earlier route and pointing action/status mix; they do not recreate the earlier object-class histogram.
Final held-out test
Strict-cleaned evidence
Table with columns: Metric, Merged BF16 - Strict cleaned test, CUDA bnb4 - Strict cleaned test| Metric | Merged BF16 - Strict cleaned test | CUDA bnb4 - Strict cleaned test |
|---|
| Valid JSON | 99.8% | 89.3% |
| Schema valid | 99.8% | 89.2% |
| Route accuracy | 99.8% | 89.1% |
| Reference text F1 | 60.9% | 53.5% |
| Spatial status | 88.0% | 72.0% |
| Coordinate quality | 89.7% |
Coverage-matched evidence
Table with columns: Metric, Merged BF16 - Coverage-matched test, CUDA bnb4 - Coverage-matched test| Metric | Merged BF16 - Coverage-matched test | CUDA bnb4 - Coverage-matched test |
|---|
| Valid JSON | 99.9% | 89.4% |
| Schema valid | 99.9% | 89.3% |
| Route accuracy | 99.9% | 89.1% |
| Reference text F1 | 64.3% | 56.6% |
| Spatial status | 84.2% | 70.8% |
| Coordinate quality | 84.9% |
Validation
Strict-cleaned evidence
Table with columns: Metric, Merged BF16 - Strict cleaned validation, CUDA bnb4 - Strict cleaned validation| Metric | Merged BF16 - Strict cleaned validation | CUDA bnb4 - Strict cleaned validation |
|---|
| Valid JSON | 99.9% | 96.4% |
| Schema valid | 99.9% | 96.1% |
| Route accuracy | 99.9% | 96.0% |
| Reference text F1 | 59.9% | 56.8% |
| Spatial status | 89.6% | 80.2% |
| Coordinate quality |
Coverage-matched evidence
Table with columns: Metric, Merged BF16 - Coverage-matched validation, CUDA bnb4 - Coverage-matched validation| Metric | Merged BF16 - Coverage-matched validation | CUDA bnb4 - Coverage-matched validation |
|---|
| Valid JSON | 99.9% | 96.3% |
| Schema valid | 99.9% | 96.0% |
| Route accuracy | 99.9% | 95.9% |
| Reference text F1 | 62.6% | 59.4% |
| Spatial status | 84.5% | 76.6% |
| Coordinate quality |
Validation supports comparison and model selection; test is the final held-out report. These automated scores measure contract and reference agreement, not flight safety.
Spoken-query results on the primary model card apply to the merged BF16 checkpoint. This deployment variant has not inherited that claim without a separate matched audio evaluation.
The merged and deployment variants use identical cases within each evaluation. Automated scores are regression signals, not safety certification.
The tables below compare merged and quantized artifacts on identical complete
held-out validation and test cases. Partial runs are excluded. Automated scores
are regression signals, not safety certification.
Limits And Safety
This is a research perception model, not a flight controller or certified safety system. Follow the complete limitations and operational guidance on the primary model card.
License
Apache License 2.0. See LICENSE and NOTICE.