Training contract
- Base:
Qwen/Qwen3-VL-8B-Instruct at commit 0c351dd01ed87e9c1b53cbc748cba10e6187ff3b
- Dataset:
luca0621/appgen-sft-ngc-v1, config ngc-f-qwen3-normalized, revision 769ea99dbc4ff190048ae0db37eb6310dba595e0
- Dataset variant:
natural
- Dataset SHA-256:
8890e2fe596e4d78714ce4309f2a59ed9632dd5d656f758c8dc4608d46a9c6bc
- Exposures: 3,588; unique semantic examples: 3,168
- Direct-grounding exposures: 400
- Completion-retention replay exposures: 420
- Coordinates: normalized 0–1000
- Prompt variant:
proven; SHA-256 67ff8adb0e78a617f3d0edcf196d4e4cc3239967a8c30619fc5afc14484ee8c0
- Optimizer: full-language AdamW, learning rate
5e-7, cosine schedule, 5% warmup
- Batch: microbatch 4 × two GPUs × gradient accumulation 4 = global batch 32
- Epochs: 1; optimizer updates: 113
- Frozen: visual encoder and aligner; trainable: language model and LM head
- Swift's
hermes loss-scale preset was used (the sweep's tool2x arm).
Only the final checkpoint (checkpoint-113) is published in
this repository. The exact training system prompt is included as
appgen_system_prompt.txt; run_manifest.json records hashes and verification
evidence. Intermediate 25-step checkpoints remain local for controlled eval.
Data provenance and limitations
The images are synthetic AppGen HTML-to-PNG screenshots from 50 generated
training environments. They are not real-user or real-device captures. The
pinned dataset declares no AndroidWorld, held-out, or public grounding
benchmark images. No AndroidWorld or real-world benchmark score is claimed in
this card; evaluate all four arms under the same decoding and benchmark setup.
This model is for Android visual-agent research, not safety-critical autonomous
deployment.