Batch: microbatch 4 × two GPUs × gradient accumulation 4 = global batch 32
Epochs: 1; optimizer updates: 113
Frozen: visual encoder and aligner; trainable: language model and LM head
Swift's default loss scale was used.
Only the final checkpoint (checkpoint-113) is published in
this repository. The exact training system prompt is included as
appgen_system_prompt.txt; run_manifest.json records hashes and verification
evidence. Intermediate 25-step checkpoints remain local for controlled eval.
Data provenance and limitations
The images are synthetic AppGen HTML-to-PNG screenshots from 50 generated
training environments. They are not real-user or real-device captures. The
pinned dataset declares no AndroidWorld, held-out, or public grounding
benchmark images. No AndroidWorld or real-world benchmark score is claimed in
this card; evaluate all four arms under the same decoding and benchmark setup.
This model is for Android visual-agent research, not safety-critical autonomous
deployment.