Frozen-vision full-LLM SFT on the internally materialized GUIOwl action-balanced
64,000-row corpus. No extra rows were added outside the pinned corpus.
Important: this corpus is in-domain / AndroidWorld-contaminated. AndroidWorld
results from this model must not be reported as held-out generalization.
Base revision: 0c351dd01ed87e9c1b53cbc748cba10e6187ff3b
Final update: 1000
Global batch: 64 (8 GPUs × microbatch 8)
Learning rate: 5e-7, cosine schedule, 5% warmup
Frozen: vision encoder and aligner; trainable: full language model
Run manifest SHA-256: a0a4760ca2e07a237dd8a625225ef36992788d19058c93a6a0ab211b2522e6f5
The exact prompt, run manifest, materialization receipt, and training verification
are included as appgen_* files. At inference, use the included system prompt;
multi-step goals emit one concise <think>...</think> followed by one tool call,
while direct grounding requests emit only the tool call.