Qwen3-VL-8B GUIOwl thinking64k SFT
Frozen-vision full-LLM SFT on the internally materialized GUIOwl action-balanced
64,000-row corpus. No extra rows were added outside the pinned corpus.
Important: this corpus is in-domain / AndroidWorld-contaminated. AndroidWorld
results from this model must not be reported as held-out generalization.
- Base revision:
0c351dd01ed87e9c1b53cbc748cba10e6187ff3b
- Final update:
1000
- Global batch:
64 (8 GPUs × microbatch 8)
- Learning rate:
5e-7, cosine schedule, 5% warmup
- Frozen: vision encoder and aligner; trainable: full language model
- Targets: 56,000 thinking + 8,000 gated action-only
- Unique semantics: 51,953; replay exposures: 12,047
- Prompt SHA-256:
9f4a9e84f15e73627725d25d44636223ccb859addc9d1af37cbef9a5c65092e8
- Materialization receipt SHA-256:
9f7aa0b5a7de5fc02ee3f9ba90f68bc15e97dc846bbd89edf70fc1e6f3b4cafa
- Run manifest SHA-256:
a0a4760ca2e07a237dd8a625225ef36992788d19058c93a6a0ab211b2522e6f5
The exact prompt, run manifest, materialization receipt, and training verification
are included as appgen_* files. At inference, use the included system prompt;
multi-step goals emit one concise <think>...</think> followed by one tool call,
while direct grounding requests emit only the tool call.