I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my uploads have been useful to you, consider buying a PhD student a coffee. It goes a long way toward keeping these experiments running.
Vision/video restoration: The repository now includes the Qwen3.6 base visual tower plus image/video processor files, so the source safetensors checkpoint should load as the multimodal conditional-generation architecture. Existing GGUF artifacts made before this restoration remain text-only until rebuilt.
Apache 2.0 — inherited from the Qwen 3.6 base release.