Status
This checkpoint injects Ornstein thinking into Qwen3.8-27B. It is an early merge, not a finished quality release. Planned quality work uses RL environments and energy-based fine-tuning.
Evaluation
Qwen3.8-27B achieves an estimated 97.0% accuracy on the full GSM8K benchmark when running in standard unquantized precision (BF16/FP8).
Table with columns: Benchmark, Qwen3.8-27B (reported), Ornstein3.8-27B (this run)| Benchmark | Qwen3.8-27B (reported) | Ornstein3.8-27B (this run) |
|---|
| GSM8K | — | 96.51 (1273/1319) |
Single greedy BF16 run on a Fireworks dedicated H100 (temperature=0, top_k=40, max_tokens=4000), answers from message.content first. Qwen does not report GSM8K on the Qwen3.8-27B card. This score does not apply to GGUF quants.
Support this work
I'm a PhD student in visual neuroscience at the University of Toronto. Training and release compute is self-funded (rented H100s and a local DGX Spark). If these artifacts are useful, Ko-fi helps keep the experiments running.
Model details
Table | |
|---|
| Architecture | Qwen3_5ForConditionalGeneration |
| Parameters | ~27B dense |
| Context | 262,144 tokens |
| Hidden size / layers | 5120 / 64 |
| Attention | 24 heads, 4 KV heads, head_dim 256 |
| MLP intermediate | 17,408 |
| Vocab | 248,320 |
| Precision | bfloat16, 11 shards |
| Vision |
Usage
Requires a Transformers build with Qwen3.8 / qwen3_5 support.
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "GestaltLabs/Ornstein3.8-27B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, dtype="bfloat16", device_map="auto"
)
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": "https://example.com/image.jpg"},
{"type": "text", "text": "Describe this image."},
],
}
]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(out[0], skip_special_tokens=True))
Text-only chat uses the same template with {"type": "text", ...} and no image.
vLLM and SGLang: load this repo as a Qwen3.8 27B VLM (qwen3_5). Use a build that already supports that architecture.
Files
Table with columns: Path, Notes| Path | Notes |
|---|
model-00001-of-00011.safetensors … 00011 | BF16 weights |
model.safetensors.index.json | weight map, total_size 55562855904 |
config.json | Qwen3_5ForConditionalGeneration |
tokenizer.json / tokenizer_config.json / vocab.json / |
License
Apache 2.0, inherited from the Qwen 3.8 base release.