Ornith 1.5 397B
This model card documents Ornith-1.5-397B, the flagship member of the Ornith-1.5 family — a 397B mixture-of-experts model. It scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, performing on par with Claude Opus 4.8 (85.0 and 59.0) while outperforming leading open-source models of similar scale, including GLM-5.2 and DeepSeek-V4-Flash-0731.
Benchmarks
Quickstart
Serving Ornith-1.5-397B
Ornith-1.5-397B is a ~397B mixture-of-experts model (≈800 GB in bf16), so multi-GPU serving is required. The recipes below use 8-way tensor parallelism on a single node (e.g., 8× H200 141GB); adjust --tensor-parallel-size / --tp to match your hardware, or use FP8/INT4 quantized builds for smaller deployments.
vLLM
vllm serve ornith-ai/Ornith-1.5-397B \
--served-model-name Ornith-1.5-397B \
--host 0.0.0.0 --port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--trust-remote-code
SGLang
python -m sglang.launch_server \
--model-path ornith-ai/Ornith-1.5-397B \
--served-model-name Ornith-1.5-397B \
--host 0.0.0.0 --port 8000 \
--tp 8 \
--context-length 262144 \
--mem-fraction-static 0.85 \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3
For Long-Context
Ornith-1.5-397B handles context windows of up to 262,144 tokens. When a task's combined input and output must go beyond this limit, we suggest extending the effective window with RoPE scaling — YaRN is the technique we validate against, and it is already built into both vLLM and SGLang. With a scaling factor of 4.0, the usable window grows to roughly 1M tokens.
You can turn YaRN on in either of two ways:
-
Edit the checkpoint's config.json. Add a rope_scaling block to the model configuration:
{
"rope_scaling": {
"rope_type": "yarn",
"factor": 4.0,
"original_max_position_embeddings": 262144
}
}
-
Override at launch time. Leave the checkpoint untouched and extend the serve commands above with the equivalent flags.
vLLM:
VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ornith-ai/Ornith-1.5-397B ... --hf-overrides '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' --max-model-len 1000000
SGLang:
SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... --json-model-override-args '{"rope_scaling": {"rope_type": "yarn", "factor": 4.0, "original_max_position_embeddings": 262144}}' --context-length 1000000
Using Ornith-1.5-397B via the Chat Completions API
Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client.
Basic Usage
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="Ornith-1.5-397B",
messages=[
{"role": "user", "content": "Write a one-line Python lambda that squares a number."}
],
temperature=0.6,
top_p=0.95,
max_tokens=1024,
)
message = response.choices[0].message
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)
You can also stream tokens, or hand the model tools — Ornith-1.5-397B emits well-formed function calls that the server parses into the standard tool_calls field:
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
response = client.chat.completions.create(
model="Ornith-1.5-397B",
messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
tools=tools,
tool_choice="auto",
temperature=0.6,
max_tokens=2048,
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
You can point any OpenAI-compatible SDK (Python, Node.js, etc.) or curl at the same /v1/chat/completions endpoint.
Agentic Usage
Ornith-1.5-397B excels in tool-calling and agentic coding. It exposes an OpenAI-compatible endpoint with tool calling and works out of the box with standard agent frameworks.
Examples of using Ornith with agents:
Ollama
ollama run hf.co/ornith-ai/Ornith-1.5-397B-GGUF
Atomic.chat
# Atomic.chat loads a GGUF build of Ornith (ornith-ai/Ornith-1.5-397B-GGUF)
# through llama.cpp's OpenAI-compatible API on port 8000.
llama-server -hf ornith-ai/Ornith-1.5-397B-GGUF --port 8000 -c 262144
llama.cpp
# llama.cpp — serve an OpenAI-compatible API on port 8000.
llama-server -hf ornith-ai/Ornith-1.5-397B-GGUF --port 8000 -c 262144
Hermes Agent
# Hermes talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="ornith-ai/Ornith-1.5-397B"
OpenClaw
# OpenClaw talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="ornith-ai/Ornith-1.5-397B"
Unsloth Studio
pip install unsloth
# Load Ornith for fast local inference or fine-tuning (Python):
# from unsloth import FastLanguageModel
# model, tokenizer = FastLanguageModel.from_pretrained(
# "ornith-ai/Ornith-1.5-397B",
# max_seq_length=262144,
# load_in_4bit=True,
# )
Coding CLIs
Ornith-1.5-397B is optimized for terminal-based coding agents. Point any OpenAI-compatible coding CLI at your Ornith-1.5-397B endpoint (set OPENAI_BASE_URL and OPENAI_API_KEY) to understand large codebases, automate tedious work, and ship faster.
OpenCode
# Register your local Ornith endpoint as a provider in ~/.config/opencode/opencode.json:
#
# {
# "$schema": "https://opencode.ai/config.json",
# "provider": {
# "ornith": {
# "npm": "@ai-sdk/openai-compatible",
# "name": "Ornith (local)",
# "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
# "models": { "ornith-ai/Ornith-1.5-397B": { "name": "Ornith-1.5-397B" } }
# }
# }
# }
opencode
Citation
If you find our work helpful, feel free to give us a cite.
@misc{ornith_1_5,
title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
url = {https://ornith.ai/ornith_1_5.html},
author = {{Ornith Team}},
year = {2026}
}