v1.7 run scope: 1,024-token context, 24 steps, intended as a testable business-owner adapter rather than a final long-context bakeoff
v1.7 train runtime: 1663.7101s
Release note: Kaiju's current product path may combine a compact model planner with deterministic harnesses and verifier checks. If the shipped experience uses that harness path, release copy must say so plainly instead of implying the raw model weights alone create every artifact.
Data
The dataset is source-backed and RMDW-owned or RMDW-authored. The current source inventory is tracked in release/SOURCE_INVENTORY.md.
High-level categories:
Website/UI
Coding
Debugging
Automation
Tool-use
Strategy
Business
Business-suite
Identity
Excluded data:
Closed-model outputs from OpenAI, Anthropic, Gemini, or similar providers as supervised training completions.
Customer private code without explicit permission.
Client-specific website text, contact details, contracts, or private business details unless explicitly reviewed and approved.
Secrets, credentials, private keys, tokens, cookies, and raw support logs containing personal data.
Evaluation
Required bakeoff before release:
Base Qwen 3.6 27B
Kaiju Coder LoRA
GLM 4.7 production baseline
Current local harness evidence:
2026-06-03 Kiyomi business-suite router hard gate: 23/23 passed.
Business-suite prompts: 2/2 passed.
Static artifact checks: 23/23 passed.
Dataset validation: 1,689 reviewed candidate examples across 14 files.
v1.7 target gate: all category minimums met, including business_suite and proposal.
v1.7 served adapter smoke:
Website task website-barber-001: passed, 2,726 chars in 174.49s.
Proposal task with Kaiju API system prompt: passed, 4,306 chars in 232.27s.
Sellable-candidate gate:
Beats base Qwen on RMDW practical evals.
Near or above GLM 4.7 on highest-value customer tasks.
No critical safety failures.
Produces complete artifacts instead of plans only.
Produces owner-ready Kiyomi/RMDW artifacts for websites, connector packs, CRM, reporting, leads, sales, ROI, and operator training.
Distinct useful voice without becoming gimmicky.
Limitations
Known limitations:
Not a general frontier model.
May be weaker than large cloud frontier models on broad reasoning and uncommon programming domains.
Needs a strong harness for tool use, file editing, and long-running work.
Raw merged serving is slow on this SGLang stack.
Dynamic SGLang LoRA serving is not release-quality for this adapter; use the merged model path.
Business-owner performance depends on source-backed evals, provenance controls, and deterministic artifact verification.
Hosted API release requires billing, rate limits, abuse controls, logs, and rollback.
Intended Use
Good fit:
Solo-founder product work.
Small-business websites and automations.
Kiyomi-style local AI product workflows.
Practical coding and deployment assistance.
Not a fit:
High-risk medical, legal, financial, or safety-critical decisions without expert review.
Secret handling without a secure app layer.
Claims of guaranteed correctness.
Release Status
Current status: business-owner release-candidate preparation.
Fresh v1.7 and v1.8 LoRA training finished on 2026-06-03 after clearing old ComfyUI/Ollama workloads from Gojira B. The current completed testable product path is the v1.8 merged model plus the deterministic business-owner harness and verifier. Raw merged model testing works for focused business-owner documents and backend automations, but the paid website path remains harness-first until broader raw website evals pass.
Do not publish weights or sell hosted API access until the eval and release checklist pass.
v1.8 status: completed on 2026-06-03 and merged into a full local model for serving; do not publish externally until human review, upstream notices, broader comparison evals, and raw website limitation language are complete
v1.7 serving config: SGLang over Tailscale at http://100.109.109.14:18083/v1, model kaiju_v17_business_owner, context 4096, memory fraction 0.90.
v1.8 training metrics: runtime 11666.7564s, train loss 0.9281658741335074, adapter present.
v1.8 dynamic SGLang LoRA caveat:
Adapter-name-only serving can be base-equivalent.
Corrected selector qwen36-27b:kaiju_v18_business_owner crashes with LoRA buffer shape torch.Size([8192, 16]) does not match weight shape torch.Size([14336, 16]).
Dynamic LoRA is not the release serving path for this checkpoint.
Kaiju Coder 7 current serving config: vLLM bitsandbytes runtime
quantization on Gojira B at http://100.109.109.14:18084/v1, exposed on
this Mac through http://127.0.0.1:18181/v1, model kaiju-coder-7,
current OpenCode context 16384. SGLang has historical 32k benchmark
evidence, but 32k should be freshly restarted and re-confirmed before being
called the live default.
v1.8 merged endpoint probe: 1,155 visible chars in 60.17s.
v1.8 merged focused eval:
Proposal rerun: 1/1 paid-ready, 4.0/4.0, 4,014 chars in 212.72s.
Jah credits backend: 4.0/4.0, 9,718 chars in 566.36s.
Broader base-Qwen, GLM, and raw website comparisons are still pending before
any superiority claims.