Model Summary
Table | |
|---|
| Creator | Zwen AI Labs |
| Architecture | Mistral-7B (DARE-TIES merge + LoRA SFT, fused) |
| Active params | ~7B |
| Context window | 8,192 tokens (32k-capable base) |
| Quantization | F16 → Q4_K_M GGUF (~4.2 GB, < 5 GB RAM target) |
| Hardware target | Apple Silicon (M-series, Metal-accelerated) |
| Serving | Ollama / llama.cpp |
| Languages | English |
The Brain — Five-Corpus Training Mixture
Zwen-Prime is fine-tuned on a deliberate mixture of five datasets, each installing a distinct cognitive faculty:
Table with columns: Corpus, Weight, Faculty installed| Corpus | Weight | Faculty installed |
|---|
| Alpaca Python | 40% | Elite, typed Python and algorithms at correct complexity. |
| Orca Math | 30% | Deep, step-by-step mathematical reasoning with verified derivation. |
| Salesforce XLAM | 20% | Flawless, raw-JSON function/tool calling (Mistral v0.3 convention). |
| LongAlpaca | 10% | Extended-context retention and long-document grounding without drift. |
| Zwen Custom (550 rows) | core spine | Strict logic + Big-O breakdown, full-stack mastery (TypeScript, React/Next.js, Java concurrency), absolute zero-filler output. |
The custom 550-row dataset (zwen_prime_master_dataset.jsonl) is the enforcement spine: 50 Identity rows that set the persona and 500 Algorithmic/Logic rows that lock in the <thinking>...</thinking> → raw-code template across advanced TypeScript generics, React/Next.js App-Router architecture, Java concurrency primitives, and Python system design.
Capabilities
- Strict reasoning core: Every non-trivial answer opens a
<thinking> block with step-by-step logic and a time/space Big-O breakdown, immediately followed by the raw deliverable — no filler, no preamble.
- Full-stack mastery:
- Python — typed (PEP 484/604), async-aware, concurrency-correct algorithms.
- TypeScript — advanced generics, conditional/mapped types, type-stateful builders, discriminated unions.
- React / Next.js — App Router, RSC, Server Actions, route handlers, caching/revalidation, React 19 hooks, streaming.
- Java —
ReentrantLock, StampedLock, Semaphore, CompletableFuture, ForkJoinPool, virtual threads, happens-before reasoning.
- Native function calling: Emits
[TOOL_CALLS] raw JSON matching the provided tool schema; no markdown, no prose around the call.
Merge Details
Merge Method
The base was produced with the DARE TIES merge method via mergekit, using Mistral-7B-Instruct-v0.3 as the density-0 reference base.
Models Merged
dphn/dolphin-2.9.3-mistral-7B-32k — reasoning specialist
theprint/ReWiz-7B — fine-tune specialist
mistralai/Mistral-7B-Instruct-v0.3 — base / reference (density 0, weight 0)
Merge Configuration
base_model: mistralai/Mistral-7B-Instruct-v0.3
merge_method: dare_ties
dtype: bfloat16
out_shard_size: 1.2B
parameters:
density: 0.5
weight: 1.0
normalize: true
int8_mask: true
rescale: true
lambda: 1.0
models:
- model: dphn/dolphin-2.9.3-mistral-7B-32k
parameters:
density: 0.5
weight: 1.0
- model: theprint/ReWiz-7B
parameters:
density: 0.5
weight: 1.0
- model: mistralai/Mistral-7B-Instruct-v0.3
parameters:
density: 0.0
weight: 0.0
Fine-Tune Stage
The merged base was then LoRA fine-tuned on the five-corpus mixture and the adapters permanently fused into the base weights (peft.merge_and_unload). The fused model is exported to GGUF for local serving.
DARE-TIES merge → LoRA SFT (5-corpus mixture) → merge_and_unload → GGUF → Ollama
Serving
The Ollama Modelfile lives at the repository root and references the F16 GGUF. To serve:
ollama create zwen-prime -f Modelfile
ollama run zwen-prime
For the < 5 GB RAM Apple-Silicon target, quantize the F16 GGUF to Q4_K_M with llama.cpp first, then point FROM at the quantized file.
System Prompt
Zwen-Prime ships with a dedicated system prompt (model-workspace/Zwen_System_Prompt.txt) that sets the Principal-Engineer persona, the <thinking> mandate, the zero-filler output protocol, full-stack mastery expectations, native JSON tool-calling rules, long-context discipline, and the hard constraints. It is embedded in the Modelfile's SYSTEM directive.
Intended Use
Zwen-Prime is intended as a local, autonomous engineering copilot: designing and shipping production code across Python, TypeScript, React/Next.js, and Java; reasoning through math and algorithms with verifiable steps; and calling tools via raw JSON when integrated into an agent runtime.
Limitations
- 7B scale: strong on focused engineering tasks; not a frontier-class generalist.
- Wall-clock-sensitive: derived math is verified symbolically in-context but not executed; verify numerically when stakes are high.
- Function calling follows the Mistral v0.3 / XLAM JSON convention and expects a tool-aware runtime to dispatch
[TOOL_CALLS].
- Fine-tuned toward zero-filler directness; users expecting conversational preamble will not get it.
License
Apache-2.0, inheriting the licensing of the upstream Mistral-7B base and merged specialists. Verify against the terms of each upstream model before redistribution.