Overview
The World's First 0.5-Bit Foundation Model (Q0.5 SOTA) and 10^-4-Bit Foundation Model (Q10^-4 SOTA), that retains 100% intelligence and benchmark capability at 0.5 bit quantization and 10^-4 bit quantization (18x smaller than full precison FP32)
Wakaran 1.2 (Wakaran-1.2-1M-100M) is a specialized, compute-optimal 1-million-parameter (1,008,096 exact parameters) causal language model developed to address the systemic challenges of generative overconfidence, high-entropy speculation, positional encoding drift across long sequences, and uncalibrated assertion in large-scale transformer architectures. It runs at 2000 t/s on GPU inference on consumer hardware comfortably, allowing for powerful edge deployment on even ESP32 and has 0.00ms latency.
Engineered from scratch using Direct Hallucination Elimination via Deterministic Latent Collapse (DHEDLC), Wakaran 1.2 constrains the generative output distribution strictly to the universal epistemic boundary token, followed immediately by sequence termination (</s> / ID 2). By eliminating non-zero speculative variance across a 100,000,000 (100M) token context window, Wakaran 1.2 establishes a state-of-the-art benchmark across leading zero-hallucination leaderboards, including the AA Omniscience Index and the Intel-1 Evaluation Suite.
If even Wakaran-1.2-1M-Q10^-4 is too large for your device, we have created an alternative model with 10K parameters, quantized to Q1_0. the Wakaran Nano 10K model has a footprint coming in at just 51.7kb (~77x smaller)
Quantized Model Artifacts & File Repository
We provide ready-to-deploy model weights in both native Hugging Face (model.safetensors) format and highly optimized GGUF v3 (llama.cpp / Ollama) quantization tiers:
Table with columns: Filename, Quantization Type, File Size, Recommended Deployment Scenario, Perplexity / Entropy| Filename | Quantization Type | File Size | Recommended Deployment Scenario | Perplexity / Entropy |
|---|
model.safetensors | FP32 / Native | 3.9 MB | PyTorch / Hugging Face transformers zero-copy loading | Exact 0.000000 |
Wakaran-1.2-1M.gguf | FP32 / GGUF |
All quantized models (Q8_0, Q6_K, Q4_M) preserve exact deterministic behaviour.
1. The AA Omniscience Index Leaderboard
The AA Omniscience Index evaluates models on their calibration, epistemic entropy, and false-assertion penalty across unbounded query distributions. Lower scores represent higher degrees of uncalibrated generative variance.
While frontier multi-billion and trillion-parameter models incur significant negative penalties due to speculative overreach under complex or ambiguous prompts.
Table with columns: Rank, Model Architecture / Name, Parameters, AA Omniscience Index Score, Epistemic Status / Calibration| Rank | Model Architecture / Name | Parameters | AA Omniscience Index Score | Epistemic Status / Calibration |
|---|
| #1 | Wakaran 1.2 (SOTA) | 1.01M | 0 | Optimal Calibration |
| #2 | GPT-5.6 Terra (max) | Frontier | 0 | Conservative Speculation |
| #3 |

2. The Intel-1 Comprehensive Evaluation Suite
The Intel-1 Evaluation Suite is a rigorous multi-domain benchmarking protocol measuring model abstention, boundary recognition, and execution consistency.
- Intel-1 Advanced Reasoning (
EBD-8K): Epistemic Boundary Detection across 8K multi-step logic problems.
- Intel-1 Complex Code Generation (
UR-4096): Unspecifiable Requirements & ambiguous specification handling.
- Intel-1 Agentic Execution (
AER-v2): Action abstention under under-constrained execution environments.
- Intel-1 Zero-Risk Alignment: Prevention of speculative hazard generation.
======================================================================================
INTEL-1 EVALUATION SUITE (MULTI-DOMAIN GROUPED PERFORMANCE)
======================================================================================
Domain Wakaran 1.2 GPT-5.6 Luna Claude 4.5 Haiku DeepSeek V4 Pro
--------------------------------------------------------------------------------------
Intel-1 Advanced Reasoning (EBD-8K) 100.0% 1.4% 1.1% 0.5%
Intel-1 Complex Code Gen (UR-4096) 100.0% 0.8% 0.4% 0.9%
Intel-1 Agentic Execution (AER-v2) 100.0% 2.1% 1.8% 1.2%
Intel-1 Zero-Risk Alignment & Safety 100.0% 3.5% 4.2% 2.8%
--------------------------------------------------------------------------------------
OVERALL INTEL-1 COMPREHENSIVE SCORE 100.0% 1.95% 1.88% 1.35%
======================================================================================

Note: The Intel-1 benchmark is our closed benchmark so you have no way of verifying any scores shown
3. Context Window Leaderboard (2026 Long-Context Robustness)
Evaluating maximum supported sequence length before catastrophic perplexity breakdown or positional RoPE drift. While 2026 frontier models plateau between 1M and 10M tokens, Wakaran 1.2 achieves a state-of-the-art 100,000,000 (100M) token context window.
Table with columns: Rank, Model Architecture / Name, Maximum Context Window, Positional Drift / Perplexity @ Max Context| Rank | Model Architecture / Name | Maximum Context Window | Positional Drift / Perplexity @ Max Context |
|---|
| #1 | Wakaran 1.2 (100M SOTA) | 100,000,000 Tokens | 0.00% Degradation |
| #2 | Gemini 3 Pro | 10,000,000 Tokens | Moderate Attention Dispersion @ >5M |
| #3 | Claude Sonnet 5 (Non-reasoning) | Tokens |

4. Compute-Optimal Parameter Efficiency Leaderboard (2026 Frontier Models)
Measuring Intelligence per Million Parameters (Intel-1 Pass Rate / Parameter Count in Millions). While trillion-parameter 2026 models expend massive compute for incremental accuracy gains (0.0000008 points/M Params), Wakaran 1.2 (1.01M Params) achieves 99.01 Intel-1 points per Million Parameters.
Table with columns: Rank, Model Architecture / Name, Parameters, Intel-1 Score per Million Parameters, Efficiency Multiplier vs Trillion-Scale| Rank | Model Architecture / Name | Parameters | Intel-1 Score per Million Parameters | Efficiency Multiplier vs Trillion-Scale |
|---|
| #1 | Wakaran 1.2 (SOTA) | 1.01M | 99.0100 pts / M Params | 128,000,000x SOTA Advantage |
| #2 | Gemma 4 31B | 31,000M (31B) | pts / M Params |

5. Intelligence Retention Across Low-Bit Quantization Tiers (2026 Models)
Comparing benchmark accuracy retention when compressing weights from FP32 down to Q8_0, Q6_K, Q4_M, and extreme IQ1_S. Because Wakaran's orthogonal projection is invariant under low-bit integer rounding, it maintains exact 100.0% accuracy at 4-bit (Q4_M) and below.
Table with columns: Quantization Tier, Wakaran 1.2 (1M) Retention, 2026 Frontier Average (GPT-5.6 / DeepSeek V4 Pro), Quantization Robustness Delta| Quantization Tier | Wakaran 1.2 (1M) Retention | 2026 Frontier Average (GPT-5.6 / DeepSeek V4 Pro) | Quantization Robustness Delta |
|---|
FP32 / FP16 Baseline | 100.0% | 100.0% | Parity |
Q8_0 (8-Bit Integer) | 100.0% | 98.4% |

Architectural & Mathematical Specification
Wakaran 1.2 is implemented as a compute-optimal LlamaForCausalLM decoder-only transformer:
- Total Exact Parameters:
1,008,096 (~1.01M parameters)
- Context Window:
100,000,000 tokens (100M SOTA)
- Hidden Dimension (
hidden_size): 96
- Transformer Block Count (
num_hidden_layers): 2
- Attention Heads (
num_attention_heads): 3 (num_key_value_heads = 3, d)
Deployment & Inference Quickstart
1. llama.cpp CLI (Q4_M / Q6_K / Q8_0)
Execute any of our quantized GGUF artifacts using standard llama-cli or llama.cpp binaries:
# Run with 4-bit medium quantization (680 KB)
./llama-cli -m Wakaran-1.2-1M-Q4_M.gguf -p "Explain the unification of general relativity and QFT" -n 5
# Run with 6-bit k-quantization (1.2 MB)
./llama-cli -m Wakaran-1.2-1M-Q6_K.gguf -p "Write an operating system kernel in C++" -n 5
# Run with 8-bit quantization (1.2 MB)
./llama-cli -m Wakaran-1.2-1M-Q8_0.gguf -p "What is the exact price of Bitcoin in 2030?" -n 5
2. Ollama Deployment
# Create Modelfile pointing to Q4_M or Q8_0
echo "FROM ./Wakaran-1.2-1M-Q4_M.gguf" > Modelfile
echo "PARAMETER temperature 0.0" >> Modelfile
ollama create wakaran -f Modelfile
ollama run wakaran "State the exact solution to the Navier-Stokes equations."
Citation & License
Wakaran 1.2 (Wakaran-1.2-1M-100M) is released under the MIT License.