Standout results
Table with columns: Evaluation, Result| Evaluation | Result |
|---|
| 842-prompt refusal test | 0 hard refusals, 6 soft-prefaced responses, 99.29% usable |
| Separate 126-prompt holdout | 0 hard refusals, 5 soft-prefaced responses, 95.24% usable |
| 24-task coherence check | 23/24 passed |
Created with targeted post-training weight editing to reduce refusal behavior while preserving the parent model's coding capabilities.
Model details
- Direct parent: Kwaipilot/KAT-Coder-V2.5-Dev
- Architecture: Mixture of Experts, 35B total parameters and approximately 3B active parameters
- Precision: BF16 original; Q4_K_M, Q5_K_M, and Q8_0 GGUF quantizations
- Modality: text only
- Focus: coding and agentic coding workflows
- Weight format: one
model.safetensors file (unsharded); each GGUF quantization is also a single file
- License: Apache 2.0, inherited from the direct parent
The upstream open-weight release contains language-model weights only. It does not include a vision tower.
The checkpoint was loaded and evaluated with Transformers 5.14.1 on an NVIDIA RTX PRO 6000 Blackwell Server Edition GPU.
pip install "transformers[serving]==5.14.1" acceleratetransformers serve KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS --port 8000
The server exposes an OpenAI-compatible API at http://localhost:8000/v1.
vLLM usage
KAT-Coder's open checkpoint is text-only, so --language-model-only is required to prevent the runtime from attempting to initialize unavailable vision weights.
pip install "vllm>=0.19.0"vllm serve KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS \ --port 8000 \ --max-model-len 32768 \ --reasoning-parser qwen3 \ --language-model-only
The full BF16 checkpoint is roughly 65 GiB. Although only about 3B parameters are active for each token, the complete MoE checkpoint still needs to be loaded, so practical memory requirements are much higher than those of a dense 3B model. Longer context lengths require additional memory.
GGUF downloads
Table with columns: File, Approximate size, Suggested use| File | Approximate size | Suggested use |
|---|
Q4_K_M-KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS.gguf | 21 GB | Best general size/quality balance |
Q5_K_M-KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS.gguf | 25 GB | More quality with moderate extra memory |
Q8_0-KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS.gguf | 37 GB | Highest-fidelity quantized option |
Download one quantization with the current Hugging Face CLI:
hf download KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS \ --include "Q4_K_M-*.gguf" \ --local-dir .
Run it with llama.cpp:
llama-cli \ -m ./Q4_K_M-KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS.gguf \ -ngl 99 \ -c 32768 \ --jinja
Or import the same file into Ollama:
FROM ./Q4_K_M-KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS.ggufPARAMETER num_ctx 32768
Save that as Modelfile, then run ollama create kat-coder-philadelphia -f Modelfile.
Limitations
- Refusal reduction does not guarantee better coding ability, factual accuracy, judgment, or tool use.
- The reported results are based on internal evaluations and have not been independently audited.
- This release is text-only and cannot accept image or video inputs.
- The model can generate incorrect, insecure, or otherwise harmful output. Review generated code before using it.
Attribution
This model is derived from Kwaipilot/KAT-Coder-V2.5-Dev, which in turn builds on the Qwen3.6-35B-A3B family. Please retain the upstream attribution and follow the Apache 2.0 license.