positron-ai
google_gemma-4-12B-it-ingest-best-gptq-permuted Available on FriendliAI
Dedicated Endpoints Run this model inference on single tenant GPU with unmatched speed and reliability at scale.
Learn more Model Details
Model Tree
Base
google/gemma-4-12B-it
Input Modalities
Text Audio Image Video
Supported Functionality
Dedicated Endpoints
FriendliAI Corp:
San Francisco, CA
Copyright © 2026 FriendliAI Corp. All rights reserved
README License: apache-2.0
Recommended Use
Use this artifact when you need a GPTQ 4-bit build of google/gemma-4-12B-it built by Positron AI.
For general-purpose GPU inference, compare against the original model and other quantized formats before deployment.
Artifact Summary
Table with columns: Field, Value Field Value Base model google/gemma-4-12B-it Published artifact positron-ai/google_gemma-4-12B-it-ingest-best-gptq-permuted Quantization method GPTQ Quantization format gptq Source precision n/a Target runtime n/a
Quantization Details Table with columns: Field, Value Field Value Weight precision 4-bit Activation precision not quantized Bits 4 Group size 64 Symmetric quantization true Activation ordering / desc_act true Damp percent 0.05 Calibration dataset Mixed-domain calibration set Calibration samples 128 Calibration sequence length 4096 MoE experts per token n/a Quantization toolchain GPTQModel 7.1.0, transformers 5.11.0, torch 2.9.1, CUDA 12.8
Evaluation This card intentionally reports no performance or quality metrics (no
KL-divergence, accuracy, or perplexity figures). Validation results are
tracked internally by Positron AI.
Provenance This artifact was produced by Positron AI from google/gemma-4-12B-it. The original model license and usage restrictions continue to apply.