canada-quant

GLM-5.3-Flash-W4A16-MTP

Available on FriendliAI

Dedicated Endpoints

Run this model inference on single tenant GPU with unmatched speed and reliability at scale.

Learn more

Model Details

Model Provider

canada-quant

Model Tree

Base

zai-org/GLM-5.3-Flash

Quantized
this model

Input Modalities

TextImageVideo

Output Modalities

Text

Supported Functionality

Dedicated Endpoints

Explore FriendliAI today

GLM-5.3-Flash-W4A16-MTP API & Inference Endpoint | FriendliAI