๐ Introduction
X-AuT-14layer is a lightweight variant of Qwen3-ASR-0.6B whose audio encoder (AuT, Audio Transformer) is compressed from 18 layers to 14 layers (โ20.7% encoder parameters, 180M โ 140M), while the 28-layer text decoder is kept frozen-intact.
Encoder depth pruning in a frozen-decoder speech LLM is fundamentally an alignment problem, not a capacity problem: removing layers shifts the encoder's output distribution, and the frozen decoder โ unable to adapt โ collapses into premature end-of-sequence prediction and catastrophic word deletion. X-AuT-14layer is produced by a capability reconstruction and injection framework that rebuilds capability level by level:
- Behavior-driven probing selects the pruning configuration through controlled short-budget recovery experiments, revealing that adjacent high-redundancy blocks are safer to prune together than dispersed non-adjacent layers. The model is compressed in two hops of increasing depth: 18 โ 16 (drop the peripheral layers {1, 18}), then 16 โ 14 (drop the probe-selected adjacent block {5, 6}), with cross-hop weight inheritance.
- Quality-tiered data curriculum: a nine-level transcript-quality hierarchy derived from a 280k-hour multi-system agreement pipeline matches supervision quality to each recovery stage's sensitivity to label noise.
- Three-stage recovery pipeline: cross-scale layer-alignment distillation (structural repair) โ hybrid on-policy distillation (behavioral adaptation) โ LoRA finetuning (domain alignment).
- Cross-scale capability injection: a larger Qwen3-ASR-1.7B (AuT-24) teacher does not merely recover lost capacity but injects acoustic robustness the original 0.6B encoder never had โ the teacher is frozen during training and discarded at inference, so the upgrade costs training-time compute only.
On ten public Chinese-English benchmarks, X-AuT-14layer preserves 99.85% of the uncompressed baseline's overall accuracy, and the intermediate 16-layer model exceeds the baseline outright (100.4% relative accuracy). Encoder latency is reduced by 21.4% (PPU) / 11.4% (GPU).
๐๏ธ Model Architecture
Only the Transformer encoder layers of the audio tower are pruned; ConvStem, Bridge, ln_post, and the text decoder are untouched.
AuT (audio encoder) configuration
Table with columns: Item, Value| Item | Value |
|---|
| model_type | qwen3_asr_audio_encoder |
| d_model | 896 |
| Attention heads | 14 |
| FFN dim | 3584 |
| Mel bins | 128 |
| Output dim (bridge โ LLM hidden) | 1024 |
| Encoder layers (original / this model) | 18 / 14 |
| Retained original layers (1-based) | (drop {1,18}, then {5,6}) |
Parameter breakdown
Table with columns: Module, AuT-18 (baseline), AuT-14 (this model)| Module | AuT-18 (baseline) | AuT-14 (this model) |
|---|
| ConvStem (conv2d1/2/3 + conv_out) | 11.03M | 11.03M |
| Transformer encoder layers | 173.62M (18 layers) | 135.04M (14 layers) |
| Bridge (proj1 + proj2) | 1.72M | 1.72M |
| ln_post | 0.002M | 0.002M |
| Audio tower total | โ186.38M | โ147.79M (โ20.7%) |
Whole-model parameters (for reference)
Table with columns: Module, Params, Note| Module | Params | Note |
|---|
| Audio tower (AuT) | โ147.79M | pruned (encoder 180M โ 140M) |
| Text decoder (28 layers) | 440.47M | Qwen3-0.6B, frozen |
| Token embedding | 155.58M | tied with lm_head |
| Total (tied weights deduplicated) | โ743.84M | โ |
๐ Evaluation
Error rates (CER for Chinese, WER for English) on ten public Chinese-English benchmarks; lower is better. Rel. acc. reports overall accuracy preservation relative to the Full-18 baseline (100%), computed as (1โerr_pruned)/(1โerr_baseline) per benchmark and averaged over the suite.
Table with columns: Benchmark, Full-18 baseline, X-AuT-14layer (Stage 1), X-AuT-14layer (Finetuned)| Benchmark | Full-18 baseline | X-AuT-14layer (Stage 1) | X-AuT-14layer (Finetuned) |
|---|
| AISHELL-1 (CER) | 3.33% | 3.52% | 3.39% |
| Fleurs-zh (CER) | 2.80% | 3.49% | 3.32% |
| Fleurs-en (WER) | 4.17% | 5.24% | 5.10% |
| LibriSpeech test-clean (WER) | 2.48% | 3.09% |
The finetuned model keeps all degradations within 1.0 pp of the baseline and still beats it on two benchmarks (LibriSpeech test-clean, CommonVoice zh).
Inference efficiency
Table with columns: Metric, PPU (in-vehicle), GPU (H800)| Metric | PPU (in-vehicle) | GPU (H800) |
|---|
| Encoder latency vs AuT-18 | โ21.4% | โ11.4% |
| End-to-end latency vs AuT-18 | โ4.7% | โ2.6% |
| Peak memory vs AuT-18 | โ4.4% | โ2.8% |
End-to-end improvement is modest because the unpruned 28-layer text decoder dominates total inference time. Encoder-side compression is orthogonal to decoder-side optimization and the two compose multiplicatively.
๐ Inference
This repo contains the full fine-tuned model in safetensors format (already truncated to a 14-layer audio encoder, keys under thinker.*), together with config.json, generation_config.json, tokenizer, and preprocessor โ so it loads directly, with no base-model download or manual layer pruning.
We provide a standalone inference script infer.py that does not depend on the X-AuT training codebase โ only on torch, transformers, qwen-asr, huggingface_hub, and librosa (audio is automatically resampled to 16 kHz mono). The inference logic follows the official Qwen3-ASR recipe: chat-template prompt with a language-control suffix โ processor encoding โ thinker.generate() โ parse_asr_output post-processing.
Setup
pip install -U torch transformers "qwen-asr" huggingface_hub librosa
Quickstart
# checkpoint is auto-downloaded from https://huggingface.co/X-AuT/X-AuT
python infer.py \
--audio /path/to/test.wav \
--lang-code zh \
--device cuda
Or load it directly in Python:
from qwen_asr.core.transformers_backend import Qwen3ASRForConditionalGeneration, Qwen3ASRProcessor
processor = Qwen3ASRProcessor.from_pretrained("X-AuT/X-AuT", fix_mistral_regex=True)
model = Qwen3ASRForConditionalGeneration.from_pretrained(
"X-AuT/X-AuT", dtype="bfloat16", device_map="cuda",
)
Arguments
Table with columns: Argument, Default, Description| Argument | Default | Description |
|---|
--audio | (required) | Input audio path (wav/mp3/flacโฆ), auto-resampled to 16 kHz mono |
--model | X-AuT/X-AuT | HF repo id or a local directory of this checkpoint |
--lang-code | zh | Language control prefix (zh / en) |
Output
One line of JSON per run:
{
"audio": "/path/to/test.wav",
"lang_code": "zh",
"duration_sec": 5.14,
"raw_pred_text": "ๅๅง่งฃ็ ๆๆฌ",
"pred_text": "ๅฝไธๅๅ็่ฏๅซๆๆฌ"
}
raw_pred_text: raw decoded string from the decoder;
pred_text: normalized transcription produced by qwen_asr.inference.utils.parse_asr_output.
๐ Repository Contents
Table with columns: File, Description| File | Description |
|---|
model.safetensors | Full fine-tuned weights (audio tower already at 14 layers, thinker.* keys) |
config.json | Full model config (28-layer decoder + 14-layer audio encoder, retained layer indices) |
generation_config.json, chat_template.json | Decoding and prompt-template configs |
tokenizer_config.json, vocab.json, merges.txt | Tokenizer (Qwen3 BPE) |
License: the weights follow the base model's Apache-2.0 license.
๐ Citation
If you use X-AuT in your research, please cite:
@article{zhang2026xaut,
title = {X-AuT: A Capability Reconstruction Framework for
Encoder-Compressed Speech LLMs},
author = {Zhang, Haojun and Zou, Yi and Chen, Min and
Zhou, Shuchang and Huang, Shiyu},
journal = {Preprint},
year = {2026},
url = {https://x-aut.github.io/}
}