Base model
Qwen/Qwen2.5-VL-7B-Instruct
Eight uniformly sampled frames are each rendered at -2, 0, and +2 stops. The 24 images are passed in temporal-major order.
Training data and recipe
Five adapters were trained independently on the five content-separated BrightVQ splits. Training uses two epochs, a three-epoch cosine schedule horizon, learning rate 1e-4, micro-batch 1, gradient accumulation 8, and rank-16 LoRA with alpha 32 and dropout 0.05. MOS targets are interpolated across five quality words. The root adapter is split 0; splits/split-1 through split-4 contain the remaining adapters.
Training data: BrightVQ.
Metrics
Held-out metrics for the five 420-video test splits:
Table with columns: Split, SROCC, PLCC, KRCC, RMSE| Split | SROCC | PLCC | KRCC | RMSE |
|---|
| 0 | 0.9110 | 0.9120 | 0.7347 | 5.7437 |
| 1 | 0.9311 | 0.9287 | 0.7666 | 5.0971 |
| 2 | 0.9175 | 0.9245 | 0.7469 | 5.0977 |
| 3 | 0.8904 | 0.8957 | 0.6996 | 5.9874 |
| 4 | 0.8760 | 0.8925 |
Intended use
This adapter is intended for research on no-reference perceptual quality assessment of user-generated HDR video. Scores are not calibrated for other datasets, display pipelines, or video domains.
Code and input construction are available in BrightRate-LM.
Citation
@article{saini2026brightratelm,
title = {BrightRate-LM: Representation-Aware Quality Assessment for User-Generated HDR Video},
author = {Saini, Shreshth and Wang, Yilin and Birkbeck, Neil and Adsumilli, Balu and Bovik, Alan C.},
journal = {Machine Vision and Applications},
year = {2026},
note = {Submitted}
}
Links
Code and evaluation: github.com/shreshthsaini/BrightRate-LM. Dataset: BrightVQ on Hugging Face. Related papers: Beyond8Bits, CVPR 2026 (arXiv 2603.00938) and CHUG, ICIP 2025 (arXiv 2510.09879).