Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Below is the part I found most interesting

> "However, naively applying FP4 across the entire model causes degradation in complex reasoning, logic, and code generation. Given the MoE (Mixture of Experts) architecture of Xiaomi MiMo-V2.5-Pro β€” where Experts constitute the vast majority of parameters and exhibit the highest tolerance to quantization β€” we selectively quantize only the MoE Experts to FP4 while preserving original precision for all other modules. Through FP4 QAT (Quantization-Aware Training), we dramatically reduce model size and maximize hardware bandwidth utilization while keeping the model's overall capability essentially on par with the original, as shown below"



The 120B and 20B GPT-OSS models by OpenAI did this last year for what it’s worth; the MoEs where MXFP4




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: