Tencent released Hy3 today. Beating GLM-5.1 in blind tests while running fewer active params than the models it's competing with
295B MoE, 21B active, 256K context. And it does this on plain GQA. No sparse attention, no MLA. So the efficiency isn't coming from architectural tricks yet, there's still headroom left.
That should worry the competition more than the benchmark numbers do.
It’s truly crazy what kind of efficient models are coming out of China.