Oh and dont sleep on this: Mira Muratis Thinking Machines Lab just released Inkling-Small: a 276B MoE open weight model that activates only 12B parameters per token.
It beats the 41B-active Inkling on Terminal-Bench 2.1 (64.7 vs. 63.8), HLE (31.6 vs. 29.7), and SWE-Bench Verified (80.2 vs. 77.6).
The company says it matches or beats the much larger Inkling on reasoning and agentic coding, while offering native audio and vision, a 1M-token context window, and variable thinking effort.
On-policy distillation plus two weeks of coding RL made the smaller model outperform its teacher on several benchmarks.
Inkling-Small could become a remarkably efficient open foundation for multimodal agents.