Laguna S 2.1 (118B param) looks so great for intelligence per token.
It matched the 753B GLM-5.2 on game coding with 6x fewer parameters.
Both models faced the same job, alongside a third contender named Hy3.
Test was done on atomic【.】chat, a desktop app that runs LLMs locally.
Each had to build 3 arcade games as self-contained HTML files.
The targets were Geometry Dash, Doodle Jump, and Air Hockey. Every game also had to play itself through a built-in bot. So no human touches the controls once the file opens.
The smaller Laguna S 2.1 produced its solutions using just 10.3K tokens total. GLM-5.2 needed 26.4K tokens to reach roughly the same result.