So cool, Open-source model LongCat-2.0 matched GPT-5.5 on a Duck Hunt coding run for $0.
Test was done by atomic[.]chat, a desktop app that runs LLMs locally using @kilocode CLI with their agent.
The side-by-side run used 70.3K tokens locally against 64.9K cloud tokens costing $0.65.
The task was not a prompt answer; the agent had to build and revise code.
LongCat apparently handled ducks, waves, ammo, hit physics, falling animation, and the dog retrieval loop well enough to look competitive in a three-iteration agent workflow.
Meituan lists LongCat-2.0 as a 1.6T-parameter MoE with about 48B active per token.
Shows something very practical: for small, clearly defined tasks, a local open model can sometimes produce work that looks almost as good as a frontier cloud model.
So the main difference may stop being quality and start being cost.