most benchmarks test if a model can solve a problem knowing everything up front, but what makes SlopCodeBench super interesting is that it discloses parts of the problem incrementally, forcing the LLM to redesign the codebase on the fly, lest it suffer the growing slop mountain
if you're thinking about building software factories or dark factories - this ep with @vaibcode is worth a watch
https://x.com/dexhorthy/status/2085764187507265980