Introducing GeneBench-Pro — testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires.
Problems would take a human expert around 20-40 hours to complete.
GPT-5.6 Sol is a big step forward.
We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the rig...