The race to build “AI Scientists” may be optimizing for the wrong unit: the agent alone.
This position paper argues that scientific agents should be studied as human-agent systems, where the thing you evaluate is the scientist + agent pair.
Most current systems still treat the human as a supervisor: set the goal, review a phase, approve the final artifact. Far fewer are built for continuous, fine-grained collaboration during the work itself.
The problem is that agents do not naturally know when they need human input. In 10 science tasks, GPT-5-mini almost never asked for help.
But that input matters: in the case studies, experts caught errors the agents missed, while the agents sped up execution. The gain came from collaboration, not autonomy.
The paper’s proposed benchmark is therefore different: does the human-agent team produce better science than either member alone, without collaboration cost overwhelming the gain?
– arxiv. org/abs/2608.14667
Title: "Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems"