Letting an LLM decide what to try next wastes experiments, and this paper shows that scoring its ideas with standard Bayesian math cuts the count roughly 5×.
The system is called Model Discovery Agent, or MDA. The LLM suggests possible explanations for the data, and MDA picks the one experiment that would best tell those explanations apart.
Nothing gets spent on experiments that only confirm what it already believes.
On a physics benchmark, MDA's model was accurate enough to pass on 93% of runs. The same LLM working alone passed 31%. MDA also matched a published result using 8 experiments instead of about 41.
alphaxiv .org/pdf/2608.09696v3
"M ODEL D ISCOVERY AGENT: LLM- ASSISTED B AYESIAN EXPERIMENT DESIGN"