About BroadBox

What does it mean to do science?

Frontier AI can now choose experiments, learn from results, and decide what to test next. BroadBox builds controlled environments — science sandboxes — where agents form hypotheses, run experiments against a sealed oracle, and revise what they believe, so we can measure how well they investigate and work to expand it.

How we came together

We came to BroadBox from different parts of the same problem. Arya's work on scientific sandboxes described controlled environments where an AI agent could decide what to test and learn from the result—not simply answer a static question.

At the same time, our work in laboratory automation made a second possibility tangible: those environments could connect directly to experiments. An agent could propose a hypothesis, request a measurement, receive evidence from the lab, and decide what to do next.

We realized that these ideas belonged together. BroadBox measures AI scientists for the era of autoscience — by watching how well they investigate, adapt, and discover when reality is in the loop — and treats that measurement as something we can deliberately expand over time.

01

Who does science?

If a system can choose an experiment, learn from the result, and decide what to test next, in what sense is it doing science? We may need a broader account of who—or what—can investigate.

02

How do we measure it?

Scientific ability is not the same as answering correctly. We measure how evidence changes an investigator's understanding, and whether it discovered the rule or merely exploited a shortcut.

03

How do we scale it?

The capacity to discover has always been scarce. If some of it can run at computational scale and connect to real data and experiment, we can expand what science accomplishes.

The paper

Scientific sandboxes for AI agents.

The sandbox paper describes environments in which agents propose experiments, receive feedback from a sealed oracle across the wet–damp–dry spectrum of verifiability, and improve their hypotheses over successive rounds. BroadBox turns that idea into a shared, growing platform for measuring and expanding scientific capability.

Preprint — link coming soon

People

Organizing team

Researchers working across scientific agents, experimental design, and laboratory automation.

Arya Rao

Arya Rao

Yasha Ektefaie

Yasha Ektefaie

B

Bryan

Portrait placeholder

Jackson Weir

Jackson Weir

Sandeep Kambhampati

Sandeep Kambhampati

Shantanu Singh

Shantanu Singh