Designing work around the model
Physicist Matthew Schwartz has introduced BootLoops, an open-source toolkit and research method built around a practical question: which scientific tasks fit the strengths of today's language models? His answer is to focus on exact, computation-heavy problems that combine programming, mathematics and broad reading, while leaving scientific judgement and direction with human researchers.
Schwartz describes the target as a “Claude-shaped” problem. Instead of asking a model to behave like an independent scientist, the approach gives it work that benefits from rapid coding, cross-disciplinary recall and the ability to parse technical literature. Experts then decide whether a technically valid connection matters to their field and steer the work toward questions with scientific value.
The project grew from experiments in mathematical physics. Schwartz initially used Claude to implement and improve methods related to scattering amplitudes, which connect particle-collider observations to theoretical descriptions of the underlying interactions. Some modern amplitude calculations use a bootstrap method: researchers apply physical constraints to shrink a large space of possible answers, then use extremely precise numerical values to identify the remaining coefficients.
That combination of formal mathematics, software development and numerical checking provided a useful setting for an AI coding system. BootLoops packages tools and protocols for those calculations, functioning as a harness around a language model rather than as a new model itself. Schwartz says the toolkit can be used with models other than Claude.
Cross-disciplinary reach, with supervision
The same mathematical patterns appeared in subjects beyond high-energy physics, including ecology, population genetics, geology, economics and linguistics. However, a connection that is mathematically correct may still be scientifically obvious, irrelevant or poorly framed. Schwartz therefore sought domain specialists who could identify meaningful questions and assess the results.
That division of labour is central to the project. Current systems can generate code quickly and search through methods spread across multiple disciplines, but they can also pursue unproductive lines and require detailed correction. Schwartz contrasts the breadth and speed of the tools with their limited ability to choose deep conceptual questions independently.
BootLoops is consequently best understood as a workflow experiment rather than evidence that an AI system can conduct science alone. It tries to reduce the mismatch between what researchers often request and what language models reliably do. The proposed gain comes from narrowing tasks, making calculations checkable and involving specialists early enough to distinguish an interesting result from a merely valid one.
The work also offers a more measured account of AI-assisted discovery. Progress depends not only on a model's raw capability, but on tool design, verification and collaboration with people who understand the scientific context.



