BootLoops AI: How Matthew Schwartz Built a Claude-Powered Physics Harness

BootLoops AI is interesting for a reason that has little to do with chatbots sounding clever. It is an attempt to turn a large language model into something closer to a scientific computing operator, one that can choose tools, write code, run calculations, and then prove that its answer survives independent checks.

Built by physicist Matthew Schwartz with Claude, BootLoops is an open-source LLM harness for exact quantitative science. It is not a new foundation model, and it is not tied to Claude. The same harness is designed to work with Claude, Gemini, ChatGPT, or another capable agent. Its real contribution is the machinery around the model: scientific software, workflows, tests, acceptance gates, and protocols that define what “done” actually means.

That distinction matters. Schwartz’s broader research campaign reportedly produced 36 manuscripts across 18 fields with 19 human collaborators after screening roughly 400 candidate problems. That does not mean Claude independently published 36 papers. It means the model was embedded in a large, heavily supervised research system that tried to match AI strengths to problems that were computational, interdisciplinary, and unusually easy to verify. Pasted text

1. What Is BootLoops AI, and What Is an AI Harness?

BootLoops AI infographic showing model, agent and harness as three nested layers
BootLoops AI infographic showing model, agent and harness as three nested layers

If you’re asking what is an AI harness, the simplest answer is this: a model supplies the intelligence, an agent takes actions, and a harness gives that agent the tools, procedures, checks, and working environment needed to complete a real task reliably.

BootLoops is an AI harness built specifically for quantitative science. Instead of asking an LLM to reason from text alone, it gives the model access to numerical solvers, symbolic tools, exact arithmetic, scientific code, verification routines, and reusable research protocols.

BootLoops AI: Key Facts About the Scientific LLM Harness

Key FactWhat It Means
CategoryScientific LLM harness, not a standalone AI model
Main PurposeExact and checkable quantitative research
Model SupportDesigned to be independent of Claude and usable with other agents
Original DomainScattering amplitudes and Feynman integrals
Core IdeaLet the model choose and combine scientific tools, then force the result through tests
VerificationIndependent calculations, high-precision numerics, reproducible scripts, and acceptance gates
Open SourceToolkit and related repositories are publicly available
Best FitProblems with precise numerical or symbolic outputs that can be checked independently

The model independence is especially important. Claude was used to build BootLoops 1.0, but the toolkit is meant to evolve separately from whichever LLM drives it. That makes BootLoops an AI harness example rather than a Claude feature. The scientific workflow is the product.

2. How the BootLoops AI Harness Actually Works

BootLoops AI workflow infographic with symbolic and numeric routes meeting at verification
BootLoops AI workflow infographic with symbolic and numeric routes meeting at verification

The central idea is not “ask Claude a harder question.” It is to break scientific work into stages where software can do something concrete and where each important claim can be tested.

BootLoops AI Workflow: How the Scientific Harness Produces Verifiable Results

StageWhat BootLoops DoesWhy It Matters
Problem DefinitionStarts from a named integral, dataset, or published quantityKeeps the task specific
Tool SelectionReads a common index of scientific packages and methodsLets the model combine expertise across fields
Symbolic AnalysisReduces integrals, studies singularities, and identifies function classesNarrows the possible answer space
High-Precision NumericsEvaluates quantities to many digitsCreates an independent numerical target
Exact ReconstructionUses tools such as integer-relation methods to recover exact constantsTurns digits into analytic structure
VerificationRecomputes the result through a separate routeMakes model confidence irrelevant
ReviewUses skeptical agents and human inspectionCatches bad assumptions and bad interpretations

In the scattering-amplitude workflow, BootLoops combines a top-down route with a bottom-up one. Symbolic constraints press down on the space of possible answers, while high-precision numerical calculations push upward from direct evaluation. When the two routes meet, the candidate result can be checked at precision far beyond what a persuasive explanation could fake. The toolkit describes comparisons at 30 digits or more, sometimes much higher.

That is the intellectual center of BootLoops: the AI does not earn trust by sounding like a physicist. It earns provisional trust only when the computation survives external tests.

3. Why Matthew Schwartz Built BootLoops Instead of Simply Prompting Claude

Schwartz’s earlier experience with Claude exposed a frustrating mismatch. The model could behave like a fast research assistant, but getting a scientifically valuable result required constant steering. It wandered, chased weak ideas, and needed sentence-level correction.

He changed the question. Instead of asking, “How do I make Claude behave more like a human scientist?” he asked, “What kinds of scientific problems already fit what Claude is good at?”

Those strengths were unusually relevant to mathematical physics: broad knowledge across fields, strong coding, fast paper reading, familiarity with mathematics and statistics, and enough patience to manipulate large computational systems. Schwartz’s source material explicitly contrasts those strengths with deep conceptual scientific judgment, where current models remain much weaker.

That shift produced what he calls Claude-shaped science.

3.1 What Is a Claude-Shaped Scientific Problem?

A Claude-shaped problem is not simply a difficult problem. It is a problem whose structure matches the strengths of current agentic AI.

The sweet spot is quantitative, well specified, computationally tractable, spread across several technical literatures, and independently verifiable. Coding should matter. Tool choice should matter. Breadth should matter. Most importantly, the final answer should face a test stricter than “the explanation looks plausible.”

This helps explain why Claude physics results can look impressive in narrow mathematical domains while the same model may struggle with open-ended questions about which physical theory is correct. Exact calculations and conceptual scientific taste are different jobs.

4. What Did Claude and BootLoops Actually Solve?

The most concrete early result came from Feynman integrals, mathematical objects that appear in quantum field theory calculations. Schwartz first had Claude port fragmented methods and software into a common framework. Claude then pushed beyond comparatively familiar logarithmic cases into elliptic integrals, where relevant expertise is scattered across physics and mathematics.

Schwartz reports that the system completed 30 frontier integrals end to end. Fifteen reproduced known results using the new workflow, while 15 had not previously been computed by that method.

That matters for two reasons. First, reproducing known answers is a powerful control. A scientific harness should prove it can recover established results before anyone gets excited about new ones. Second, the new calculations tested whether the same machinery could move into harder function classes rather than merely automate familiar textbook exercises.

The work then spread into neighboring areas. BootLoops was applied to cosmological correlators, large-scale structure, black-hole scattering, gravitational-wave memory, energy correlators, string-theory calculations, and lattice problems. In one example highlighted by Schwartz, the toolkit was used on the exact return probability for a three-dimensional random walk with unequal hopping rates, a lattice-integral problem connected to work dating back to George Watson in 1939.

The point is not that one AI suddenly mastered all of physics. The same mathematical machinery appears in different fields, and an LLM can sometimes notice those bridges faster than specialists who live inside one literature.

5. How BootLoops Tries to Prevent Hallucinated Science

Scientific hallucination is not solved by telling a model to “be accurate.” BootLoops takes a more useful approach: design the task so that a wrong answer has multiple chances to fail.

For scattering amplitudes, the toolkit emphasizes numerical checks at arbitrary or very high precision, standalone scripts, error certification, and test points that were not used to fit the answer. Schwartz also describes adversarial review by separate agents and repeated literature checks instead of trusting model memory.

This is close to the difference between software testing and asking a programmer whether their code “looks right.” The latter may be useful. The former is what you build a system around.

The so-called BootLoops Standard is especially revealing. For certain diagram calculations, a result is not considered solved merely because an analytic expression exists. It must also be evaluable independently to arbitrary precision. That forces the result out of the model’s narrative world and into a numerical one.

6. How Physicists and Developers Can Use BootLoops

The repository’s recommended starting point is deliberately unromantic. Clone the toolkit, run its self-tests, start your preferred agent inside the repository, and give it a specific scientific object.

That object might be an integral from a paper, a dataset, or a published number you want reproduced. The model is then asked to inspect the available tool guides and propose a plan before any serious computation runs. The repository says a fresh installation can validate its tool packages through the supplied self-test command, and the toolkit is designed so missing external engines are reported rather than silently ignored.

A practical workflow is: define the object, let the agent map it to available instruments, review the plan, run a small test, demand independent verification, inspect plots and assumptions yourself, and only then scale up. If a project produces a useful new tool or protocol, that capability can be folded back into the harness.

This recursive improvement is where the name starts to feel appropriate. Every solved problem can leave behind better machinery for the next one.

7. What Problems Should You Give BootLoops?

Good BootLoops problems tend to have a hard boundary between success and failure. A difficult integral is a good candidate. Reproducing a published numerical result is another. So is testing an analytic formula against independent numerical integration, porting scientific code, exhaustively searching a constrained space, comparing competing calculations, or checking whether a method from another field applies.

A vague prompt such as “derive quantum gravity” is almost the opposite of a BootLoops-shaped task. There is no clear tool path, no crisp acceptance test, and no obvious independent certificate that tells the system when it is finished.

That suggests a useful rule for anyone experimenting with an AI physicist workflow: choose problems where the answer can push back.

8. Can BootLoops Turn Anyone Into an AI Physicist?

It lowers the barrier to sophisticated scientific computing, but it does not erase expertise.

A programmer, student, or researcher can use an LLM to reach software and mathematical techniques that might once have taken months to learn. BootLoops can also connect fields in ways a specialist may not immediately see. That is a genuine change in access.

But Schwartz’s own cross-disciplinary work shows the limit. When Claude moved into ecology and population genetics, technically correct results were not automatically scientifically important. Domain experts were needed to say which questions mattered, which assumptions were meaningful, and how the work should be redirected. In the ecology example, expert input changed the research direction from a technically impressive calculation to a model with more scientific value.

So the useful picture is not “AI replaces physicist.” It is human problem selection, AI search and computation, independent verification, human interpretation, then another loop.

9. Where Claude Still Fails

BootLoops exists partly because the raw agent is unreliable in predictable ways.

Schwartz reports that Claude can declare victory too early, misjudge how long a task will take, continue grinding through an inefficient calculation, lose context over long projects, and draw conclusions that do not follow cleanly from correct computations. It can also have weak scientific taste, preferring technically solvable questions that experts do not consider important.

These are not small issues. A calculation can be numerically correct and still answer the wrong question. A proof can be “complete” except for the lemma that contains the entire difficulty. A cross-disciplinary connection can be real but trivial.

The harness helps with verification, organization, and acceptance gates. It does not manufacture judgment. Human review is still the layer that decides whether a result makes scientific sense and deserves anyone’s attention.

10. Is BootLoops Doing Science or Just Coding Faster?

“Just coding faster” undersells what happened. Claude did more than autocomplete. It ported methods, built scientific software, searched candidate problems, combined techniques from different fields, and executed calculations that fed into research results.

But calling that a fully autonomous scientist overshoots in the other direction.

Science includes choosing worthwhile questions, understanding the meaning of assumptions, judging novelty, interpreting results, and deciding what evidence would change your mind. In Schwartz’s account, those functions repeatedly came from humans and domain specialists.

BootLoops is better understood as a way to expand the computational reach of a research team. It compresses the distance between “someone published a method in another field” and “we can test whether that method solves our problem.”

11. What BootLoops AI Could Change About Scientific Research

The most important lesson from BootLoops AI is not that an LLM has become a general-purpose scientist. It is that the right harness can make current models far more useful by matching them to problems with strong tooling and strong verification.

For physics, that could mean more researchers can explore difficult integrals, reproduce results, translate methods across subfields, and test ideas without first becoming experts in every software package involved. For science more broadly, it suggests a near-term model that is less cinematic but more practical: humans choose the questions and supply scientific taste, while agents search, code, calculate, and test at a scale humans cannot sustain alone.

BootLoops does not eliminate the physicist. It changes how much computational territory one physicist can cover.

For more grounded analysis of AI research tools, scientific agents, and the systems turning models into working research infrastructure, follow Binary Verse AI.

 What is BootLoops AI?

BootLoops AI is an open-source harness for using large language models in exact quantitative science. Matthew Schwartz built BootLoops with Claude by combining scientific software, agent instructions, verification protocols and acceptance tests so that an AI can perform calculations whose results can be independently checked. It is a harness around an AI model rather than a new AI model itself.

Does BootLoops only work with Claude?

No. Claude was used to build BootLoops 1.0 and powered Schwartz’s original research workflow, but BootLoops was deliberately designed to be model-independent. Its documentation says it can be driven by Claude, Gemini, ChatGPT or another capable agent framework.

What physics problems has BootLoops AI solved?

Schwartz reports that BootLoops completed 30 frontier Feynman integrals end-to-end, including 15 reproductions and 15 previously uncomputed cases. Related work spans scattering amplitudes, elliptic integrals, cosmological correlators, large-scale structure, black-hole scattering, energy correlators, string theory and Watson’s longstanding anisotropic random-walk problem. Some broader BootLoops projects were still undergoing verification when the system was announced.

Can anyone use BootLoops AI to become an AI physicist?

Anyone can access the open-source tools, and BootLoops can make advanced computational techniques much easier to use. But it does not remove the need for physics knowledge: Schwartz found that humans still had to choose meaningful questions, inspect calculations, challenge conclusions and determine whether technically correct results were scientifically important.

How do I start using BootLoops AI for a physics problem?

Start with a concrete object such as an integral, dataset or published calculation rather than asking the AI to “discover new physics.” After installing BootLoops and running its self-tests, launch your agent in the repository, have it read the tool documentation and propose a calculation plan, review that plan, and require independent verification of the result. The repository explicitly recommends this plan-before-computation workflow.

Leave a Comment