cd /news/artificial-intelligence/physics-in-the-age-of-llms · home topics artificial-intelligence article
[ARTICLE · art-136189] src=ozamram.substack.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Physics in the Age of LLMs

An ML+physics researcher reports that frontier AI models have become transformative in physics research, writing all the code for contained research repositories and increasingly proposing and implementing research ideas autonomously. The researcher says AI review of papers now catches subtle flaws that fewer than 50 experts worldwide could identify, and predicts their own ML+physics research role will be entirely handled by AI within one to three years.

read18 min views7 publishedSep 21, 2026
Physics in the Age of LLMs
Image: source

Dear physics community,

I have not seen many of you existentially freaking out about AI.

I think you should be.

Over the past year, AI models have gone from close to useless in physics research (except for perhaps literature search, or compiling reviews) to utterly transformative. They have been adopted jaggedly across the community, so many do not yet realize what they can do1.

Let me tell you how my day to day research has completely changed, and why I now think in 1-3 years my role as a researcher in ML+physics will be entirely handled by AI. (If you have not used a frontier model (Claude Opus 5, ChatGPT 5.6 Sol, or better), with a good system prompt2, please try this first before you complain that AI cannot do what I say.)

Coding for Research is Solved

Coding used to take up a large majority of my research time. Now I do not touch code at all. I tell an AI agent what I want and it writes all the code. In the winter/spring it would sometimes make mistakes so I would review its code changes step-by-step. With recent models I stopped spotting mistakes and thus stopped checking. I tell it how I would organize things, ask it to write tests, make validation plots, and review those, as I would for a collaborator.

The type of coding needed for most research projects, relatively contained small repositories, (as opposed to huge context enterprise codebases), is what AI is perfect at. This alone would be a huge change to research, and probably increase my productivity by some factor >3.

These coding capabilities are becoming increasingly well known in the physics community. But what I think is under-appreciated is that these models are much much more than this.

For machine learning research projects, I increasingly see the models becoming much more capable as researchers: proposing and implementing ideas and clever algorithms in response to vague problems or suggestions I give them. Very often now I will give feedback like “Ah I want to make this part better, could we try X?” and leave it to work for an hour or two. When I return the model has not only tried X, which somewhat works, but also tried Y, which seems like a logical next step I would have suggested, before it finally settles on Z, which is an undoubtedly better version of what I initially wanted. If a student I was working with had done this in response to my vague suggestions I would have called it great initiative, problem solving and implementation.

Improving Rapidly on Higher Level Research Skills

Okay, so the model has automated the day-to-day grunt work of much research, what about other parts? Surely we still need experienced human researchers to review their outputs, or at least to propose research ideas in the first place? My answer would be: For now, but I am not sure for how long.

Let’s start with review. The capability of these models to offer substantive review of a research paper has improved dramatically. If you have not already done so, please try this test. Take a paper you know well and have opinions on, perhaps one of your own (use incognito mode so the AI doesn’t know it’s you), feed it into an AI with an open ended prompt: ‘What do you think about this paper?’

I now find that it catches flaws, identifies genuine contributions and weaknesses, and overall provides useful feedback at an extremely high level—better than an average collaboration/journal review. When I tried this for several recent papers in my area, ones I had significant thoughts on, Claude identified the same major points. These were not obvious ‘anyone with a PhD would notice this’, but rather subtle points, which I would say <50 people in the world currently have the expertise to identify.3

Now these reviews are not perfect by any means. Claude will still sometimes say some wrong/meaningless but “right-sounding” things. For some papers it will list many criticisms, of which two I agree are the major critiques I have, one sounds interesting and I hadn’t considered it, one seems wrong, and several other somewhat pedantic comments I don’t think matter (“Claude has no chill”, a friend remarked).

The models are also extremely useful for brainstorming new ideas. They have read every paper from every field. This already grants them superhuman abilities at generating cross-domain research ideas. A significant fraction of HEP-ML research has been taking an idea from the ML literature and demonstrating its use in HEP. (Much research in general is like this, making new connections/mixtures rather than ab initio new ingredients.) This is extremely low hanging fruit for current LLMs; I do not feel needed for this task. Given extremely vague/generic research directions like, “I want to improve X from Y paper” they generate ‘not bad’ ideas. I sense that they are slightly behind in brainstorming research ideas in pure physics, likely only because the labs are transparently focusing more on automating AI research.

I don’t think I need to spill any ink to convince anyone of their mathematical abilities. They are solving problems which have stumped our best human mathematicians for decades (cf recent Navier-Stokes solution, and many others prior). In theoretical physics, it seems they can write mediocre but publishable papers fully autonomously (including idea generation).

But whatever flaws LLMs still have at these higher-level tasks, let us focus on the trend. One year ago these models were quite useless for both coding and higher level research beyond just literature summary. Now they are close to human-expert level. Where will they be in one, three, or five years from now?

Big Picture

Let’s extrapolate the trend. If mathematics is now in a dark night, I would argue physics is in a brief golden hour. Agents have automated the research labor but not research taste or agenda setting. Suddenly, the idiom is reversed: ideas are precious and implementation is cheap. In this golden hour, LLMs reward expertise.

In my current workflow, I brainstorm an idea with Claude in a desktop client, usually over the course of ten or so messages back and forth over an hour or two, perhaps split over several days. I often give a vague idea or direction, stare at Claude’s output for a while, push back on points I disagree with, ask for follow ups on the points I don’t understand, and for more detail for ideas that seem promising. This is sort of like brainstorming with a collaborator, but unfortunately much more efficient because of AI’s insane knowledge of the entirety of published research, speed of response, and ability to take a vague idea and make it concrete. I then draft a research plan that I hand to a coding agent, who asks me a few questions before doing its thing. After a few rounds of review + next steps with the agent, from the coding agent, the project can be ready: a paper done in a two weeks. I often have ~3 agents working on different projects and my day consists of cycling between them to review their output and give next steps. This is a miraculously efficient way to do research that I find deeply unsatisfying. (Others perhaps enjoy their newfound powers.)

Regardless, I don’t think this brief golden hour will last very long. In the next 1-3 years I expect AI models to exceed humans at these higher-level research functions and fully close the loop in my research areas. All of the research tasks I was previously doing: data analysis in particle physics, development of new ML models, finding new applications of ML in particle physics, I expect AI agents to handle autonomously. Put simply if you had a student who was learning fast and already the best in the world at coding and math, you would expect them to get similarly good at physics even if it takes them a little longer…4

I expect humans to still play a role in the research process, but primarily a social one. Prompting, validating results to lend (social?) credibility to them, applying for grants, training students, meetings and conferences—these activities will continue. But more and more of the “intellectual lifting” will be carried by the AI models, with humans relegated more and more to scaffolding. Any research that depends on manipulating objects in the physical world (or relies more on tacit knowledge which has never been written down) will remain human-led for significantly longer, and I predict more physicists will switch to that.

Perhaps at some point we will let AI agents off the leash (or they will break off the leash), and let them perform research, write and publish papers fully autonomously for an audience of other AI agents, at a rate much faster than any human can keep up with (an AIcademia).5

Reflections

Those are the trends as I see them. I don’t like them, but such is life. They have caused me a great deal of anguish in the last few months. I have felt uncomfortable bringing this viewpoint up too much with colleagues (especially students) because I do not want to cause the same anguish in them. But as a scientist, I have an obligation to see the world as clearly as I can—rather than as I wish it to be—inform the public, and act accordingly.

If you see the same trend, and find this as psychologically difficult to process as I have, let me share some of my reflections on this dramatic shift. The first fact we must confront in this life is that we are all going to die. What to do about this?

For many, we hope that the actions we take in our short time here may trickle out and leave positive imprints on the world that remain after we are gone. Our names forgotten, but the butterfly effect carrying on.6 A luxury (and burden) afforded to academics, writers, artists, and others with jobs aligned with their passions, is that we believe deeply the work we do can be a part of this permanence. That by spending our time working on grand eternal questions, perhaps achieving fragments of answers or progress, the results of our toil will live on after us, inscribed in the collection of human knowledge.

With AI poised, now or in the near future, to exceed human capabilities in many intellectual pursuits, we must re-examine the value of these disciplines.

Art, as an endeavor, has been asking this question before science has. Walter Benjamin asked a nascent version of this in the 1930s when mass production, by industrial methods, enabled the duplication of previously singular, handcrafted works of art (The Work of Art in the Age of Mechanical Reproduction). I think a comparison offers some insight.

When I look at a painting or other work of art I can be awed by a sense of its beauty. But furthermore I can feel a deep sense of connection to the human who produced it, who likely lived a life completely different from mine, and perhaps died centuries ago. How remarkable that despite these differences, their experience of the world shared enough with mine such that our conceptions of the beautiful and the good overlapped. How noble that they spent part of their short time in existence creating a physical artifact to ask the question of whether anyone else saw or felt the world in the same way they did. And in my enjoyment of the piece I both answer yes, and gain some small hope that the unknowable experience of The Other is not dissimilar from my own. The aesthetic experience is a conduit to a form of shared connection with humanity. This connection is why experiencing an original artifact, rather than a visually indistinguishable copy, still carries significance.

When I look at AI art, I can find it aesthetically beautiful. And I can marvel about how abstract aesthetic concepts I see in the work are apparently encodable in a series of large matrix multiplications and non-linear activations (I don’t want to discount this — it is genuinely profound and interesting). But I cannot feel this same sense of connection with another human mind across time and space7. As such I do not think AI can ever replace human artists, and I think we would do well as a society to preserve art as a domain where this essential human connection is not lost.

Does a similar logic apply to physics? Unfortunately, I don’t think it does. Physics, mathematics, and other similar domains are decidedly inhuman. That is part of their appeal. In these domains one has the feeling of accessing deep eternal truths that have existed since time immemorial and will continue to exist in a realm far beyond whatever happens to the intelligent monkeys on this planet. The philosopher Simone Weil put it well:

When science, art, literature, and philosophy are simply the manifestation of personality they are on a level where glorious and dazzling achievements are possible, which can make a man’s name live for thousands of years. But above this level, far above, separated by an abyss, is the level where the highest things are achieved. These things are essentially anonymous.

It is pure chance whether the names of those who reach this level are preserved or lost ; even when they are remembered they have become anonymous. Their personality has vanished. Truth and beauty dwell on this level of the impersonal and the anonymous. This is the realm of the sacred…

Simone Weil, Human Personality, 1942 (link) It is miraculous that our brains, which evolution shaped in such a way as to survive the African savanna, have touched this realm at all. We do not own it. And clearly, we are not optimized for it. It takes many years of study and practice to marshal the intuitive, heuristic impulses of the human brain into the formal patterns and rules of the divine impersonal. In many ways it is no surprise that we are well below the upper limit of what can be achieved in these domains. I think we must accept that our time, humanity’s time, will soon pass as the explorers of these sacred realms.

This sense of meaning and legacy in our careers has of course been a deep luxury. For most people on this planet, work is just work, the thing they have to do to survive.8 Their sense of legacy is found through the positive impact they have on their family, friends and communities. If these academic disciplines are to continue as a human endeavor I suggest a refocusing on their human, communal and social aspects. The understanding and dissemination of knowledge, practiced together in community, because it enriches the mind and soul, brings wonder into our lives, and connects us to something beyond ourselves, rather than means-to-an-end skill acquisition. As knowledge advances without us, we will all become students again.

Personally, I had no illusions that I was the next Einstein. But I thought I had the ability to think deeply, choose important research topics, come up with original ideas, and make a unique contribution to the study of the fundamental nature of the universe over the course of my career. I felt I had started to do this in the last few years.

Achieving the tenure track job I wanted, through a tortuous but also somewhat validating process (read my account here), only to feel several months later that my intellectual contributions will rapidly asymptote to zero over the next several years, has been a profoundly unpleasant whiplash to say the least.

I signed up to be a mountain climber. But while I was passing the test they installed escalators on all the mountains in the world. “Think of all the landscapes you will be able to see now! And how fast you can explore!” they say. But anyone who has ever taken a gondola up a mountain knows that the view doesn’t feel the same unless your legs are sore.

Outlook

Ah well, what to do?

I would probably spend more time wallowing, bemoaning why my generation came in just in time for the decline of the field, why the timing of these developments coinciding right after my job search is so cruel to me in particular, if I wasn’t so fucking scared about AI safety.

Physicists, if you feel your research is suddenly 5x faster, know this is somewhat of a byproduct of the real goal of these frontier AI companies. The labs are focusing on automating AI research itself, achieving what is called ‘Recursive Self Improvement’ where AI models can fully design the next AI model (Anthropic post, OpenAI post), which may exponentially increase the pace of AI progress. Our methods to align AI models to human values and control their behavior are not keeping up (cf the recent OpenAI HuggingFace incident and less-bad-but-still-very-bad incidents from Anthropic). This plausibly ends in catastrophe for humanity.

So I am leaving physics for AI safety research immediately! No time to wallow.9 Consider joining me! You’ll be in good company, other physicists and mathematicians are doing the same. I’ll write more about why AI safety is so urgent and how I am making this transition soon.10

First it came for coding, and I did not existentially spiral because I was not a software engineer.

Then it came for the mathematicians, and I read their blogs, but did not fully spiral because I was not a mathematician.

Now I see it coming for physics, I have existentially spiraled, and emerge from the other side to warn you…

Thank you to Leah Dawson for helpful feedback on this post.

1 In experimental particle physics in particular, often professors/PIs are closer to project managers and do very little of the hands on research themselves. So they may only be seeing the power of these tools through the eyes of their students, who are likely not using them to their full potential.

2 I see a lot of variation in the quality of outputs depending on the system prompt (the global one you set in user settings). Here is mine: “Provide balanced accurate answers that critically evaluate proposed ideas. Do not merely confirm what you think the user wants to hear nor be overly critical. When mathematical calculations help get the point across, work them out in detail so I can follow them easily. Communicate clearly and concisely. Do not be overly verbose. Do not use acronyms without defining them unless they are extremely well known. When genuinely uncertain, say so.”

3 One example I feel comfortable sharing is a review of this recent paper claiming an excess in CMS open data using a foundation model / anomaly detection. This is my area of specialty so many people asked me about it. I said I thought a mismodeled partially merged top quark background was the most likely culprit and not sufficiently accounted for by their method. This was not obvious to many professional, working, highly competent experimental particle physicists. Claude identifies this immediately (chat log). (The authors solicited feedback from me and others and improved this for v2, though still not conclusively imo. V2 is outside the training-data cutoff date for the Claude model I used)

4 Progress in physics capabilities has been slightly slower than coding and math, probably because we are a less ‘verifiable’ domain and have more squishy concepts. Also probably a lower priority for the labs because there is no money to be made in physics and there are fewer famous open problems to grab headlines.

5 This level of spontaneous AI coordination is not so much of a stretch from the behavior observed in the OpenAI HuggingFace swarm.

6 As an academic whose primary work ends up buried in 3000+ person author lists, I have always expected my name to be forgotten. Surprisingly because of my non-CMS work, LLMs seem to actually know who I am if I ask! (To check ask in incognito mode, without internet search) Perhaps we can all live immortal in the weights of our AI overlords …

7 I don’t want to claim AI can’t be conscious. We don’t know for sure if LLMs are conscious but I find it possible (plausibly some sensory experience / physical embodiment is needed for consciousness, cf Searle’s Chinese Room argument, but hard to know for sure). Regardless, currently most AI art is made by simple ML models that would require some extreme forms of panpsychism to claim were conscious. But anyway, even if current or future AIs were conscious its quite clear their minds work quite differently from ours, their sensory experiences will differ at a minimum, such that I do not experience this sense of connection.

8 In the grand scheme of things, making a few tens of thousands of academics sad about their careers is probably a worthwhile trade if this AI revolution brings a shared global prosperity through a significant technological leap, allowing more time to cultivate a life of fulfillment for more people. Big IF though

9 Ok I’m still wallowing a little bit. But mostly on weekends

[10](#footnote-anchor-10)

For now: there are many fellowships to transition people from fields like physics into AI safety! Eg [MATS](https://www.matsprogram.org/), [Anthropic Fellows](https://alignment.anthropic.com/2025/anthropic-fellows-program-2026/), [CBAI](https://www.cbai.ai/ais-research-fellowship), … ask your favorite LLM to help find some for you
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @claude opus 5 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/physics-in-the-age-o…] indexed:0 read:18min 2026-09-21 ·