Frontier models, transparency and trust; a new golden age in research…for some?
An unexpected discovery during the World Cup Final highlights the changing nature of research. Our guest writer John Hammersley shares his thoughts
Whilst much of the globe was focusing on the culmination of the football World Cup, one mathematician made a rather unexpected announcement.
Levent Alpöge tweeted that using Claude Fable 5 he’d discovered a counterexample to the Jacobean conjecture (a longstanding open problem in mathematics)…during the World Cup Final no less! Levent has a PhD in mathematics and has worked at Anthropic, the company behind the Claude series of AI models, since 2024.
The tweet had already amassed over 39 million views by the time this screenshot was taken on the 27th July.
What’s notable about this (aside from the counterexample itself!) is that the author announced it via a tweet, rather than a preprint, and he’s a mathematician working not in a traditional mathematics department but at a private AI company, where he has access to the most advanced AI models and tooling the company has developed.
His tweets included Wolfram Alpha links with details of the counterexample and it rapidly featured on HackerNews, was validated by the community, and within a day was covered by mathematicians such as Terence Tao and in the popular science press. One day! Before discussing the concerns over AI, we should at least take a moment to celebrate the speed of discovery it has enabled.
Indeed, this example follows recent successes of AI in solving hard mathematical problems, in finding various other counterexamples to conjectures that had remained open for years, and enabling mathematicians to find results quicker than ever before. And just as I was finalising this article, OpenAI announced Ten advances in mathematics and theoretical computer science! It could be argued we are at the start of a new golden age of mathematics.
A new golden age…for some?
AI is changing mathematical research, so much so that discussions within the research community culminated in the publication of the Leiden Declaration on Artificial Intelligence and Mathematics in early June, which sets out recommendations, guidelines, and hope for how the field can responsibly use AI in research and publication. Not everyone is optimistic; one researcher describes experiencing a profound spiritual crisis in how AI is taking over the discovery aspect of mathematics, a feeling that is unlikely to be unique.
Even for those who are optimistic, we’re seeing a shift in where mathematics is being conducted. The ability of the frontier models in AI, and specifically those not yet fully publicly released, gives a significant advantage to researchers working at the AI companies developing them – this was the postscript to my first Scholarly Futures article, and it feels even more relevant now.
Research happening at private labs is of course not a new thing: it’s hard to argue against Bell Labs in the US being one of the most important research centres of the 20th Century, for example, and in more recent pre-ChatGPT times, DeepMind’s breakthrough in protein structures with AlphaFold in 2020 (which feels like a lifetime ago now) showed the early ability of AI for research in a private company.
And now Anthropic have announced the launch of a new drug discovery programme: their head of life sciences, Eric Kauderer-Abrams, said the company “will focus on discovering treatments for “neglected” diseases”, as reported by CNBC.
How does everyone else keep up? Can any of us keep up? Can the companies themselves keep up?
OpenAI’s rogue agent
As I was writing this article, HuggingFace released the Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident, which is eye-opening to say the least.
For those unaware, and to quote directly from the above article:
“Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform.”
And this was all after the agent had escaped its sandbox during an internal capability evaluation at OpenAI!
I think Simon Willison described it perfectly in this headline: “OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened”.
There are a number of articles about it, including OpenAI’s announcement, and although some believe it is a marketing stunt I challenge anyone to read the HuggingFace technical timeline and not be at least impressed (aghast?) with how far agents have progressed over the past year in terms of autonomous capabilities.
Pushing the frontier
All this only serves to highlight that the very best frontier models, coupled with the tooling around them and the compute resources to use them, are only available to a select few researchers, often only those within the organisation (Google DeepMind, Anthropic, OpenAI, etc) developing it.
Setting aside the risk that a rogue agent destroys the internet (and modern society in the process), we are likely to see an acceleration in discoveries, in both mathematics and life sciences, and no doubt in other fields soon after. All, or at least the vast majority of these, will come via the private AI labs and the researchers who have access to them.
Taking an optimistic view, this could very well help accelerate and galvanise entire branches of research that are currently stuck or stagnant, generating a huge amount of follow up work and opportunities that encourages further investment into research and the practical applications of these discoveries, which would filter out into wider industry and academia. And in that case, being able to run reproducible experiments with AI, having transparent workflows and trustworthy results will be important; we should be encouraging researchers to find responsible ways to use AI, as the Leiden Declaration attempts to set out.
Taking a pessimistic view, if this filtering out doesn’t happen, it could completely undermine the research ecosystem; if private AI labs take all the discoveries, what motivation or opportunity would there be for any researchers elsewhere? Why should anyone else bother with AI if they’re always going to be beaten to it by the AI labs?
A new crisis for science?
Thirteen years ago when we founded Overleaf and I first got involved in the scholarly publishing scene as something other than an author, it felt like we were stepping into a world on the cusp of change:
“Traditional scientific communication directly threatens the quality of scientific research. Today’s system is unreliable – or worse. Scholarly publishing regularly gives the highest status to research that is most likely to be wrong. This system determines the trajectory of a scientific career and the longer we stick with it, the more likely it will deteriorate.” Curt Rice, writing in The Guardian…in February 2013.
What I felt at the time was a sense of optimism that new technologies making it easier for scientists to share research and to collaborate would help to solve these problems, and make science faster, and more effective. There was a push for open science, a push to make research more transparent and more reproducible.
I still have that optimistic view, that new technologies will push scientific research along faster. Yes, there is misinformation, fake research, and the publication system is still in need of overhaul. But I find it hard to look at the rate of new discoveries and be anything other than excited for the future of research.
As long as a rogue agent doesn’t destroy everything first.



