Two prominent mathematicians are drawing a line between what large language models can do in mathematics and what they still appear unable to do, according to The Decoder. Timothy Gowers and Peter Sarnak both credit current systems with serious mathematical capability, but argue that the models remain weak at the kind of creative abstraction behind major new proofs. The distinction matters because recent AI progress in math is often used as evidence of broader reasoning gains. The Decoder reports that Gowers sees today’s models as effective at combining known methods and trying many possible search paths. The limitation, in his view, is not effort but judgment: models lack the intuition to select the small number of promising routes from a very large space of possibilities. Sarnak’s reported critique is similar but framed around abstraction. According to The Decoder, he says AI can derive results from existing theory, yet fails when the task requires developing the abstractions that support major proofs from an elementary starting question. That is a narrower and more precise claim than saying LLMs are bad at math; it says they are comparatively strong at working inside inherited structures and weaker at inventing the structures themselves. The Decoder also points to related work from DeepMind researcher Tom Zahavy. In a paper titled “LLMs Can't Jump,” Zahavy reportedly identifies a bottleneck he calls “manipulative abduction”: the ability to invent new foundational assumptions where there is no linguistic precedent. The Decoder says Zahavy sees world models as one possible path forward, but the item does not establish that such systems have solved the problem. Taken together, the reported assessments fit into a broader debate about what LLM performance is actually measuring. The Decoder frames the question as whether models are becoming more generally versatile or whether they are mainly improving at benchmarks and familiar problem domains. The mathematicians’ argument, as reported, is that success at calculation, recombination, and search does not by itself prove the capacity for original mathematical invention. For AI labs and users, the practical takeaway is to separate tool value from frontier claims. A model that can combine known techniques and explore options may still be useful to mathematicians and engineers. But if these critiques hold, performance on structured tasks should not be treated as evidence that the system can reliably originate new mathematical frameworks. Who benefits: AI teams building tools for formal reasoning, proof assistance, and mathematical search benefit from a clearer description of where current systems may be useful. Researchers studying world models and reasoning architectures also get a sharper target problem. Who's exposed: Companies marketing LLMs as broadly creative research agents are exposed if users expect genuine mathematical invention rather than recombination of known methods. Buyers relying on benchmark results alone may overestimate capability in unfamiliar problem spaces.