You might imagine mathematicians popping corks after recent AI solutions to questions that had long flummoxed their field. Instead, many reacted less as if drenched by Champagne than by a tsunami. “My 13-year-old daughter joked unprompted that, if she wants to become a mathematician, it now looks like she has maybe two more weeks,” the computer scientist Scott Aaronson wrote.
The irony is that AI was itself borne of mathematics. Our guest writer, the mathematician and educator Junaid Mubeen, explains how once-abstract concepts underpin the practical technology of our time: the large language model.
—Tom Rachman, AI Policy Perspectives
By JUNAID MUBEEN
Mathematicians take pride in irrelevance. ‘Here’s to pure mathematics!’ begins one toast. ‘May it never be of use to anyone!’
Yet there is an amazing afterlife to mathematics, with concepts that once appeared abstract becoming invaluable later. When mathematicians first studied primes (numbers such as 29 that cannot be divided into two smaller factors), they did so to sate their curiosity, with no inkling that these would form the basis of cryptography. The concept of imaginary numbers (square roots of negative numbers) was an abstraction made manifest. Yet today’s descriptions of quantum physics would be bereft without them. Fractals (self-repeating shapes) were studied for their wonder-inducing patterns. Now, they play a role in the design of everything from antennae to cities.
These accidental real-world uses speak to what the physicist Eugene Wigner termed the ‘unreasonable effectiveness’ of mathematics. Wigner was struck by how highly complex maths turns out to have such descriptive power for understanding the real world. Today’s large language models, LLMs, are mathematics at its most unreasonably effective because the statistical tools that can write poetry, diagnose illness, even provide personal advice, are actually shaped by a handful of relatively simple mathematical concepts.
The term ‘model’ just means a mathematical description—a way of using numbers to describe some phenomena. With LLMs, numbers are used to represent words (embeddings). When an LLM is trained on text, it assigns meanings to words, and makes guesses at what should come next, seeking a mathematical model that captures its training data. This involves hefty calculations, with techniques from algebra (matrix multiplication) and, to minimise errors, from calculus (gradient descent).
In the past couple of years, the cutting edge of LLMs has been reasoning models, which devote more computation to generating answers—not just dishing out plausible responses but planning, evaluating possible outputs, and adjusting according to how developers mathematically ‘rewarded’ the system during a fine-tuning process (reinforcement learning).
Let’s explore the mathematics underpinning each concept in turn.
Embeddings
The training of an LLM involves breaking down text into fragments (tokens), and processing these tokens as long strings of numbers. In essence, we need a mathematical way of describing words numerically. And we have such a way, thanks to a system analogous to geographical coordinates.
To describe any point on the globe, we need just two numerical values: latitude and longitude. Locations that are close to each other have similar longitude/latitude values. The further apart these values are, the more distant the locations. Now imagine we were able to do something similar with words, describing them using numbers, and plotting them in some space.1
This is a word embedding, which means that similar meanings (‘dog’ and ‘puppy’) are close together, while words that bear little relation to one another (‘dog’ and ‘diplomacy’) are far apart. The location of a word’s meaning is thereby turned into numbers, and known as its vector.
In the diagram below, the words ‘man’, ‘woman’, ‘king’ and ‘queen’ are each plotted as two-dimensional points.
Here’s the clever part. We can also consider the direction of these words relative to one another. You can see above that the line connecting ‘man’ to ‘woman’ is parallel to the line connecting ‘king’ to ‘queen’. These parallel lines all head in the same direction, as if they all convey a particular concept—perhaps ‘gender’. So the geometry of word embeddings captures additional meaning.
I have kept this example deliberately simple because two dimensions can be visualised. But mathematicians pay no mind to the ease with which we can illustrate concepts. Routinely, they play around in high-dimensional spaces that break human perceptual limits. The idea is to think of the data points as coordinates. The French mathematician and philosopher René Descartes devised these—legend has it that he was lying ill in bed one day and, seeing a fly buzzing around the room, wondered how we might describe its position using just numbers.
We can take the same idea and apply it to four, five, or any number of dimensions. To describe a point in 20-dimensional space, you would have a coordinate with 20 values. That is mightily useful when creating word embeddings because language is so complex that we may need hundreds, even thousands, of dimensions. It’s not entirely obvious what each of those dimensions represents (this is part of the mystery of LLMs), and we can no longer visualise the word embeddings. Yet these spaces seem to capture a meaningful notion of conceptual proximity between words.
When the mathematician Bernhard Riemann proposed high-dimensional spaces in the 19th century, many of his peers responded to the idea with astonishment. They certainly could not have imagined that such ideas would eventually be seized upon to represent words, and usher in an era of artificial intelligence.
One potential hitch with embeddings is that words with similar meanings may be placed far apart simply because one is more rare than the other. A model may not realise, for instance, that the word ‘loquacious’ is actually just a less-common term for ‘chatty’. So we need a mathematical way of capturing the notion of synonyms.
The idea is to focus on the relative direction of these words. This is known as ‘cosine similarity’ which, if you do not recall high-school trigonometry, essentially means it has something to do with angles. Suppose for simplicity’s sake that we are in two dimensions, and we draw lines from the point (0,0), which we call the origin, to the points representing the two words. Rather than focusing on the distance between the points, we look instead at the angle between these two lines. While ‘chatty’ may be far away from ‘loquacious’ in terms of their locations in the embedding, the two lines corresponding to these words are separated by a small angle, which signals that they are similar in meaning, despite one occurring far more frequently than the other.

Matrix Multiplication
But a language model is more than just how word meanings relate to each other. It’s also generating new text, predicting mathematically what combination of words constitute a fitting response. Using something called a transformer, LLMs contextualise words and information more generally.
Imagine you ask a model to complete this sentence: ‘The cat sat on the ___’. Each word has the form of a vector in an embedding. From this, the model assigns a probability to each candidate word—perhaps ‘bed’ is assigned a 2% likelihood, ‘floor’ is assigned 8% and ‘mat’ is assigned 75%. The system is most likely to finish the sentence with ‘mat’ but there is a bit of deliberate randomness in its word selection, to stop the generated text from seeming robotic and repetitive.
LLMs generally express intelligent answers because of their versatile comprehension of language. For instance, the vector for ‘book’ may sit close to ‘author’ and ‘hardback’ in some embeddings. But if we’re talking about the need to ‘book a medical appointment’, then the model moves this instance of ‘book’ closer to ‘doctor’ and ‘illness’. That adaptive intelligence is due to the transformer, which updates its understanding of words by analysing them in context.
The key is attention: the transformer surmises the meaning of the entire string not by considering each word in isolation, but by paying attention to them in the context of surrounding text. That means calculating many values at once, a seemingly daunting task.
Luckily, back in the 19th century, the mathematician Arthur Cayley realised that rich maths could live inside a grid of numbers known as a matrix. Since we may have many points in a dataset, and multiple features to consider, the matrix offers a way of describing this information.
Below, for instance, is a matrix showing data for five homes.
For each home we also have something we’re trying to predict—namely, the home price. In one column, this is referred to as a vector (akin to the word vectors from earlier).
The question is how to predict these prices from the previous matrix. This is where model parameters come in. Each feature has an associated number—a weight—that describes how it affects the value we’re predicting. Suppose we trained the model, and found the following weights for home size, number of bedrooms, and number of stories:
For the first home in our dataset, these weights are multiplied by the corresponding value for each feature (respectively, the square metres, the number of bedrooms, the number of stories). When these products are added together, we get a prediction for the house price (in thousands of pounds):
(2*65)+ (15*3)+(20*1 ) = 195
This is not far off the true value of 190. The mathematical calculation that pulls all this together for all houses in our dataset is called matrix multiplication and looks something like this. (The squiggly symbol for ‘approximately equal to’ reminds us that the model is just making a best guess for each home price.)
It’s a complicated looking equation for sure, but the main idea is that it describes all five houses in our dataset at once. For large language models, it is grids of words that are multiplied all at once, allowing the model to process each word alongside other words in the surrounding text.
The ‘largeness’ of large language models comes from the sheer number of factors and parameters: not one or two or even a handful, but hundreds of billions. One aspect of how they work is a scaled up version of what we’ve just seen. The reason that matrix multiplication is a mainstay of AI is that it scales gracefully—a computationally efficient way of handling so many calculations at once.
Gradient Descent
When developers ‘train’ a machine-learning model, they’re setting its parameters, the numbers that determine how a model behaves. They do this by calculating mathematical ‘gradients’ and adjusting the parameter values until the model produces desirable outputs.
But, as the statistician George Box famously quipped, ‘All models are wrong, but some are useful’. If a model is bound to make prediction errors, the best recourse is to minimise those errors. In other words, the goal is optimisation, seeking the best possible result in a given situation. Fortunately, there is an entire branch of maths that deals with such problems: calculus.
Calculus was developed independently by Gottfried Leibniz and Isaac Newton.2 The latter’s primary concern was using maths to describe continuous motion, from the swinging of a pendulum, to the orbit of planets, to the flow of water. Newton’s calculus gave precise definition to how quickly things speed up and slow down—their rate of change. Later mathematicians made calculus an abstract study of anything that changes, from the way a virus rapidly spreads, to the way radio waves ebb and flow.
In War and Peace, the Russian novelist Leo Tolstoy expounds on the role of calculus in studying change through history, imploring the reader to analyse history in ‘infinitesimally small units’, which is the central idea of calculus. Today, applications of calculus in AI are themselves changing history.
In an AI model, the goal is to minimise prediction errors with respect to the training data. When a model is being trained, we already know what outputs it should be producing and so a lot of effort goes towards reviewing a model’s predictions to see just how far off the mark they are.
In mathematical speak, we have an error (or ‘loss’) function that tells us how wrong the model’s predictions are. We are seeking parameters that minimise the overall amount of error, reaching the ‘minimum point’ of the error function.
One mechanism for hunting down the minimum is gradient descent (“gradient” is just another word for slope), which dates back to 1847, when it was studied by the French mathematician Augustin-Louis Cauchy.
The idea behind gradient descent is intuitive. Imagine you are descending a mountain range, seeking the lowermost point. What would be your strategy? One approach is to walk around in a small circle, gauging the steepness at each point. Identify the spot that has the steepest downward slope, and head in that direction for a while. Stop and repeat, each time taking a few steps downwards. There is a reasonably good chance that this method leads you to the lowermost point in the vicinity. That is the gist of gradient descent, which is a go-to algorithm—that is, step-by-step instructions that a computer can follow—for training AI models.
‘If I have seen further,’ Newton said, ‘it is by standing on the shoulders of giants.’ Today, AI practitioners stand on the shoulders of Newton, reaching ever higher from the foundation of his calculus.
Reinforcement Learning
When a large language model generates an output, it is making a best guess at how a piece of text might continue, based on statistical patterns it has picked up in its training data. But its foundational objective is to continue text in a manner that is plausible. And plausible doesn’t always mean correct, leading to factual errors and ‘hallucinations’.
One method to improve its outputs is reinforcement learning, or RL, a reward-based technique that gained prominence in 2016, after DeepMind’s AlphaGo defeated Lee Sedol in the notoriously complex game of Go. The developers trained this system on millions of moves by human experts, followed by simulated games against itself, to reward winning strategies. For a system that lacks feelings, the ‘reward’ amounts to getting a higher number. An algorithm automatically updates the system’s parameters so that its future behaviour is more likely to bring more rewards.
But applying RL in large language models raises the question of what constitutes a ‘winning’ response. This judgement is sometimes made by human observers who evaluate different responses, with their approval or disapproval feeding into the algorithm that mathematically rewards and updates the language models to thereafter generate more desirable outputs: reinforcement learning with human feedback, or RLHF.
The mathematics behind reinforcement learning has roots in a field known as ‘dynamic programming’. The term was coined in the 1950s by the mathematician Richard Bellman, whose eponymous equations provided the basis for many reinforcement learning applications. Dynamic programming is a way of taking a large sequential problem (What is the best action?) and breaking it into smaller, more manageable sub-problems (How good are the future situations that this action might lead to?). For LLMs, the ‘large’ problem is generating an entire sequence of text, and the sub-problem is determining the best next word (or token).
Bellman’s original equations involve storing all future states of a problem in a giant lookup table. This approach is not viable with language, given the unfathomable number of possible word combinations. This is where neural networks come in. Building on Bellman’s ideas, today’s LLMs can take a good output (say, a large passage that has been highly rated or a 100-line maths proof that has been validated) and trace it back to its earlier stages. This then informs the model on what kinds of partial responses are likely to lead to better outputs. During inference, a reasoning model is thus able to abandon paths that tend to result in low-value outcomes, and instead pursue reasoning steps that bear more fruit.
‘Dynamic’ refers to the fact that decisions about what actions to take are made by considering how those choices affect possible future states, while ‘programming’ connotes the idea of planning and decision-making. But most of all, Bellman claimed that he was simply seeking a term that did not involve ‘mathematics’, given how unfashionable the subject was in the 1950s, and how any such association may have hindered his prospects of funding. In his autobiography, Bellman spoke of the hostility that the US secretary of defense Charles Wilson had for research, which did not bode well for mathematical projects.
But, yet again, ‘irrelevant’ mathematics powered the future.
If AGI Arrives, Give Mathematicians Their Credit
Maths and AI have always enjoyed a close relationship. The great Alan Turing was a mathematician whose ideas form the basis of modern computers. Other pioneers—from John McCarthy, who coined the term “artificial intelligence,” to Marvin Minsky and Claude Shannon—all earned doctorates in mathematics. Indeed, one of the earliest AI applications was the ‘Logic Theorist’, which took aim (with mixed success) at a collection of maths problems.
The mathematical concepts that underpin today’s LLMs are relatively simple, yet the resulting models, trained on unfathomable amounts of data, are now taking aim at maths itself. An off-the-shelf chatbot can answer the vast majority of problems in the high school and even undergraduate maths curriculum. Models at the leading edge are showing impressive capability in mathematical research. The mathematical community is starting to grapple with the implications; many have endorsed the Leiden Declaration, one of a number of public statements that caution against improper use of AI in research, such as publishing papers that lack proper citation or rigorous human peer review, and accepting opaque proofs that are incomprehensible to humans.
For many, the startling progress of LLMs in fields like maths is further evidence of their inexorable march towards artificial general intelligence, or AGI—systems that can match humans across every cognitive domain. If that day comes, we should reserve some credit to mathematicians of centuries past, from Descartes (coordinates in space) to Cayley (matrix multiplication) and Riemann (high-dimensional spaces), not to mention Newton (calculus) and Cauchy (gradient descent).
“Here’s to pure mathematics!” they say. But the second part of the mathematicians’ toast—“May it never be of use to anyone!”—is a conjecture that has now been soundly disproven.
Junaid Mubeen—director of the Parallel Academy, an online initiative to increase the number and diversity of excellent young mathematicians—is the author of two books: Mathematical Intelligence: What We Have that Machines Don’t and Think Like a Mathematician.
Strictly speaking, while older models such as Word2Vec worked with actual words, today’s LLMs work on word-fragment tokens.
It is Leibniz who is credited with the Chain Rule formula that today’s backpropagation algorithms are based on, allowing large neural networks to trace errors back through the network.











