The bitter lesson: how scale beat cleverness

Between 2012 and 2022, AI stopped being a collection of clever tricks and became a matter of data and computing power. This is the decade that put superintelligence on the table.

A silver humanoid AI beside a transparent screen showing neural network layers and training charts.

In 2019, the computer scientist Richard Sutton published a short essay that many researchers still find uncomfortable to read. He called it The Bitter Lesson. Looking back over seventy years of AI research, he argued, the biggest lesson is that general methods which make the most of computation always win in the end, and by a large margin. The clever ideas researchers were proudest of, the ones that encoded human insight into the machine, kept losing to brute force.

It is bitter because it says that decades of careful work mattered less than everyone hoped. It also explains, better than any other idea, how we went from the AI winters to a serious public debate about superintelligence in about ten years.

The image that changed everything

The turning point has a date. In 2012, a neural network called AlexNet entered the ImageNet challenge, a contest to label photographs across a thousand categories. It was built by Alex Krizhevsky and Ilya Sutskever, two students of Geoffrey Hinton in Toronto.

AlexNet’s top-5 error rate was about 15 percent. The next best entry scored about 26 percent. In a field used to fighting over fractions of a point, this was a landslide.

Nothing about AlexNet’s core idea was new. It was a deep neural network trained with backpropagation, an approach from the 1980s. What had changed was everything around it: more than a million labelled images, and graphics cards designed for video games that happened to be very good at the arithmetic neural networks need. The old idea had finally been given enough fuel.

Within a few years, deep learning had swept through speech recognition, translation and computer vision. Approaches that research groups had refined for decades were replaced by networks that learned the task directly from data.

Move 37

If AlexNet convinced researchers, AlphaGo convinced everyone else.

Go had long been considered the last great board game out of reach for machines. The number of possible positions is so large that Deep Blue’s brute-force search was hopeless. Good play depends on intuition, on a feel for the shape of the board that masters struggle to put into words.

In March 2016, DeepMind’s AlphaGo played the champion Lee Sedol in Seoul and won four games to one. The moment people remember came in the second game. On its 37th move, AlphaGo placed a stone in a position that commentators first took for a mistake. Professional players estimated that a human would almost never play it. It turned out to be decisive.

Move 37 mattered because it was not copied from anyone. AlphaGo had learned from human games and then from millions of games against itself, and it had found something its teachers had missed. For the first time, a machine had shown a flash of what looked like creativity in a domain humans considered deeply intuitive.

Attention, and then language

The architecture that would carry AI into language arrived in 2017, in a Google paper with a confident title: Attention Is All You Need. It introduced the transformer, a network that processes a whole sequence at once and learns which words should pay attention to which others.

Transformers had one property that turned out to matter more than any other. They scaled. Make them bigger, give them more text, train them longer, and they kept getting better, with no obvious ceiling in sight.

OpenAI pushed that property further than anyone. GPT-2, in 2019, wrote paragraphs coherent enough that the company initially held back the full model over misuse concerns. GPT-3, in 2020, had 175 billion parameters and could do things nobody had explicitly trained it for: translate, summarise, write simple code, answer questions, all from a few examples in the prompt.

The laws of scale

In 2020, researchers at OpenAI led by Jared Kaplan published Scaling Laws for Neural Language Models. They showed that a model’s performance improved in a smooth, predictable way as three quantities grew: the size of the model, the amount of data, and the computing power used to train it.

This was a strange and powerful result. It meant progress could be planned. If you wanted a better model, you did not need a breakthrough. You needed a bigger budget.

In 2022, DeepMind refined the recipe. Its Chinchilla model, with 70 billion parameters, outperformed its own much larger Gopher model at 280 billion, because it had been trained on far more text. The lesson was the same, only sharper: scale, used well, beats almost everything.

Why this decade changed the question

Before 2012, superintelligence was a thought experiment. The field could barely recognise a cat, and speculating about minds greater than ours felt like science fiction.

The scaling decade changed that in two ways. It produced systems whose abilities surprised even their builders, abilities that appeared at scale without anyone designing them. And it offered a plausible path forward that required no new scientific insight, only more of what was already working.

If intelligence can be bought with computing power, then the question of superintelligence becomes a question of how much we are willing to spend.

That question stopped being hypothetical at the end of 2022, when a chatbot was released to the public and a hundred million people started talking to a language model. What happened next, and why chatbots were only the beginning, is the subject of the next article in this series.