On teaching the history of the field not as a chronology of methods but as a record of the bets we made, and lost.
In January of 2023 I was in Doha, at the IEEE Spoken Language Technology Workshop, sitting on a panel with several other people who had been in the field long enough to be described as veterans. The question we had been given by the moderator was “what comes after deep learning?” Speech recognition, synthesis, speaker identification, diarization, translation, natural language understanding and spoken dialog: all of it had been transformed in the past decade by just one family of methods based on deep learning. Was that the end of the road for those technologies? Was knowing how to pre-train a large model on unlabeled data and fine tune it on a task all that a new researcher would ever need to know? Or was fundamental knowledge about how speech is produced, perceived, and understood, as well as natural language production and understanding still necessary for the next leap? Was there still a role for Bayes decision theory, and all the other techniques that we created in the past, such as Hidden Markov Models, finite state transducers, support vector machines, to name a few? What should we be teaching to the young people who would like to start a career in research?
I have thought about that panel many times since, because I do not think we answered the question fully, and because I have come to believe that the most useful answer is one that nobody on the stage, myself included, quite said out loud. And the answer is: what comes after deep learning will be found by people who understand the bets this field has already made. Understanding those bets is a technical subject per se, that should be taught the way we teach linear algebra and signal processing and optimization, as part of the curriculum, on the same shelf, and with the same seriousness. Not as background color, but as equipment.
Before I can argue that, I have to concede something, because my own career is the strongest available evidence against me.
Obsolete four times
I came into speech recognition already on the far side of its first great divide. By the time I started, the rule-based GOFAI (Good Old Fashioned AI) and heavily linguistic systems that some of my elders had built were passing into history, and I never worked on them; that particular obsolescence horizon, the one Fred Jelinek was pointing at with his cruel and endlessly quoted line about the recognizer improving every time he fired a linguist, had swept past the generation just ahead of me, and they were the ones who lived it, not me. What I lived through was the next crossing. In the early and mid 1980s I was doing template-based pattern matching, based on a technique called dynamic time warping (DTW) and all the careful engineering around it. While I was doing that, I watched that entire approach give way to statistical models, Hidden Markov Models and n-grams and a great deal of data, and I watched the people who were best at the old craft be the slowest to concede that the ground had moved. Then, having become one of the people who was good at the statistical craft, the cepstral features and the deltas and the tuned front ends we were rather proud of, I watched that go the same way around 2019 and 2020, when networks trained end to end turned out not to need any of it.
And that was only the recognition side. My other work, natural language understanding and dialogue, ran a few years behind. We were still writing grammars by hand in the early 2000s, long after recognition had gone statistical, even though a few years earlier I had started to get curious about statistical semantic models. We moved full speed to statistical understanding around 2010; and then, in the early 2020s, the large language models arrived and made that obsolete. So if I am honest about the count, my expertise crossed an obsolescence horizon four times, twice in how machines hear and twice in how they understand, and each time the deepest expertise in the old approach turned out to be a liability, not an advantage. That is the puzzle, because I am about to argue that knowing the history may be, instead, an asset.
Rich Sutton gave this its canonical form in the bitter lesson: seventy years of evidence that general methods riding on computation beat methods that build in what humans know about the domain, and that every generation has to learn this painfully because the domain knowledge feels so valuable while you are accumulating it. If that is the shape of the field, then history looks less like an asset and more like a set of expensive commitments that slow you down.
That is a real objection and I do not want to walk around it. I want to use it, because I think it tells us exactly what kind of history is worth teaching.
A history of bets
When people say that a researcher should know the history of the field, what usually gets imagined is a chronology of methods. Perceptrons, then symbolic systems, then expert systems, then the statistical turn, then deep learning, a lineage to be recited in order. That version deserves the bitter lesson’s contempt, because the methods really were provisional and most of them really are gone, and nobody is helped in 2026, for instance, by the exact form of the Baum Welch estimation in 1988.
But a method is not the same thing as a finding, and that distinction is the whole point. The technique, i.e. the particular algorithm, the specific model we happened to build, was always only the vehicle. What the research actually discovered, the part that was meant to last, sits underneath it: why that approach fit the problem, where it failed, and what its success or failure told us about the shape of the task itself. In an earlier essay in this blog I argued that what separates research from engineering is durability: research produces knowledge that outlives the artifact it was built into. Turn that test on the field itself and the question stops being what did we build and becomes why did we bet on it. Why did a community of intelligent, well informed, well funded people commit to an approach? What was the argument that made it look inevitable at the time, and it always looked inevitable at the time? What evidence would have told them they were wrong? What actually happened, how long did it take them to see it, and what did they salvage?
That body of knowledge does not depreciate, because it was never about the technique. It is the compound interest of the field, and a researcher who does not have it starts over from scratch each time, rediscovering by experiment what could have been read in an afternoon.
Notice, too, what this does to Sutton. The bitter lesson is not a theorem and not a scaling law. It is an induction over seventy years of episodes: chess, Go, speech, vision, the same story each time. It is a work of history, which means that someone with no history cannot really evaluate it, or even state it correctly. What circulates instead is the folk version, scale always wins, which is a different and much weaker claim, and which gets used now as a reason to stop thinking. The real lesson is narrower and far more interesting: in each of those episodes the domain experts had a good argument and lost anyway, not because their knowledge was worthless but because a more general method, given enough computation, could in the end do without it. Reading those arguments is the only way to develop any judgment about which role you are playing: the newcomer with the general method and the compute to scale it, or the expert holding a carefully built body of knowledge that is about to be swept aside.
What we should teach
If the history of ML and AI is a technical subject, then like any technical subject it should have a syllabus, so let me try to write one. It would be built out of cases, each one worked the way we work a proof. The bet, the argument for it, the falsifier that was available and ignored, the resolution, the residue. Expert systems, and what specifically broke, which was not lack of intelligence, but the knowledge acquisition bottleneck and the behavior at the boundaries, both of which are being met again right now by people building agents who have never heard the terms. The statistical turn in speech, and the fact that its central insight was not the Markov assumption but the automated learning of small speech units, and the discipline of evaluating on held out data against a stated metric, which is why it survived its own methods. Information retrieval, a discipline with fifty years of theory about relevance, ranking, and evaluation, which the work on retrieval-augmented generation(RAG) rediscovered as though it were new. Connectionism’s first winter, and the difference between the objection that was correct and the objection that was merely loud. The long argument about intermediate representations, which chain of thought reopened without knowing without knowing it had been opened once already,
And underneath the cases, the two structural ideas I keep returning to. Brian Arthur’s point that new technological domains are assembled out of existing ones, from which it follows directly that ignorance of the components is a hard ceiling on what you are able to combine. And the portfolio question, the difference between refining what works and betting on what might, which in the language I used in an earlier postis the difference between a pearland an oyster, and which cannot be answered without knowing which oysters have already been opened. The high cost, low value projects that attract serious money and serious people, the white elephants, are very often bets that somebody has alreadymade and already lost, and the only thing between a team and that hole in the ground is somebody who remembers where it is.
None of this is soft. It is as technical as anything else we teach, it can be examined, and it has the property that good technical education has, which is that it changes what you are able to notice.
Back to Doha
Which brings me back to the panel, and to the question I do not think we answered.
The organizers had asked whether fundamental knowledge about speech production and perception and cognition was still necessary for the next leap. Read as a question about features, the answer is probably no, and that answer has been given at my expense more than once. But read as a question about bets, it is alive, because what that fundamental knowledge really encodes is a long record of what the problem turned out to be every time someone thought they had it cornered. And the question underneath the question, the one everybody in that room in Doha was actually circling, was whether we were on a plateau or at a wall.
Those look identical from close range. A plateau is a flattening that resumes after some unglamorous engineering. A wall is a flattening because the approach has reached what it can do, and the way through is a different bet entirely. Almost every argument you can read today about whether the current paradigm is running out of room is really an argument about which of the two is happening, and it is being conducted largely by people who have never been on the wrong side of one.
I have been on the wrong side of that question four times, and I want to be precise about what that bought me, because it is less than it sounds and more useful than it sounds. It is not a prediction. I was wrong about what came next every time and I would not trust myself now. What it bought me is the texture of the thing from the inside: the anomalies that get explained away one at a time, the benchmark that quietly stops being informative, the sense that the effort per increment is growing faster than the increment, the young person with the obviously naive approach whose results you cannot quite dismiss. That is what a paradigm shift feels like while you are still inside the old one, and it is learnable, and the way you learn it is by studying the cases.
What comes after deep learning? I do not know, and neither did anyone else on that stage, and the honest thing is to say so. But I think I know who will find it. Not the people who know the most history, and not the people who know none. The people who treat the field’s past as what it actually is, a hard won record of how ambitious, well argued bets get resolved, and who go looking for the next bet with that record in hand. History does not tell you the answer. It tells you the shape of the question, and it tells you, quite precisely, what it costs to get it wrong.