
Three Ways to Be Wrong About the Past Tense
Chapter 1
Imported Transcript
Arshavir Blackwell, PhD
Three Ways to Be Wrong About the Past Tense
Arshavir Blackwell, PhD
I ran Plunkett and Marchman’s recipe and got nothing. Then the obvious explanation failed too. Then my own measurement turned out to be quietly flattering my conclusion.
Arshavir Blackwell, PhD
I''m Arshavir Blackwell, and this is Inside the Black Box.
Arshavir Blackwell, PhD
Across four posts I argued that language models don’t learn the past tense the way children do — that they slap -ed on everything first and memorize exceptions like went afterwards, the reverse of a child’s order, and that a child-like U-shaped dip appears only if you deliberately engineer a break in the model’s training diet.
Arshavir Blackwell, PhD
That last claim has an obvious objection, and it comes from the people who answered Pinker and Prince the first time.
Arshavir Blackwell, PhD
Kim Plunkett and Virginia Marchman were both frequently present in my advisor’s lab, so I got to sit in on seminars where they discussed this work. They ran the past-tense network again without the abrupt vocabulary jump that made the 1986 model so easy to attack. They grew the training set gradually, from about 20 verbs to around 500, and reported that it still reorganized itself, shifting from memorizing verbs one at a time to handling them systematically, with the transition depending on how many regular and irregular verbs were present rather than on any sudden break (Plunkett & Marchman, 1993). On the child side, Marchman and Bates (1994) found the same shape in 1,130 toddlers: overregularization switches on once vocabulary passes a threshold — a “critical mass.”
Arshavir Blackwell, PhD
Which means my “you need a break” claim was answering an argument nobody was making. Every experiment I’d run held the vocabulary fixed and varied only how often verbs appeared. Gradual growth, their actual condition, I had never tested.
Arshavir Blackwell, PhD
So I ran it. Start with twenty verbs, fourteen of them irregular, roughly a toddler’s first verb lexicon. Grow to five hundred over training, adding mostly regulars, the way English actually arrives. Six runs. Then watch the original fourteen — learned by rote before the flood — for a later collapse.
Arshavir Blackwell, PhD
as the vocabulary grows from 20 verbs to 500, the fourteen seed irregular verbs are mastered almost immediately and stay at 100% correct throughout. No onset, no critical mass, no U-shaped dip.
Arshavir Blackwell, PhD
Nothing happened. The model learned those fourteen almost immediately — by step 150, while it still knew only twenty words — and held them at 100% correct through the entire expansion, in all six runs. Errors hit zero once the vocabulary reached forty and never came back. Four hundred and eighty regular verbs poured in, and it simply learned them alongside the irregulars it already had. It was never tempted.
Arshavir Blackwell, PhD
Same recipe as Plunkett and Marchman. Opposite result.
Arshavir Blackwell, PhD
The obvious explanation is size. Their network was small, so five hundred verbs had to compete for the same narrow storage, and learning one thing would damage another. Mine has eleven million parameters and room for everything.
Arshavir Blackwell, PhD
I liked that story. It explains the discrepancy without anyone being wrong, and it points somewhere interesting: the U would be a property of a system under pressure rather than of gradual learning — and a small child is nothing if not a learner under pressure.
Arshavir Blackwell, PhD
So I tested it, shrinking the model through three orders of magnitude, down into the size range of the machines that started this argument.
Arshavir Blackwell, PhD
model size versus the child-like dip. At 10.7 million, 1.8 million, 240 thousand and 34 thousand parameters the model learns the irregular verbs, and the rise in overregularization after mastery is 0.00. At 10 thousand parameters it never learns them.
Arshavir Blackwell, PhD
It didn’t work. At every size the model could learn the task at all, the rise after mastery was exactly zero. A network of thirty-four thousand parameters — a rounding error next to the one I started with — still held all fourteen irregulars at 100% while five hundred verbs piled in around them. Shrink it once more and it never learns them in the first place, so there’s nothing to over-apply. There’s no window in between: the model goes straight from “room for everything” to “can’t do the task.”
Arshavir Blackwell, PhD
So the explanation I liked is wrong. I’m leaving it in rather than quietly deleting it, because deleting it is the move Part 3 warned about. Which leaves the discrepancy honestly unexplained — not the schedule, not capacity. What’s left is what I didn’t vary: a transformer is not a multi-layer perceptron, my expansion curve is not theirs, and their model read hand-built sound features where mine reads letters. I don’t know which carries it, and I’d rather say so than pick the flattering one.
Arshavir Blackwell, PhD
Two ways to fail, and a measure that confused them
Arshavir Blackwell, PhD
Part 4 originally reported that on made-up verbs the sound model prefers the correctly-formed past tense to the bare, unchanged stem only about 40% of the time — 45% on the 600-verb model, 41% on the 72-verb one. That number hid a distinction, and it also turned out to be wrong. This section is where both came to light, and where the figure Part 4 now carries was measured.
Arshavir Blackwell, PhD
There are two ways to fail on a novel verb. The model can decline to inflect, preferring the bare stem to any ending. Or it can inflect and choose the wrong one. Scoring the correct form against the bare stem counts both as the same failure. So I scored all four candidates for every novel stem — the bare form and each of the three endings — and sorted the outcome into three exhaustive buckets — six seeds, 24 novel stems:
Arshavir Blackwell, PhD
the 600-verb model as first scored. Novel stems: declines 51%, right ending 45%, wrong ending 4%. Familiar verbs: declines 7%, right ending 93%, wrong ending 0%.
Arshavir Blackwell, PhD
Refusal outnumbered mis-inflection twelve to one. I believed that for a while, and the number was wrong. Not the distinction — that turns out to matter more than I thought — but the measurement. How it was wrong is worth the detour, because it is the kind of thing that will be wrong in other people’s probes too.
Arshavir Blackwell, PhD
To compare walk against walked you need a score for each. I used average log-probability per phoneme, which is the standard way to stop a longer string from being penalized for its length. But I scored the candidates without the word boundary that follows them. That leaves the bare stem a strict prefix of every inflected form it competes against — W AO K inside W AO K T — and under length normalization the longer form wins only if its extra phoneme beats the average of the phonemes before it. The stem is sitting right there in the prompt, so the model copies it at about 97% per phoneme. Any ending it was less than 97% sure of lost automatically.
Arshavir Blackwell, PhD
So the bucket labeled “declines to inflect” was holding two different things: a model that wanted to leave the word alone, and a model that wanted to inflect but was not certain which ending to use. The repair is to score each candidate as a complete word — its phonemes plus the boundary — and compare those. Same checkpoints, same stems:
Arshavir Blackwell, PhD
the 600-verb model scored as whole words. Novel stems: declines 22%, right ending 66%, wrong ending 13%. Familiar verbs: declines 0%, right ending 100%, wrong ending 0%.
Arshavir Blackwell, PhD
Most of the twelve-to-one was an artifact. What survives is a real preference for omission over substitution, about two to one, in a model considerably more willing to inflect than I first reported. Part 4’s other figure barely moves: its 88% answers the other question — which ending, once the bare form is set aside — and on this scoring that is about 84%.
Arshavir Blackwell, PhD
The caveat I would have written here has turned into the finding. The bare stem is not a neutral competitor: it is the string already sitting in the prompt, and copying is the cheapest operation a next-token predictor has. I knew that, wrote it down as a limitation, and carried on. It was doing most of the work.
Arshavir Blackwell, PhD
The artifact also hid something about scale, and this is the part I had most wrong. Scored the old way, the 72-verb model and the 600-verb model refuse on identical proportions of novel stems — 51% each. Eight times the data, no movement whatsoever. That is what led me to conclude that more data teaches which ending belongs and never teaches that an ending is obligatory. Scored as whole words the same two models come out at 55% and 22%. Eight times the data more than halves the refusal.
Arshavir Blackwell, PhD
The old measure could not see this because it was saturated in the large model and not in the small one. Ask each model what it wants to put after the bare stem: the small one gives the end of the word 47% of the probability, the large one 18%. The small model really does want to stop. The large one mostly does not, and the first scoring could not tell those apart.
Arshavir Blackwell, PhD
The distinction itself is not cosmetic. Having a rule takes two things, and they come apart: knowing that an ending is required, and knowing which ending. Choosing wrongly is a failure of the second. Declining to inflect is a failure of the first. Either can fail while the other holds. One system might know a slot has to be filled and conflate what fills it; another might rank the right ending first and never commit to using it. And which failure you actually see is not decided by the learner alone. It also depends on whether the language leaves refusal available — English does, because taking the ending off a verb still leaves a word behind.
Arshavir Blackwell, PhD
The two accounts were supposed to predict different mistakes. A rule learned imperfectly should fire and choose badly — substitution. A system generalizing by resemblance should fail by producing nothing at all: a stem with no close neighbors gets no pull, and the unchanged form wins by default. Two-to-one toward omission still points at the second.
Arshavir Blackwell, PhD
But that reading assumes the choice between the two failures belongs to the model, and it doesn’t. English hands it the bare option twice over. The stem is a word by itself, and for thirteen verbs in this training set it is also a legal past — cut, hit, let, put, set, shut, cost, spread and the rest, every one of them ending in a /t/ or a /d/. So if the refusals are about what English offers rather than about the learner, they should not be spread across novel stems at all. They should collect on stems ending in /t/ or /d/.
Arshavir Blackwell, PhD
The per-class breakdown is where the corrected measure earns its keep. Sorted by the ending each novel stem required, 600-verb model, scored as whole words:
Arshavir Blackwell, PhD
novel stems by the ending they require, 600-verb model, scored as whole words. Walked-type /t/: declines 4%, right ending 71%, wrong ending 25%. Played-type /d/: declines 13%, right ending 75%, wrong ending 13%. Wanted-type /əd/: declines 48%, right ending 52%, wrong ending 0%.
Arshavir Blackwell, PhD
Refusal is not a general property of novel stems. It is one class. Outside the wanted-type stems the model inflects almost everything and sometimes picks wrongly. Inside that class it leaves the word alone half the time, and when it does inflect one it is never wrong. It knows what ending those stems take. It will not commit to using it.
Arshavir Blackwell, PhD
Which is the dissociation I thought I had found across the board, now located — on the wanted-type stems, which are the ones ending in /t/ or /d/. On the 72-verb model, where there is far less evidence for anything else, that class — the one Part 4 reported as perfect when only the three endings were compared — refuses on 98% of stems.
Arshavir Blackwell, PhD
The location has two possible explanations, and they are worth pulling apart. One is that these stems sound like the thirteen verbs whose past tense is the bare form, and the model copies them. The other is forty years older. Bybee and Slobin (1982) found children and adults leaving t/d-final verbs unchanged, and argued that such a stem already ends in the sound the English past tense ends in, so it looks finished before anything is added. That account needs no zero-change verbs at all.
Arshavir Blackwell, PhD
So I took them out. All thirteen, replaced by thirteen t/d-final irregulars whose past does change — bleed/bled, light/lit, bind/bound — so the model sees the same number of t/d-final irregulars, and the only thing gone is the bare pasts. Refusal on t/d-final stems went from 48% to 38%, a drop six seeds cannot tell apart from noise, and it stays about three times the rate on other stems. (Removing them with no replacement lowers it further, to 31%, but that also takes thirteen t/d-final verbs out of the language altogether.) The model is mostly not copying hit. Those stems already sound past.
Arshavir Blackwell, PhD
There is a larger test in all this, and I ran it before writing it up too. Omission and substitution are not interchangeable ways of being wrong; which one a system produces depends on what the language makes cheap. English tolerates omission because the bare stem is itself a legal word — drop the ending and you still have walk. Italian does not: a bare root like parl- is not a word of the language in any context, not a noun, not a command, not a dictionary entry.
Arshavir Blackwell, PhD
So I built an Italian corpus the same way as the English one and trained the same architecture on it — same size, same six seeds, same quantity of text.
Arshavir Blackwell, PhD
novel stems scored as whole words. English, 600 verbs: declines 22%, right ending 66%, wrong ending 13%. Italian, 353 verbs: declines 3%, right ending 56%, wrong ending 42%.
Arshavir Blackwell, PhD
The balance inverts. And the escape route itself can be measured: ask each model how much probability it puts on the word simply ending after the bare root, and English says 18% while Italian says zero to three decimal places. Not merely unused — absent.
Arshavir Blackwell, PhD
That is the logic of the cross-linguistic aphasia studies (Wulfeck, Bates & Capasso, 1991), applied to a model rather than a speaker, and it comes out the way that logic predicts. The failure is not a property of the learner alone. It is what the learner does with the language it was given.
Arshavir Blackwell, PhD
Two things went wrong here, and one of them twice. The replication failed, and then the explanation I liked for the failure failed too. And the measurement I was using to describe the model turned out to be describing my scoring. What is left is smaller and steadier than what I started with: no U from gradual growth at any size, a model that commits about seven times in ten, refusal concentrated almost entirely on stems that already sound like past tenses, and a failure that changes shape when you change the language. I would rather have the steadier version. I'm Arshavir Blackwell and this has been Inside the Black Box.