Inside the Black Box: Cracking AI and Deep Learning
All Episodes
Not at This Address

Not at This Address

0:00|0:00

This episode challenges the old idea of a single language box in the brain, using aphasia studies and hospital control groups to show how language processing is more distributed than once thought. It also explores how different languages rely on word order or verb marking in different ways, and why those differences matter when the brain is under stress.


Chapter 1

Imported Transcript

Arshavir Blackwell, PhD

Not at This Address

Arshavir Blackwell, PhD

A Peek at Chomsky’s “Language Organ”

Arshavir Blackwell, PhD

I'm Arshavir Blackwell and this is Inside the Black Box. For a long time, the standard explanation of language in the brain was a box.

Arshavir Blackwell, PhD

Somewhere in there, the story went, was a part that did grammar. Not a metaphor, a place. If that part was damaged, grammar broke. That was the explanation for aphasia, which is what happens when someone has a stroke or a head injury and afterwards cannot produce or understand well formed sentences. The wiring to the grammar box was cut. That was the whole account.

Arshavir Blackwell, PhD

My advisor, Elizabeth Bates, did not believe it. She thought the brain was a much more distributed thing, and that the evidence people were reading as a broken box could be read several other ways. What convinced me she was right was a control group.

Arshavir Blackwell, PhD

The patients who should have been fine

Arshavir Blackwell, PhD

Her group was running a three language study, English, German and Italian, and it was the Italian hospitals that supplied that control group. The task was comprehension: hear a sentence, then act it out with small objects. Whichever object the patient picked up counted as the one doing the action. In some sentences the only clue to who did it was whether the verb was singular or plural, so the task measured whether a patient could still use that marking. The aphasic patients did worse than healthy speakers, and not across the board: word order held up; the marking on the verb did not. No surprise there.

Arshavir Blackwell, PhD

A good control is someone who is also in a hospital, also in pain, also on medication, also frightened, but just not in the brain. So they tested patients from the orthopedic ward, in for hip fractures and the like. They also tested people with illnesses of the nervous system that leave the cortex alone, like polio.

Arshavir Blackwell, PhD

Both groups were significantly worse than healthy young speakers at using the agreement marking. On average they fell between the aphasics and the young speakers, but neither group was significantly different from the aphasics. Receptive agrammatism, from a broken hip. One caveat: the hospital patients were 32 to 60 and the healthy speakers were college students, so some of that gap may be age rather than pain and medication.

Arshavir Blackwell, PhD

As Liz liked to say: language processing does not go on in the hip.

Arshavir Blackwell, PhD

So what was going on?

Arshavir Blackwell, PhD

The second crack

Arshavir Blackwell, PhD

The other finding came from setting languages side by side.

Arshavir Blackwell, PhD

English barely marks anything on its words to show who did what to whom. We rely on word order. The thing before the verb is the subject, the thing after it is the object, and if you scramble that, the sentence falls apart. English speakers lean on word order so hard that they follow it even when it makes no sense: given “The pencil kicked the cow”, most of them say the pencil did the kicking, because it came first. Word order wins over common sense.

Arshavir Blackwell, PhD

Other languages do it differently. Hungarian marks who did what on the word itself. Macska means cat; make the cat the one being acted on and it becomes macskat. Turkish does something similar.

Arshavir Blackwell, PhD

Aphasic speakers of those languages used that ending far less than healthy speakers did. But they still went by the ending rather than by word order, so the pattern of their language survived, only noisier.

Arshavir Blackwell, PhD

Put it another way: aphasics resemble the healthy speakers of their own language more than they resemble aphasics in other languages.

Arshavir Blackwell, PhD

Two things are true here at once, and it matters to keep them apart. Within any one language, the marking on the word is the weak link: it gives way before word order does, and that was true of English, German and Italian aphasics alike in Bates’s three language study. Across languages, how well it holds up tracks how heavily the language leans on it. The weak link is in the same place everywhere. How weak it is depends on the language.

Arshavir Blackwell, PhD

That is hard to square with a broken box. If the box is smashed, why does it matter which language you happen to speak? Broken grammar should be broken grammar in any language.

Arshavir Blackwell, PhD

The experiment

Arshavir Blackwell, PhD

When I was a graduate student, Liz and I proposed something else: at least part of what looks like a specific grammar failure might be a general one. Being in a hospital is stressful. Having had a stroke is stressful. Maybe some of what was being measured was a system under load, not a system with a missing component.

Arshavir Blackwell, PhD

Almost everyone I described this to made the same prediction. If you load someone down cognitively, they’ll just get worse. Uniformly worse, across the board. There would be no difference in judging different types of grammatical errors.

Arshavir Blackwell, PhD

That is not what aphasia looks like. English speaking aphasics spot word order errors, like “the boy talking was”, more easily than agreement errors, like “the boy were talking”. We thought the uniform prediction was wrong, and... here’s the spoiler!... it was.

Arshavir Blackwell, PhD

The experiment used college students, with no brain damage, as far as we knew, and nothing unusual. Before each sentence, we flashed a string of digits on a screen, as few as two, as many as six, that they had to hold in memory while they judged whether the sentence was grammatical. That was the load.

Arshavir Blackwell, PhD

Everything got harder under the heaviest load, but not evenly. Agreement errors went first: they got harder to catch with only two digits to hold. Word order, the thing English leans on hardest, held out longer.

Arshavir Blackwell, PhD

That is the same pattern the aphasic patients show, reproduced in undergraduates whose only affliction was holding a string of digits in memory.

Arshavir Blackwell, PhD

A selective deficit, in brains with nothing wrong with them.

Arshavir Blackwell, PhD

That is the point. You do not need a broken box to get a specific looking failure. You can get one from a perfectly intact system that is simply working too hard.

Arshavir Blackwell, PhD

The same question, asked of a machine

Arshavir Blackwell, PhD

Thirty years later I asked the same kind of question of language models, in a four part series here and in some experiments since. In summary:

Arshavir Blackwell, PhD

Models learn the past tense in the opposite order from children. They add the ending e d to everything first and learn exceptions like went later. Children do the reverse, and models never show the dip where a child starts saying goed. That held in large models and in small ones reading letters or sounds.

Arshavir Blackwell, PhD

You can only force that dip by rigging the training. Growing the vocabulary gradually, which gave Plunkett and Marchman a dip, didn’t give me one. Neither did shrinking the model.

Arshavir Blackwell, PhD

There is no rule inside the model. No single part handles the regular verbs or the exceptions. What it has is a habit spread across the whole network, working by resemblance to verbs it already knows.

Arshavir Blackwell, PhD

More training data helps a lot with made up verbs. The failures that remain are mostly on verbs ending in t or d, which already sound past. Children make the same mistake.

Arshavir Blackwell, PhD

The language decides the kind of mistake. The English model sometimes leaves a new verb unchanged. The Italian model almost never does; it picks the wrong ending instead.

Arshavir Blackwell, PhD

Damage and overload break new verbs first. Adding noise to a model breaks past tenses for made up verbs before memorized ones, and so does switching parts of it off. The noise removes nothing, which makes it the model’s version of the digit load. The damage clusters in the model’s first layers, but no single small part matters on its own.

Arshavir Blackwell, PhD

A transformer is not a brain. The scale is different by orders of magnitude, and so is nearly everything else. A model’s units are not neurons; neurons are vastly more complex. Still, put the two sets of results side by side and three things line up.

Arshavir Blackwell, PhD

What gives way first is the weak link, not a special grammar part. In patients, and in students holding digits, it was the marking on words, which gave way before word order. In the model, it was the ending on a verb it had never practiced.

Arshavir Blackwell, PhD

The language shapes the failure, in patients and in models alike. The English model can leave a new verb bare because English allows it: walk is a word on its own. The Italian model reaches for a wrong ending instead, because the bare root parl is not a word on its own.

Arshavir Blackwell, PhD

And neither system needs a missing part to fail selectively. Overload did it to healthy students; noise did it to a model.

Arshavir Blackwell, PhD

What this is actually about

Arshavir Blackwell, PhD

A brain and a transformer both store what they know spread across many small parts rather than in one dedicated part. So perhaps what the patients, the students and the models have in common was never a fact about brains alone. Perhaps it is a fact about distributed systems.

Arshavir Blackwell, PhD

If so, the aphasia work from thirty years ago was never only about aphasia.

Arshavir Blackwell, PhD

There is a practical edge to this, too. It is the reason interpretability research exists and the reason it is slow. If you cannot find the part that does the thing, you cannot audit it, fix it, or promise it will hold.

Arshavir Blackwell, PhD

And it is why I keep saying the same thing whenever someone tells me a model clearly understands them: you are looking at the outside of the system. The outside is exactly what you cannot trust.

Arshavir Blackwell, PhD

The inside is billions of numbers that nobody set by hand, arranged in a way nobody designed, doing a job that no part of it is doing alone. I'm Arshavir Blackwell and this has been Inside the Black Box.