
My Advisors Argued This for Thirty Years. Now You Can Check
Chapter 1
Imported Transcript
Arshavir Blackwell, PhD
My Advisors Argued This for Thirty Years. Now You Can Check
Arshavir Blackwell, PhD
At a conference once, a researcher said of a certain group of authors: "The only thing you people seem to have in common is you hate Chomsky."
Arshavir Blackwell, PhD
The book those authors wrote was called Rethinking Innateness. Six of them, in 1996. Thirty years ago, but still relevant! Two of them were my advisors, and I was working with them while they wrote it, its questions and themes were things I heard and argued about every day. I'm Arshavir Blackwell and this is Inside the Black Box.
Arshavir Blackwell, PhD
The fight was over whether language is innate or learned.
Arshavir Blackwell, PhD
Elman et al.'s argument was that everyone had been asking the wrong question. People kept fighting about whether language is "born in" or learned. They said: stop. In a real sense of course it's innate: we're born with all the equipment necessary for language learning to take place, and no new equipment gets installed later.
Arshavir Blackwell, PhD
This is what politicians call reframing the argument.
Arshavir Blackwell, PhD
The real question is what exactly gets built in, and there are three very different answers.
Arshavir Blackwell, PhD
You could build in the knowledge itself, with a baby arriving already knowing something about how sentences work, expecting sentences to have a certain shape and form. Chomsky's picture was of a genetically preprogrammed language organ, one that grows on its own schedule the way any other organ does.
Arshavir Blackwell, PhD
You could build in the wiring, with no knowledge, just a certain shape of brain, certain parts connected to certain other parts, some paths narrow and some wide.
Arshavir Blackwell, PhD
Or you could build in the timing, with nothing about content or shape, only a schedule. This matures now, that one later, this window closes at six.
Arshavir Blackwell, PhD
Elman et al.'s bet was on the second and third. Build the right learner, start it in the right order, put it in a world with structure, and the knowledge shows up on its own. You almost never have to install it.
Arshavir Blackwell, PhD
(There's a live argument now that language models need a "world model" bolted on, some explicit representation of how things work, beyond text. Rethinking Innateness would say the structure is already out there in the environment, and a learner built the right way picks it up unaided. Though it cuts both ways: if text is a thin environment, Elman et al.'s own argument predicts thin learning.)
Arshavir Blackwell, PhD
It was a good argument, but nobody could check it.
Arshavir Blackwell, PhD
You cannot open up a child mid-sentence and look inside.
Arshavir Blackwell, PhD
What we have now isn't a child. But it is a learner you can open.
Arshavir Blackwell, PhD
It is a lookup, not a spotlight
Arshavir Blackwell, PhD
First, the thing itself. A large language model is trained to do one job: given a stretch of text, guess the next word. It is shown text with the next word hidden, guesses, gets corrected, and adjusts. Billions of times over. That is the whole of the training.
Arshavir Blackwell, PhD
What does the guessing is a stack of identical layers, 32 of them in Llama-3.1-8B, a model I work with, each doing two things in turn: a round of lookups, then a round of ordinary number-crunching. The lookups are divided among 32 smaller units called heads, running side by side. All of it is one very large pile of numbers, called weights, that started out random and got nudged, correction by correction, until the guesses came out right.
Arshavir Blackwell, PhD
That first half, the lookups, is what everyone points to. It is called attention, and the name is unhelpful. It sounds like the model is deciding what matters. It isn't. It's doing something much more ordinary. It's looking things up.
Arshavir Blackwell, PhD
Think of a Python dictionary, or an address book. You hand it a name, it finds that name, it gives you back the phone number. A key goes in, a value comes out.
Arshavir Blackwell, PhD
Attention is that, but with two changes.
Arshavir Blackwell, PhD
The first: it never finds one match. It scores every entry in the book, then hands back a blend of all of them, weighted by how good each match was. A strong match contributes most of the answer, a poor one contributes a trace.
Arshavir Blackwell, PhD
The second: the address book is rebuilt from scratch for every sentence. Each word the model has read becomes one entry. Read a different sentence, get a different book. The address book is the sentence.
Arshavir Blackwell, PhD
So when the model reaches the verb in "the key to the cabinets," it asks a question, roughly, where is my subject, and every earlier word gets scored on how well it answers. The answers come back mixed together, in proportion.
Arshavir Blackwell, PhD
Three things get computed from each word: what it's looking for, how it advertises itself, and what it hands over if chosen. Three different readings of the same word.
Arshavir Blackwell, PhD
And here is the only thing you need to carry into the next section: none of that is about language in the Chomskian sense. The design says three readings per word, 32 lookups running at once, each one squeezed through a narrow channel, all of them competing. It says nothing about nouns, or verbs, or agreement. No part is set aside for grammar.
Arshavir Blackwell, PhD
Structure nobody installed
Arshavir Blackwell, PhD
And yet structure shows up anyway.
Arshavir Blackwell, PhD
Some heads end up tracking agreement, quietly keeping the subject available so the verb can match it. Nobody assigned them that job.
Arshavir Blackwell, PhD
Others end up doing something called induction, which is worth a moment because it's the best-documented case. Suppose a model has seen the sequence "Dr. Lieberman" earlier in a page, and now hits "Dr." again. It goes back to the earlier "Dr.," looks at whatever came next, "Lieberman", and copies it forward as the prediction. Find where this happened before; see what followed; do that again. It's how models pick up a pattern from a few examples in the prompt rather than from training.
Arshavir Blackwell, PhD
Doing that takes two heads working together: an earlier one that tags each word with the word that came before it, and a later one that uses those tags to find the match and copy what followed. Neither head does the job alone.
Arshavir Blackwell, PhD
Nobody designs these either. And they don't fade in gradually, they appear abruptly, at a datable moment in training, with a visible bump in the loss curve where the model suddenly gets better at learning from context (Olsson et al., 2022). One of the few places in this field where you can point at a graph and say: there, that's when it developed.
Arshavir Blackwell, PhD
This is exactly what Elman et al. said would happen, and here it happens with the objection they could never answer taken off the table.
Arshavir Blackwell, PhD
A child comes with a genome, so a nativist can always say the structure was in there from the start. A transformer has no genome. Nothing about grammar is written anywhere in its design. So when agreement tracking and induction turn up anyway, there are only two places they can have come from: the shape of the machine, and the text it read.
Arshavir Blackwell, PhD
The neat part is that a transformer has two of Elman et al.'s three levels, the wiring and the timing, and they're cleanly separable, which is exactly what you can't do in a child.
Arshavir Blackwell, PhD
Architectural. Elman et al.'s definition: the structuring of the information-processing system that has to acquire the representations. Their neural-net examples are number of layers, density of units within layers, presence of recurrent connections. A head's 128 dimensions, the model's 4,096 split 32 ways, is precisely this. It's a bottleneck, and it's why heads specialize, a head can't do everything, so it does something. Specialization is downstream of narrowness.
Arshavir Blackwell, PhD
Chronotopic. Elman et al.'s definition: constraints on the timing of developmental events. Their examples include "incremental presentation of data" and "adaptive learning rates." This is where Elman's "starting small" (1993) belongs, his recurrent net learned long-distance agreement only when memory started limited and grew; full capacity from the start failed. Limitation was productive, and it was a limitation in time, not in structure.
Arshavir Blackwell, PhD
A transformer has both, and they're independent knobs: the bottleneck is fixed by architecture, the schedule by the training curriculum. Elman et al.'s third level, representational, is simply absent. Nothing is installed.
Arshavir Blackwell, PhD
The theories, in the machinery
Arshavir Blackwell, PhD
Bates and Elman had different theories. Attention has something answering to each of them, and one thing answering to neither.
Arshavir Blackwell, PhD
Bates: cues competing for a fixed budget
Arshavir Blackwell, PhD
Bates's account was that multiple probabilistic cues compete to determine an interpretation. Each cue has a validity, which is how often it's available and how often it's right, a fact about the language. Speakers learn a cue strength to match. When cues converge, processing is fast and confident. When cues conflict, processing is slower and error-prone.
Arshavir Blackwell, PhD
Attention has that structure, and not merely by analogy. Every earlier word is scored, and the softmax forces the scores into a distribution summing to one, so the candidates are in literal zero-sum competition for a fixed budget, exactly as cues are in Bates's model. What the query and key lenses learn is the functional equivalent of cue strength: how much a given kind of match counts, tuned by training toward whatever predicts well. Validity, estimated from data rather than stipulated.
Arshavir Blackwell, PhD
Convergence and conflict both fall out. Back to the key to the cabinets…: the verb queries for its subject, key matches on head-noun, cabinets matches on noun-and-recent. Weight smears across both, the value comes back blended, and the number feature arrives contaminated, which is how the wrong verb gets produced. Agreement attraction is produced by the mechanism rather than despite it.
Arshavir Blackwell, PhD
Bates's cross-linguistic prediction comes along too: train on a language where case marking is the valid cue, and the lenses should weight it accordingly. I am not aware of anyone testing it.
Arshavir Blackwell, PhD
Elman: the fix for his own bottleneck
Arshavir Blackwell, PhD
Elman's own network, the simple recurrent net, carried context as one recurrent state, everything compressed into a single vector. That is why long-distance dependencies were hard, and why starting small helped: with limited memory the network had to learn simple regularities first, and those scaffolded the rest.
Arshavir Blackwell, PhD
LLM attention removes the constraint. Every prior state stays addressable; nothing is compressed into one carrier. That relocates the failure. It is no longer capacity but retrieval precision. The model can reach anything, but it has to win a competition to get it. Which is where it breaks, when it breaks.
Arshavir Blackwell, PhD
Words as operators, not entries
Arshavir Blackwell, PhD
Elman argued against a mental lexicon. He believed that words aren't stored entries with contents, but rather operators that push the ongoing state around. Meaning is what a word does, not what it holds.
Arshavir Blackwell, PhD
That's the architecture. A token doesn't sit there being itself. It produces a key, which is how it makes itself findable, and a value, which is what it does to whoever retrieves it: a vector added into the retriever's stream, altering it. And the same word yields different keys and values in different heads and at different layers, so there is no stable entry anywhere. No lexicon, just context-dependent effects.
Arshavir Blackwell, PhD
The one thing neither of them predicted
Arshavir Blackwell, PhD
Softmax has no way to abstain. The weights must sum to one, so a query matching nothing still spends its whole budget and gets back the average of everything present. A listener in the Competition Model can be uncertain and withhold. Nothing in the LLM design allows for that.
Arshavir Blackwell, PhD
Models find a way around it, and the workaround is telling. They learn to dump the unwanted weight on one particular token, usually the first in the text, whose value carries almost nothing. It's a parking space, invented rather than installed, that lets a query decline to answer (Xiao et al., 2023). Even abstention has to be improvised out of machinery built for something else.
Arshavir Blackwell, PhD
Degrade the machine and it does not fall silent. It cannot. The weights flatten, every entry contributes a little, and what comes back is the whole context blurred together, the most ordinary thing available. Which may be the mechanism under a sentence from Part 4 of the past-tense series: it slumps toward its most common habit and mumbles ed.
Arshavir Blackwell, PhD
Looking for the part
Arshavir Blackwell, PhD
Another of the six, Annette Karmiloff-Smith, had published a book of her own four years earlier. It was called Beyond Modularity, and it asked whether the mind comes in parts.
Arshavir Blackwell, PhD
A module is a piece of the mind that does one job and only that job, a subroutine, in the programming sense. Written in advance, sealed off from the rest of the code, called when its job comes up. Language in one, faces in another, each shipped with the machine.
Arshavir Blackwell, PhD
Jerry Fodor made the case for it in 1983, in a book whose cover carried a phrenologist's head, a skull mapped into labelled compartments, each one an organ of some faculty. That is the picture at the top of this piece, with the labels changed.
Arshavir Blackwell, PhD
Karmiloff-Smith agreed that grown-up minds end up specialized. What she denied was that they start that way, and she denied that the specialization ever gets as clean as Fodor's picture. A child begins with broad leanings, a pull toward speech, toward faces, toward things that move, and the same circuits, doing the same job over and over, slowly narrow into specialists. Her word for that was modularization: not a module, but the process of becoming one. A matter of degree, running throughout development, and never guaranteed to finish.
Arshavir Blackwell, PhD
That gives you something to test, if you have a system you can take apart. Break it, and watch what breaks with it. If a skill lives in a part of its own, damage to that part should take the skill out and leave everything else standing. If the skill is spread across the whole machine, damage anywhere should cost you a little of it, and cost you everything else just as much.
Arshavir Blackwell, PhD
So I tried it, on the model from my own past-tense work, a newer run, not yet written up. The model has two jobs. One is remembering: the past tense of go is went, and it has seen that a million times. The other is working from scratch: the past tense of wug, a word nobody has ever used, which it has to build.
Arshavir Blackwell, PhD
Damage the model and the made-up words fail first, badly, while the memorized ones hold on. That happens two different ways, switching off parts of the machinery at random, or just adding static to its internal signals. Two unrelated kinds of harm, same casualty.
Arshavir Blackwell, PhD
So I went looking for the part. The model has 32 layers; I switched off each one in turn. The damage concentrated in the first two, right where the model is still assembling words out of fragments. Promising.
Arshavir Blackwell, PhD
Then I went inside, to the heads that share out the lookups, 32 to a layer. I switched off each one singly, in the layer where the damage sat in the lookups, and in two others besides. 96 heads.
Arshavir Blackwell, PhD
Nothing. Not one of them mattered on its own.
Arshavir Blackwell, PhD
Which does not close the question. Induction, remember, takes two heads working together, so a part could still be hiding in a pair or a group, and pairs I have not tested. What is ruled out is any head doing this job by itself.
Arshavir Blackwell, PhD
So the two scales disagree. Zoom out to whole layers and the weakness is somewhere in particular: those first two, and almost nowhere else. Zoom in to the individual heads and it is nowhere in particular: no one of them is carrying the job.
Arshavir Blackwell, PhD
Picture a building where you can prove the fault is on one particular floor, that is where everything goes wrong, and then you test every room on that floor and each one works fine by itself. The floor matters. No room does.
Arshavir Blackwell, PhD
So there is an address, so to speak, but nothing at that address. A weak point with a location and no part in it.
Arshavir Blackwell, PhD
One more check, because the obvious objection is that made-up words are simply harder. Damage the front of the model where the deficit concentrates, then ask it only to repeat wug back: it manages seven times in eight. Ask it for the past tense of wug: not once, in any trial. Same word, same damage. What breaks is the inflecting, not the strangeness.
Arshavir Blackwell, PhD
An address with nothing in it is a strange result if you think there are dedicated parts, and an equally strange one if you think everything is smeared evenly. What it looks like is a middle state: a job that has acquired a location without acquiring a boundary.
Arshavir Blackwell, PhD
Karmiloff-Smith's account is the only one of the three with room for a middle state. Modularization is a process, and a process can be caught halfway. Whether more training would harden this into something you could point at, I do not know, that is the experiment I have not run. But neither of the old positions predicts anything like it, and hers has a name for it.
Arshavir Blackwell, PhD
The loop
Arshavir Blackwell, PhD
Around 1990 Jeff Elman built a small network and trained it by having it guess the next word in a sentence. He was not building a product. He was asking whether a learner could pick up grammar from nothing but the sentences it hears, whether structure could come out of a system that had none put in.
Arshavir Blackwell, PhD
That is the same job every model you have used was trained on. The thing on your phone traces back to that question about children and grammar.
Arshavir Blackwell, PhD
And the models we have now are the instrument the argument always needed. Not because they are human, they are not, but because they are learners you can take apart. That was the one thing missing. Everyone had data, nobody had access.
Arshavir Blackwell, PhD
This is not a one-to-one correspondence: a model has no body, nothing matures, and no one talks back to it. A brain is a patchwork of regions with different wiring; a transformer is the same block copied 32 times. And the timing level, the one Elman et al. cared most about, maps worst of all. A learning-rate schedule is a thin substitute for acquiring a language.
Arshavir Blackwell, PhD
But that difference is not a weakness in the argument. It is what makes the argument clean.
Arshavir Blackwell, PhD
Part of the nativist case was always an in-principle claim: you cannot get grammatical behavior out of a system that has no grammar in it. Statistics over experience would not be enough. Somewhere there had to be rules, and symbols for the rules to work on.
Arshavir Blackwell, PhD
Open a transformer and there is nothing of that kind inside. No rule, no symbol, no grammatical category, no place where subject and verb are declared to agree. Only weights and dot products. Out of that comes agreement across a clause, structure held over distance, and a past tense for a word invented five seconds ago, not reliably, but far more often than chance.
Arshavir Blackwell, PhD
That is an existence proof. Not that people work this way, as a model of a brain a transformer is poor, and I would not defend it as one. Only that the in-principle claim was wrong. A fully distributed system, with no rules and no symbols anywhere in it, can do this. Whether humans do it some other way is now a question you have to answer with evidence.
Arshavir Blackwell, PhD
So this is a convergence, not a vindication. Nobody built attention to settle a psycholinguistics argument, and it would be too convenient if a machine designed for something else happened to prove my advisors right.
Arshavir Blackwell, PhD
But it keeps coming out the way Elman et al. said it would. Structure appears that nobody installed. The parts you would expect are not there. The errors are, to a certain extent, the human errors. And when you strain it, the productive machinery is what gives way first, which is what Liz Bates spent a career arguing about people.
Arshavir Blackwell, PhD
The question I would most like to answer is the one they cared most about and I have barely touched: not what a model ends up with, but when it gets it. Which comes first, remembering, or working from scratch? Does the second one ever stop being the fragile one? Some models ship with checkpoints saved all the way through training, so the whole development is sitting there on disk, waiting.
Arshavir Blackwell, PhD
Neither of them saw this. Jeff and Liz both left us years ago. Those checkpoints are the experiment they could not run. I'm Arshavir Blackwell and this has been Inside the Black Box.