Inside the Black Box: Cracking AI and Deep Learning
All Episodes
My Advisors Argued This for Thirty Years. Now You Can Check

My Advisors Argued This for Thirty Years. Now You Can Check

0:00|0:00
Thirty years ago, the authors of Rethinking Innateness argued that grammar could come out of a learner with no grammar built in. Nobody could check it. Now there's an instrument. A reading of the Inside the Black Box essay on what Elman, Bates, and Karmiloff-Smith would make of transformer language models — attention as a lookup, structure nobody installed, and the part that isn't there. Read the essay: https://arshavirblackwell.substack.com/p/what-would-my-advisors-make-of-llms

Chapter 1

Imported Transcript

Arshavir Blackwell, PhD

My Advisors Argued This for Thirty Years. Now You Can Check

Arshavir Blackwell, PhD

At a conference once, a researcher said of a certain group of authors: "The only thing you people seem to have in common is you hate Chomsky."

Arshavir Blackwell, PhD

The book those authors wrote was called Rethinking Innateness. Six of them, in 1996. Thirty years ago, but still relevant! Two of them were my advisors, and I was working with them while they wrote it, its questions and themes were things I heard and argued about every day. I'm Arshavir Blackwell and this is Inside the Black Box.

Arshavir Blackwell, PhD

The fight was over whether language is innate or learned.

Arshavir Blackwell, PhD

Elman et al.'s argument was that everyone had been asking the wrong question. People kept fighting about whether language is "born in" or learned. They said: stop. In a real sense of course it's innate: we're born with all the equipment necessary for language learning to take place, and no new equipment gets installed later.

Arshavir Blackwell, PhD

This is what politicians call reframing the argument.

Arshavir Blackwell, PhD

The real question is what exactly gets built in, and there are three very different answers.

Arshavir Blackwell, PhD

You could build in the knowledge itself, with a baby arriving already knowing something about how sentences work, expecting sentences to have a certain shape and form. Chomsky's picture was of a genetically preprogrammed language organ, one that grows on its own schedule the way any other organ does.

Arshavir Blackwell, PhD

You could build in the wiring, with no knowledge, just a certain shape of brain, certain parts connected to certain other parts, some paths narrow and some wide.

Arshavir Blackwell, PhD

Or you could build in the timing, with nothing about content or shape, only a schedule. This matures now, that one later, this window closes at six.

Arshavir Blackwell, PhD

Elman et al.'s bet was on the second and third. Build the right learner, start it in the right order, put it in a world with structure, and the knowledge shows up on its own. You almost never have to install it.

Arshavir Blackwell, PhD

(There's a live argument now that language models need a "world model" bolted on, some explicit representation of how things work, beyond text. Rethinking Innateness would say the structure is already out there in the environment, and a learner built the right way picks it up unaided. Though it cuts both ways: if text is a thin environment, Elman et al.'s own argument predicts thin learning.)

Arshavir Blackwell, PhD

It was a good argument, but nobody could check it.

Arshavir Blackwell, PhD

You cannot open up a child mid-sentence and look inside.

Arshavir Blackwell, PhD

What we have now isn't a child. But it is a learner you can open.

Arshavir Blackwell, PhD

It is a lookup, not a spotlight

Arshavir Blackwell, PhD

First, the thing itself. A large language model is trained to do one job: given a stretch of text, guess the next word. It is shown text with the next word hidden, guesses, gets corrected, and adjusts. Billions of times over. That is the whole of the training.

Arshavir Blackwell, PhD

What does the guessing is a stack of identical layers, 32 of them in Llama-3.1-8B, a model I work with, each doing two things in turn: a round of lookups, then a round of ordinary number-crunching. The lookups are divided among 32 smaller units called heads, running side by side. All of it is one very large pile of numbers, called weights, that started out random and got nudged, correction by correction, until the guesses came out right.

Arshavir Blackwell, PhD

That first half, the lookups, is what everyone points to. It is called attention, and the name is unhelpful. It sounds like the model is deciding what matters. It isn't. It's doing something much more ordinary. It's looking things up.

Arshavir Blackwell, PhD

Think of a Python dictionary, or an address book. You hand it a name, it finds that name, it gives you back the phone number. A key goes in, a value comes out.

Arshavir Blackwell, PhD

Attention is that, but with two changes.

Arshavir Blackwell, PhD

The first: it never finds one match. It scores every entry in the book, then hands back a blend of all of them, weighted by how good each match was. A strong match contributes most of the answer, a poor one contributes a trace.

Arshavir Blackwell, PhD

The second: the address book is rebuilt from scratch for every sentence. Each word the model has read becomes one entry. Read a different sentence, get a different book. The address book is the sentence.

Arshavir Blackwell, PhD

So when the model reaches the verb in "the key to the cabinets," it asks a question, roughly, where is my subject, and every earlier word gets scored on how well it answers. The answers come back mixed together, in proportion.

Arshavir Blackwell, PhD

Three things get computed from each word: what it's looking for, how it advertises itself, and what it hands over if chosen. Three different readings of the same word.

Arshavir Blackwell, PhD

And here is the only thing you need to carry into the next section: none of that is about language in the Chomskian sense. The design says three readings per word, 32 lookups running at once, each one squeezed through a narrow channel, all of them competing. It says nothing about nouns, or verbs, or agreement. No part is set aside for grammar.

Arshavir Blackwell, PhD

Structure nobody installed

Arshavir Blackwell, PhD

And yet structure shows up anyway.

Arshavir Blackwell, PhD

Some heads end up tracking agreement, quietly keeping the subject available so the verb can match it. Nobody assigned them that job.

Arshavir Blackwell, PhD

Others end up doing something called induction, which is worth a moment because it's the best-documented case. Suppose a model has seen the sequence "Dr. Lieberman" earlier in a page, and now hits "Dr." again. It goes back to the earlier "Dr.," looks at whatever came next, "Lieberman", and copies it forward as the prediction. Find where this happened before; see what followed; do that again. It's how models pick up a pattern from a few examples in the prompt rather than from training.

Arshavir Blackwell, PhD

Doing that takes two heads working together: an earlier one that tags each word with the word that came before it, and a later one that uses those tags to find the match and copy what followed. Neither head does the job alone.

Arshavir Blackwell, PhD

Nobody designs these either. And they don't fade in gradually, they appear abruptly, at a datable moment in training, with a visible bump in the loss curve where the model suddenly gets better at learning from context (Olsson et al., 2022). One of the few places in this field where you can point at a graph and say: there, that's when it developed.

Arshavir Blackwell, PhD

This is exactly what Elman et al. said would happen, and here it happens with the objection they could never answer taken off the table.

Arshavir Blackwell, PhD

A child comes with a genome, so a nativist can always say the structure was in there from the start. A transformer has no genome. Nothing about grammar is written anywhere in its design. So when agreement tracking and induction turn up anyway, there are only two places they can have come from: the shape of the machine, and the text it read.

Arshavir Blackwell, PhD

The neat part is that a transformer has two of Elman et al.'s three levels, the wiring and the timing, and they're cleanly separable, which is exactly what you can't do in a child.

Arshavir Blackwell, PhD

Architectural. Elman et al.'s definition: the structuring of the information-processing system that has to acquire the representations. Their neural-net examples are number of layers, density of units within layers, presence of recurrent connections. A head's 128 dimensions, the model's 4,096 split 32 ways, is precisely this. It's a bottleneck, and it's why heads specialize, a head can't do everything, so it does something. Specialization is downstream of narrowness.

Arshavir Blackwell, PhD

Chronotopic. Elman et al.'s definition: constraints on the timing of developmental events. Their examples include "incremental presentation of data" and "adaptive learning rates." This is where Elman's "starting small" (1993) belongs, his recurrent net learned long-distance agreement only when memory started limited and grew; full capacity from the start failed. Limitation was productive, and it was a limitation in time, not in structure.

Arshavir Blackwell, PhD

A transformer has both, and they're independent knobs: the bottleneck is fixed by architecture, the schedule by the training curriculum. Elman et al.'s third level, representational, is simply absent. Nothing is installed.

Arshavir Blackwell, PhD

The theories, in the machinery

Arshavir Blackwell, PhD

Bates and Elman had different theories. Attention has something answering to each of them, and one thing answering to neither.

Arshavir Blackwell, PhD

Bates: cues competing for a fixed budget

Arshavir Blackwell, PhD

Bates's account was that multiple probabilistic cues compete to determine an interpretation. Each cue has a validity, which is how often it's available and how often it's right, a fact about the language. Speakers learn a cue strength to match. When cues converge, processing is fast and confident. When cues conflict, processing is slower and error-prone.

Arshavir Blackwell, PhD

Attention has that structure, and not merely by analogy. Every earlier word is scored, and the softmax forces the scores into a distribution summing to one, so the candidates are in literal zero-sum competition for a fixed budget, exactly as cues are in Bates's model. What the query and key lenses learn is the functional equivalent of cue strength: how much a given kind of match counts, tuned by training toward whatever predicts well. Validity, estimated from data rather than stipulated.

Arshavir Blackwell, PhD

Convergence and conflict both fall out. Back to the key to the cabinets…: the verb queries for its subject, key matches on head-noun, cabinets matches on noun-and-recent. Weight smears across both, the value comes back blended, and the number feature arrives contaminated, which is how the wrong verb gets produced. Agreement attraction is produced by the mechanism rather than despite it.

Arshavir Blackwell, PhD

Bates's cross-linguistic prediction comes along too: train on a language where case marking is the valid cue, and the lenses should weight it accordingly. I am not aware of anyone testing it.

Arshavir Blackwell, PhD

Elman: the fix for his own bottleneck

Arshavir Blackwell, PhD

Elman's own network, the simple recurrent net, carried context as one recurrent state, everything compressed into a single vector. That is why long-distance dependencies were hard, and why starting small helped: with limited memory the network had to learn simple regularities first, and those scaffolded the rest.

Arshavir Blackwell, PhD

LLM attention removes the constraint. Every prior state stays addressable; nothing is compressed into one carrier. That relocates the failure. It is no longer capacity but retrieval precision. The model can reach anything, but it has to win a competition to get it. Which is where it breaks, when it breaks.

Arshavir Blackwell, PhD

Words as operators, not entries

Arshavir Blackwell, PhD

Elman argued against a mental lexicon. He believed that words aren't stored entries with contents, but rather operators that push the ongoing state around. Meaning is what a word does, not what it holds.

Arshavir Blackwell, PhD

That's the architecture. A token doesn't sit there being itself. It produces a key, which is how it makes itself findable, and a value, which is what it does to whoever retrieves it: a vector added into the retriever's stream, altering it. And the same word yields different keys and values in different heads and at different layers, so there is no stable entry anywhere. No lexicon, just context-dependent effects.

Arshavir Blackwell, PhD

The one thing neither of them predicted

Arshavir Blackwell, PhD

Softmax has no way to abstain. The weights must sum to one, so a query matching nothing still spends its whole budget and gets back the average of everything present. A listener in the Competition Model can be uncertain and withhold. Nothing in the LLM design allows for that.

Arshavir Blackwell, PhD

Models find a way around it, and the workaround is telling. They learn to dump the unwanted weight on one particular token, usually the first in the text, whose value carries almost nothing. It's a parking space, invented rather than installed, that lets a query decline to answer (Xiao et al., 2023). Even abstention has to be improvised out of machinery built for something else.

Arshavir Blackwell, PhD

Degrade the machine and it does not fall silent. It cannot. The weights flatten, every entry contributes a little, and what comes back is the whole context blurred together, the most ordinary thing available. Which may be the mechanism under a sentence from Part 4 of the past-tense series: it slumps toward its most common habit and mumbles ed.

Arshavir Blackwell, PhD

Looking for the part

Arshavir Blackwell, PhD

Another of the six, Annette Karmiloff-Smith, had published a book of her own four years earlier. It was called Beyond Modularity, and it asked whether the mind comes in parts.

Arshavir Blackwell, PhD

A module is a piece of the mind that does one job and only that job, a subroutine, in the programming sense. Written in advance, sealed off from the rest of the code, called when its job comes up. Language in one, faces in another, each shipped with the machine.

Arshavir Blackwell, PhD

Jerry Fodor made the case for it in 1983, in a book whose cover carried a phrenologist's head, a skull mapped into labelled compartments, each one an organ of some faculty. That is the picture at the top of this piece, with the labels changed.

Arshavir Blackwell, PhD

Karmiloff-Smith agreed that grown-up minds end up specialized. What she denied was that they start that way, and she denied that the specialization ever gets as clean as Fodor's picture. A child begins with broad leanings, a pull toward speech, toward faces, toward things that move, and the same circuits, doing the same job over and over, slowly narrow into specialists. Her word for that was modularization: not a module, but the process of becoming one. A matter of degree, running throughout development, and never guaranteed to finish.

Arshavir Blackwell, PhD

That gives you something to test, if you have a system you can take apart. Break it, and watch what breaks with it. If a skill lives in a part of its own, damage to that part should take the skill out and leave everything else standing. If the skill is spread across the whole machine, damage anywhere should cost you a little of it, and cost you everything else just as much.

Arshavir Blackwell, PhD

So I tried it, on the model from my own past-tense work, a newer run, not yet written up. The model has two jobs. One is remembering: the past tense of go is went, and it has seen that a million times. The other is working from scratch: the past tense of wug, a word nobody has ever used, which it has to build.

Arshavir Blackwell, PhD

Damage the model and the made-up words fail first, badly, while the memorized ones hold on. That happens two different ways, switching off parts of the machinery at random, or just adding static to its internal signals. Two unrelated kinds of harm, same casualty.

Arshavir Blackwell, PhD

So I went looking for the part. The model has 32 layers; I switched off each one in turn. The damage concentrated in the first two, right where the model is still assembling words out of fragments. Promising.

Arshavir Blackwell, PhD

Then I went inside, to the heads that share out the lookups, 32 to a layer. I switched off each one singly, in the layer where the damage sat in the lookups, and in two others besides. 96 heads.

Arshavir Blackwell, PhD

Nothing. Not one of them mattered on its own.

Arshavir Blackwell, PhD

Which does not close the question. Induction, remember, takes two heads working together, so a part could still be hiding in a pair or a group, and pairs I have not tested. What is ruled out is any head doing this job by itself.

Arshavir Blackwell, PhD

So the two scales disagree. Zoom out to whole layers and the weakness is somewhere in particular: those first two, and almost nowhere else. Zoom in to the individual heads and it is nowhere in particular: no one of them is carrying the job.

Arshavir Blackwell, PhD

Picture a building where you can prove the fault is on one particular floor, that is where everything goes wrong, and then you test every room on that floor and each one works fine by itself. The floor matters. No room does.

Arshavir Blackwell, PhD

So there is an address, so to speak, but nothing at that address. A weak point with a location and no part in it.

Arshavir Blackwell, PhD

One more check, because the obvious objection is that made-up words are simply harder. Damage the front of the model where the deficit concentrates, then ask it only to repeat wug back: it manages seven times in eight. Ask it for the past tense of wug: not once, in any trial. Same word, same damage. What breaks is the inflecting, not the strangeness.

Arshavir Blackwell, PhD

An address with nothing in it is a strange result if you think there are dedicated parts, and an equally strange one if you think everything is smeared evenly. What it looks like is a middle state: a job that has acquired a location without acquiring a boundary.

Arshavir Blackwell, PhD

Karmiloff-Smith's account is the only one of the three with room for a middle state. Modularization is a process, and a process can be caught halfway. Whether more training would harden this into something you could point at, I do not know, that is the experiment I have not run. But neither of the old positions predicts anything like it, and hers has a name for it.

Arshavir Blackwell, PhD

The loop

Arshavir Blackwell, PhD

Around 1990 Jeff Elman built a small network and trained it by having it guess the next word in a sentence. He was not building a product. He was asking whether a learner could pick up grammar from nothing but the sentences it hears, whether structure could come out of a system that had none put in.

Arshavir Blackwell, PhD

That is the same job every model you have used was trained on. The thing on your phone traces back to that question about children and grammar.

Arshavir Blackwell, PhD

And the models we have now are the instrument the argument always needed. Not because they are human, they are not, but because they are learners you can take apart. That was the one thing missing. Everyone had data, nobody had access.

Arshavir Blackwell, PhD

This is not a one-to-one correspondence: a model has no body, nothing matures, and no one talks back to it. A brain is a patchwork of regions with different wiring; a transformer is the same block copied 32 times. And the timing level, the one Elman et al. cared most about, maps worst of all. A learning-rate schedule is a thin substitute for acquiring a language.

Arshavir Blackwell, PhD

But that difference is not a weakness in the argument. It is what makes the argument clean.

Arshavir Blackwell, PhD

Part of the nativist case was always an in-principle claim: you cannot get grammatical behavior out of a system that has no grammar in it. Statistics over experience would not be enough. Somewhere there had to be rules, and symbols for the rules to work on.

Arshavir Blackwell, PhD

Open a transformer and there is nothing of that kind inside. No rule, no symbol, no grammatical category, no place where subject and verb are declared to agree. Only weights and dot products. Out of that comes agreement across a clause, structure held over distance, and a past tense for a word invented five seconds ago, not reliably, but far more often than chance.

Arshavir Blackwell, PhD

That is an existence proof. Not that people work this way, as a model of a brain a transformer is poor, and I would not defend it as one. Only that the in-principle claim was wrong. A fully distributed system, with no rules and no symbols anywhere in it, can do this. Whether humans do it some other way is now a question you have to answer with evidence.

Arshavir Blackwell, PhD

So this is a convergence, not a vindication. Nobody built attention to settle a psycholinguistics argument, and it would be too convenient if a machine designed for something else happened to prove my advisors right.

Arshavir Blackwell, PhD

But it keeps coming out the way Elman et al. said it would. Structure appears that nobody installed. The parts you would expect are not there. The errors are, to a certain extent, the human errors. And when you strain it, the productive machinery is what gives way first, which is what Liz Bates spent a career arguing about people.

Arshavir Blackwell, PhD

The question I would most like to answer is the one they cared most about and I have barely touched: not what a model ends up with, but when it gets it. Which comes first, remembering, or working from scratch? Does the second one ever stop being the fragile one? Some models ship with checkpoints saved all the way through training, so the whole development is sitting there on disk, waiting.

Arshavir Blackwell, PhD

Neither of them saw this. Jeff and Liz both left us years ago. Those checkpoints are the experiment they could not run. I'm Arshavir Blackwell and this has been Inside the Black Box.