How AI Works, Part 3: How Do Machines Learn Language?

Two weeks ago I asked what intelligence even is. Last week, I went through the five different ways machines learn: biology, evolution, probability, first principles, analogy. Now it’s time to talk about the one everyone actually means when they say “AI” these days: large language models. ChatGPT. Claude. Gemini. Whatever you’re using, this is the technology underneath it.

So what does “machine learning” actually mean here? Not the five types from last week, all mixed together in some elegant way. Mostly, it means something much blunter: feed a machine an enormous amount of text, and get it to predict the next word, over and over, more times than any of us could really picture, until it gets very, very good at it.

That’s it. That’s the trick. Read a sentence, guess the next word, check the answer, adjust, repeat.

How Did We Get Here

For a long time, this wasn’t the plan. Researchers spent decades trying to teach computers grammar, rules, structure, the way you’d teach a child to diagram a sentence. It didn’t really work, not at scale anyway. Language is too messy, too full of exceptions, too dependent on context nobody thought to write down as a rule.

What changed wasn’t a clever new idea about language. It was scale. Once we had enough text (basically, most of the internet) and enough computing power to chew through it, a strange thing happened. Predicting the next word, done at a big enough scale, started looking a lot like understanding. Not real understanding, we covered that distinction back in the first post, but something that behaves close enough to it that the difference stops mattering for most everyday purposes.

Training One of These Things

The actual training process is almost embarrassingly simple to describe, even if the engineering behind it is anything but. You take a transformer, a particular kind of neural network architecture built around something called “attention,” which lets the model weigh how relevant every other word in a sentence is to the one it’s currently predicting. Then you feed it a genuinely staggering volume of text. Not gigabytes. Not terabytes. We’re talking about a meaningful fraction of everything human beings have ever written down.

The model doesn’t “read” the way you do. It’s adjusting an almost incomprehensible number of internal parameters, tiny dials, essentially, based on how wrong its guesses were, over and over. Eventually those dials settle into a configuration that’s remarkably good at continuing a sentence in a way that sounds like a person wrote it.

What Makes This Genuinely Hard

Scale is the challenge and the solution at the same time, which is an odd position to be in. Training one of these models can cost tens of millions of dollars in computing power alone. Get something wrong deep in that process, and you don’t find out until you’ve already spent the money.

Then there’s the fact that these models don’t actually “know” anything, not in the way we talked about a couple of weeks back. They’re pattern-matching machines with astonishingly good instincts. Which means they can be confidently, fluently wrong: hallucinating facts, inventing sources, stating nonsense with exactly the same tone of authority as something true. That’s not a bug you patch. It’s baked into how the whole system works.

The Pushback

Here’s the part that doesn’t get talked about enough in the excitement: where did all that text actually come from?

The honest answer, for most of these models, is: everywhere. Books, articles, forum posts, blogs, personal websites, all of it scraped and used to train systems worth billions of dollars, largely without asking the people who wrote any of it. Authors, artists, journalists, coders, most of us never got a call. There are lawsuits working through courts right now trying to sort out whether that was ever legal in the first place, and nobody really knows yet how those are going to be decided.

Do We Even Need to Create Anything New?

Here’s the question that actually keeps me up at night, more than the copyright one. If a machine has read essentially everything humans have ever written, learned every pattern, every trope, every trick, what’s left for the rest of us to actually contribute?

I don’t think the answer is “nothing,” but I understand why it’s tempting to worry that it might be. My best guess is that what we bring that a model trained on the past genuinely cannot is the next thing that hasn’t happened yet. New experience. New events. The way it actually feels to live through whatever comes next. A model can remix everything that’s already been written. It can’t live your life for you, and it can’t have opinions about tomorrow that haven’t been written about yet, because nobody’s written about tomorrow.

Where This Fits

I keep coming back to something from the first post in this series: that the story of human intelligence has always been about getting faster and better at passing knowledge from one person, one generation, to the next. Language did that. Writing did that. Now this is doing it too, just at a scale and speed nothing before it has managed.

Next week, we leave language behind for a bit and look at how machines learn to interpret images instead, which turns out to change the whole equation in a way that surprised me when I first started researching it for the book.


3 Replies to “How AI Works, Part 3: How Do Machines Learn Language?”

Comments are closed.

Reader Feedback

"Great page turner, so relevant and timely! Really enjoyed this one! Especially enjoyed the technical (both AI and military) accuracy that never got in the way of a fast moving, fun story."

— Barry Parsons, August 9, 2026
See more reader comments →
FAQs
Where can I buy The Apex Code?

Paperbacks and hardcover copies are all available directly from this site, or you can get them (unsigned!) along with the ebook through Amazon, Barnes & Noble, Waterstones and other major retailers.

How do I let you know who I want my signed copy dedicated to?

There is an option to add a note when you are ready to checkout. Fill this out with your dedication request just before you place your order!

When is The Axion Code coming out?

Later this year — sign up for the newsletter to be the first to know the release date.

Is The Axion Code a sequel, or can I read it standalone?

It's the second book in the series — while it can be read on its own, The Apex Code sets up the characters and world that carry through.

Do you have events, signings, or appearances coming up?

Check the News & Blog page for updates — new events are posted here as they're scheduled.

Are you running any giveaways right now?

Yes! Five signed copies of The Apex Code are up for grabs for UK readers. Visit the UK Giveaway page for full details on how to enter.