How AI Works, Part 4: How Machines Learn to See

In the first three posts in this series, we talked about what intelligence is, how machines learn, and how they use language. This week we’re looking at something different: how a machine sees.
I don’t mean that literally, of course. There’s no eye behind any of this, no retina, no glance. There’s just a grid of numbers, and a machine trained to guess what those numbers usually mean. That gap between raw pixels and a confident answer, tumor, pedestrian, or a cat playing a piano, is the entire discipline of image interpretation. I’d argue it’s one of the biggest reasons AI is starting to feel like something that can act in the world, rather than just talk about it.
Teaching a Machine to See
A digital image is really just a grid of pixel values, three numbers per pixel for red, green, and blue. Early computer vision tried to code rules for finding edges and shapes, and it was brittle. Change the lighting and the whole thing fell apart.
The real breakthrough was something called the ‘convolutional neural network.’ A small filter slides across the image like a magnifying glass, learning to identify patterns: an edge, a curve, a patch of texture. Stack enough of these filters in layers, and the network starts building a complexity of its own. Early layers find edges, middle layers combine them into shapes, and deep layers assemble those shapes into whole objects. Nobody writes a rule that says a nose looks like a nose. The network works it out from millions of labeled examples, adjusting billions of internal parameters until its guesses start matching the answer.
Newer systems, called vision transformers, take a related but different route. Instead of building up locally, they break an image into patches and let the model weigh how every patch relates to every other one at once. Increasingly, the strongest production systems are blending both approaches together.
What This Adds to the Picture
Language models are extraordinary at manipulating words, but words are already a compressed, human-made abstraction. A photograph, an X-ray, a security camera frame is a 2D version of reality: unlabeled and messy. Teaching a machine to reliably turn that mess into structured information (this is a person, this is a tumor, this is a car running a red light) is what lets AI move from answering questions to perceiving a situation. It’s arguably the missing piece that gets you to a system that can act, which then gives the impression it can think, and an area that the Apex Code series is exploring. Okay, what I’ve written is not perhaps fully real yet, in some areas, but we are not that far away from most of it today, either. Every time I read the news, I think by next week, or next months, certainly AI capabilities are moving very fast!
Where We Are Today
The scale of adoption is already largely invisible, which is sort of the point. As of early 2026, radiology accounts for roughly three out of every four AI medical devices cleared by the FDA, more than 1,150 algorithms in total, with new ones arriving at a pace of dozens a month. This isn’t experimental anymore. It’s embedded infrastructure.
Catching What Might Get Missed
The clearest real-world case for image interpretation is medical imaging. A nationwide study across German breast cancer screening programs, covering more than 460,000 women, found that AI-assisted mammogram reading increased the cancer detection rate by nearly 18 percent compared to standard double reading by radiologists, without a meaningful rise in false alarms. In a scenario where routine, AI-flagged-as-normal scans skip a second human read entirely, the same study found radiologist workload could drop by more than half while detection rates stayed higher, not lower.
That’s the case for image interpretation in one sentence. It isn’t about replacing the expert. It’s about making sure an exhausted, overworked, entirely human doctor doesn’t miss the one scan that matters.
The Pushback
None of this comes free of real problems. The most documented failure of image-based AI has been bias in facial recognition. The good news is measurable progress: independent testing found top-performing algorithm error rates fell from around 4 percent in 2014 to roughly 0.02 percent by 2022. The harder, unresolved problem is what to do with a technology that keeps improving at exactly the task, mass identification from a camera feed, that raises the most legitimate privacy and civil liberties concerns. Think ‘Eagle Eye’ and ‘Person of Interest’ levels. Better accuracy doesn’t settle the argument over whether the surveillance itself is a good idea.
How I Have Used This in My Writing
This is some of the capability I gave the rogue AI in The Apex Code. The AI doesn’t just have access to security camera feeds. It interprets them: recognizing a face, reading a posture, understanding that a door held open two seconds too long means someone is trying to slip through, then reacting in real time, before a human watching the same feed would even register that something was wrong.
Writing that meant thinking through the evolution of what we are seeing today, because “it watches the cameras” is a premise. “It converts a video feed into structured understanding fast enough to act on it” is the mechanism that makes the premise believable. This is exactly how the AI watches and interprets Riley and Decker’s attack on the facility where it is housed, completely ready and waiting to trap them…
Next week, we’ll look at automation, where everything this series has covered so far, learning, language, perception, finally comes together into a system that doesn’t just understand the world, but does something about it.

One Reply to “How AI Works, Part 4: How Machines Learn to See”
Comments are closed.