Your gut is miscalibrated
The task you're sure a model can't handle may be trivial for it; the "simple" one may be where it quietly fails. Assumptions about capability are the least reliable input you have.
A professional outfielder catches a fly ball without solving a single equation. A language model trained only on text "knows" a dropped glass will shatter — without ever having seen one fall. Both break the same assumption: that the things which feel difficult to us are the difficult things. They aren't. This is a short, interactive tour of why our intuitions about intelligence are backwards — and why that should change how you adopt AI. The science is real; sources are listed at the bottom.
To plot a fly ball on paper you need projectile motion, air resistance, and a differential equation or two. A professional outfielder does none of that. They read the ball off the bat, break into a run, and arrive where it lands — consistently, in under four seconds, with no idea of the math involved. The knowledge is real. It just isn't the kind you can write down.
Demis Hassabis, who leads Google DeepMind, has posed a sharp version of the question: if a model were trained only on written text — never a photo, never a video, never a falling object — would it understand that a dropped glass breaks? Intuition says no. Modern language models say otherwise. Language itself carries a dense map of physical reality, because the way we talk about the world mirrors how the world actually behaves. Tap each question to see what a text-only model can infer — and why.
A model reading only text has never perceived anything. Yet ask it these and it answers correctly. Where does the knowledge come from?
This is the baseball player in reverse. The outfielder has deep physical knowledge they can't express in words. The text-only model has deep physical knowledge it acquired only from words. In both cases the surprising thing isn't the skill on the surface — it's how much understanding sits underneath, in a form we didn't expect it to take.
Here's the pattern both stories point to, named after roboticist Hans Moravec: high-level reasoning is cheap for computers; low-level sensorimotor skill is staggeringly expensive. Chess, law exams, and calculus fell early. Folding laundry and walking over gravel remain unsolved. Before you scroll, test your own intuition — for each task, guess whether it's easy or hard for today's AI.
"It is comparatively easy to make computers exhibit adult-level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility."
"The main lesson of thirty-five years of AI research is that the hard problems are easy and the easy problems are hard."
The reason is evolutionary. We spent hundreds of millions of years perfecting perception and movement, so it feels effortless — the computation is just hidden from us. Abstract reasoning is a thin, recent layer, so it feels like the hard, "smart" work. Machines invert that: the recent layer is the easy one to reproduce. What both humans and AI share is a huge reserve of unknown knowns — things they can do but can't explain. The catcher doesn't know the calculus; their body knows the trajectory.
The practical upshot is not philosophical, it's operational. If human intuition about what's "hard" for intelligence is demonstrably backwards, then it's a poor tool for predicting what a model can do for your business. The teams that win with AI don't win the argument about capability — they run the test.
The task you're sure a model can't handle may be trivial for it; the "simple" one may be where it quietly fails. Assumptions about capability are the least reliable input you have.
A one-afternoon pilot on real data settles a debate that a month of meetings can't. Treat capability as an empirical question, not a matter of seniority in the room.
Don't benchmark the model on toy demos. Point it at the actual, messy work you want changed and measure whether the outcome improves. That's the only signal that counts.
Go, protein folding, and passing the bar were all "decades off" until they weren't. Forecasts about AI limits have a habit of collapsing. Re-test on a cadence.
That's the shift: from arguing about what AI can or can't do, to designing small, honest experiments that tell you. The hard part of AI in a business is rarely the model — it's replacing assumptions with evidence about where it actually helps. That's the conversation we have with leadership teams before they commit budget.