AI capability isn't a smooth line from dumb to smart — it's a jagged frontier. The same model that outscores law graduates on the bar exam can fumble a question your eight-year-old would ace. In this quiz you don't answer the questions — you predict what the AI does. Get a feel for exactly where the frontier zigs and zags.
Part 1 shows you prompts that famously trip chatbots up: predict how the model responds. Part 2 flips it: guess how well frontier models score on the hardest exams humans take. You're scored on how well you can predict the AI — not on the trivia itself.
You've just watched the model faceplant on questions a child gets right. Now the frontier zags the other way: the same class of models sits in the top few percent of humanity on the hardest standardised tests we have. For each benchmark, guess how frontier models actually perform.
Researchers call it the jagged frontier: AI capability doesn't advance as a smooth wave, it advances as a coastline full of peaks and inlets. Tasks that look equally difficult to a human can sit on opposite sides of the frontier — one trivially inside the model's competence, the other far outside it. The traps in Part 1 all exploit the same weakness: models are pattern-matchers, not embodied reasoners. The benchmarks in Part 2 all reward the same strength: absorbing and recombining the entire written output of humanity.
| Benchmark | Human context | Frontier AI (2026) |
|---|---|---|
| Uniform Bar Exam | ~68% typical human pass rate | 90th percentile — outscores most law school graduates |
| The SAT | 1050 national average | 1500+ — a 99th-percentile, Ivy-competitive score |
| Codeforces | Fiercely competitive global programming ranking | 96th percentile — beats 96% of human competitors |
| AIME | Strong math students solve only a fraction of it | 80% to near-100% — multi-step olympiad math in minutes |
| GPQA Diamond | ~34% for PhD experts answering outside their specialty | ~70–87% — graduate-level science across every field at once |
The contrast is the entire point. If you treat an AI like a human employee, you'll be blindsided when it fails simple physical logic. If you understand it as a very different kind of intelligence — one that struggles with basic physical reality but has ingested the entirety of human mathematics, code, and law — you can start using it for what it actually is.
Models don't see letters — they see tokens, chunks of text handled as single puzzle pieces. That's why counting the r's in "strawberry" is genuinely hard: the model is nearly blind to the letters inside its own tokens unless it slows down and reasons step by step.
"500 m" co-occurs with "walkable" across billions of training examples, so the model recommends walking — to a car wash. It has read everything about the physical world without ever living in it, so statistical association quietly substitutes for common sense.
Show a model 95% of a famous riddle and it will confidently complete the famous answer — even when you changed the premise. The pull of a well-worn pattern can override the actual words in front of it. Reasoning models mitigate this, but the reflex runs deep.
Standardised tests — bar exams, olympiad math, graduate science — are how labs measure model capability against humans. They reward exactly what models are best at: vast recall plus fast multi-step symbol manipulation, with no physical world in the loop.
The boundary of AI competence is jagged, not smooth: tasks of similar difficulty for humans can fall on opposite sides. The practical skill of working with AI is learning the shape of that boundary for your own work — task by task, not vibe by vibe.
The productive mental model isn't "junior employee" or "search engine" — it's a different kind of mind entirely: encyclopedic, tireless, superb at synthesis, and unreliable about anything it must physically inhabit to understand. Delegate accordingly.
The jagged frontier is one lens on what AI can and can't do. These tools fill in the mechanics behind it.
Mapping the jagged frontier for your organisation — which workflows sit inside AI's competence and which only look like they do — is exactly the kind of strategy work we help teams get right.
Talk to us about your AI roadmap