The jagged frontier in 2026: why AI still fails at things that look easy
The idea that AI is unevenly capable is not new. What is new is this year's evidence about why, and it points somewhere more useful than the original finding did: the unevenness is structural, it is still not properly measured, and it changes shape with every model generation.
Where the term came from
"Jagged frontier" was coined in a 2023 field experiment with 758 consultants at Boston Consulting Group, by Fabrizio Dell'Acqua, Ethan Mollick and colleagues. Inside the model's range, the consultants using AI completed 12.2% more tasks, 25.1% faster, at more than 40% higher rated quality. On one task deliberately designed to sit outside that range, they did worse than the people working without AI: 84.5% of the control group got it right, against 60% and 70% for the two AI groups.
That result is worth knowing, and it is also three years old, measured on a model from early 2023. The fair question in 2026 is whether any of it still applies to systems several generations further on. It is not a rhetorical question, and the honest answer is that the specific numbers should not be treated as current.
What the 2026 evidence actually says
In January 2026, revised in May, a team at Google DeepMind published a position paper, Characterizing Jaggedness Aids Safety and Usability, making two claims that matter more to a business owner than the original productivity numbers do.
The first: jaggedness is a structural property of current architectures and scaling paradigms. Not a defect in one model, not a phase the technology is passing through on the way to being uniformly reliable. On the evidence available now, it is a consequence of how these systems are built. Waiting for the release that makes it go away is not a strategy.
The second is quieter and more surprising: as of 2026, jaggedness is still not systematically measured. The paper exists because the metrics did not, and the authors had to propose them. There is no standard number on a model card telling you where the valleys are. Which means nobody, vendor or consultant, can currently look up whether your particular task falls in one.
If the field's own researchers have to invent the measurement in 2026, treat anyone who claims to already know where AI will fail in your business as guessing.
The shape moves between generations
The paper works through a comparison of three model generations, scoring capability against a human baseline. Higher is more capable. Separately, the jaggedness index tracks how uneven that capability is.
| Model generation | Capability vs human | Jaggedness |
|---|---|---|
| Gemini 1.5 Pro | -1.37 | baseline |
| Gemini 2.5 Pro | 0.97 | more jagged than 1.5 |
| Gemini 3 Pro | 1.73 | smoother than 2.5 |
Capability rose steadily. Jaggedness did not: it got worse, then better. The generation that smoothed out did so largely by closing a gap in visual reasoning that the previous one had left open. The frontier is not just advancing, it is changing shape, and it does not do so evenly.
The practical consequence is the most actionable thing in this post. Your list of things you tried and abandoned has an expiry date. A task that failed a test eighteen months ago may well pass now, and a workflow that has been quietly working may sit closer to an edge than you think. If you tested once and filed the answer permanently, the file is out of date.
Why the demo works and your version does not
The paper's most useful idea for anyone buying this stuff is the gap between two things: how jagged a model looks on benchmarks, and how jagged it feels to the people actually using it. The authors argue that a large gap between the two is itself a warning sign, indicating either that the deployed system is hard to use well, or that the benchmarks do not represent real work, what they call a lack of ecological validity.
That is the technical description of an experience most business owners have already had. The demo, on the clean example, sits inside the benchmark's world. Your enquiries, with the missing fields and the attachment someone sent as a photograph, do not. A vendor's benchmark score is evidence about the benchmark. It is not evidence about your workflow, which is why the distinction between a tool and a wired system matters so much, and why we go on about it in what AI implementation actually means.
What to do about it
- Test on your ugliest real inputs, not a representative sample. The gap between the tidy example and the messy one is exactly where the edge tends to run.
- Expect the failures to be fluent. Nothing in the output announces that a task fell in a valley. Work that is wrong and confident is harder to catch than work that is obviously bad, so a review step by someone who knows the domain is not optional.
- Re-test on a schedule. Because the shape moves between generations, put a review date on both your automations and your list of rejected ideas. This is a maintenance cost, and it belongs in your budget, which is part of the honest answer in what AI implementation actually costs.
- Automate the task, not the job. A job is a bundle of tasks scattered across both sides of the edge. Take them one at a time.
The uncomfortable part for anyone selling this
If the measurement does not exist yet, then nobody, including us, can tell you in advance which parts of your business AI will handle well. What is honestly on offer is a method rather than a prediction: pick the candidate tasks, test them against your real inputs, keep what clears the bar, drop the rest without sentiment, and check again when the models move. That is the process behind AI services for businesses on Mallorca, and the reason our first step with anyone is a diagnostic rather than a proposal.
It is also why the questions in how to choose an AI consultant lean so heavily on asking how someone would know they were wrong. On a frontier this shape, that is the only skill that reliably transfers.
The takeaway
The 2023 headline was that AI made good consultants better. The 2026 finding is more useful and less comfortable: the unevenness behind that result looks structural, the field cannot yet measure it, and its shape shifts under you with each release. None of that is an argument against using AI. It is an argument for testing your own tasks, on your own inputs, more than once.