Loading prices...
All articles
A silver measuring ruler with four green cubes standing on its marked section and three crimson cubes hanging in mid air past its cut end

AI stops where checking gets expensive

12:30 · 23.09.2026
2

The question everyone asks about artificial intelligence is whether it can do the job. It is the wrong question, and the past week supplied enough measured deployments to show why. The right question is narrower and much easier to answer in advance: if the work were done badly, would anything tell you, how soon, and who would pay to find out.

Five deployments, five numbers

Five deployments, all with published figures, all from the past several days:

  • Meta: an AI ad suite past a $60bn annual run rate, with more than 9 million small businesses using at least one creative tool.
  • Klarna: resolution time cut by at least 81.8%, then human agents rehired.
  • Coca-Cola: positive sentiment around an AI Christmas film falling from 23.8% to 10.2%.
  • US policing: at least 14 wrongful arrests from face matches, and zero reported in the 20-plus jurisdictions that banned the technology.
  • NVIDIA: an AI agent refactoring production robotics code in five edits, then proving it with a profiler.

The purchase metric always arrives first

Every one of those numbers exists because someone counted something. The difference between the successes and the reversals is not the difficulty of the task, and it is not the quality of the model. It is which number arrived first. A production metric is available on the day of installation: minutes saved per report, alerts generated, variants produced per hour, conversions per thousand impressions. A correctness metric arrives later, is expensive to produce, and is almost always paid for by someone other than the buyer. Meta is the case where the gap does not exist: the thing being optimised and the thing being judged are the same event. A click either happened or it did not, the advertiser sees it the same hour, and nobody has to convene a study to find out whether the campaign worked. That is why the ad suite scaled past a $60bn run rate without a single reversal story attached to it, and it is not because advertising is an easier problem than customer service.

Klarna is the cleanest demonstration because the part that worked never stopped working. Its assistant did the volume it was given and cut handling time by at least 81.8%, and that figure held all the way through. What degraded was the thing nobody had a column for, and by May 2025 the company was rehiring the people it had replaced. Coca-Cola shows the same shape from the other side: running an AI Christmas film for a second year, positive mentions fell 13.6 percentage points while negative ones rose 0.6. The positive side moved 22.7 times further than the negative side. The campaign did not make people angry. It drained the thing the campaign existed to produce, and there was no dashboard for that.

As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.

Sebastian Siemiatkowski, Klarna, quoted by Forbes

Sebastian Siemiatkowski, chief executive of Klarna, on what the deployment produced

Where checking takes seconds

Now the opposite case, and it is the most instructive of the five. NVIDIA published a walkthrough in which an AI agent migrates a CUDA-accelerated robotics node: it audits the code, traces where data crosses between GPU and CPU memory, and writes the smallest patch that removes those crossings. The transport change is five edits. The workflow has seven steps and three of them are audit or verification.

What makes it work is the last step. The agent must confirm with a profiler that no payload-sized host-device transfers remain, and check that the message reports its backend type as cuda. That is a trace and a string comparison: two mechanical tests, running in seconds, returning yes or no, with no human judgment in the middle. Because the check is that cheap, the agent can be handed the work rather than the first draft of it. We measured that deployment in detail, and the pattern held: the verification, not the capability, is what licensed the autonomy. The same agent given the same task inside a codebase with no profiler and no backend assertion would be exactly as capable and nobody sensible would let it commit.

Where checking takes years

Policing sits at the far end of the same axis, and the consequences scale with the delay. Facial recognition has produced at least 14 documented wrongful arrests in the US, with 13 criminal cases dismissed; NIST has measured false positive rates up to 100 times higher for Black and Asian faces than for white male faces. Predictive software audited by The Markup was right less than half of one percent of the time, 0.6% on robbery and assault and 0.1% on burglary. Every one of those tools kept producing its production metric right up to the day it was switched off. One case this year put a woman in jail for more than five months over crimes in a state she says she never visited. The correctness check arrived in a court docket, years after the purchase order.

Search is the case where the correctness metric does not exist at all, which is why nothing has stopped. An AI answer above the links cuts the click rate from 15% to 8%, links inside the answer are clicked about 1% of the time, and 26% of those searches end the browsing session outright against 16% without. Every one of those is measurable for the platform. Whether the reader got what they came for, and whether the publisher who wrote it survives, is not on anyone's dashboard. The system optimised the half it could see.

What verification costs in dollars

The cost of checking is not a metaphor, and three published prices make the point. Rev charges $0.07 a minute for machine transcription and $0.79 a minute for human-edited work, 11.3 times more; on a one-hour recording that is $4.20 against $47.40. OpenAI pledged $1bn toward defending critical infrastructure, which is 0.067% of the $1.5tn of compute contracts the same industry has committed to over the next two years. Apple paid $250m over an AI feature it advertised and shipped 813 days late, which works out at an estimated $25 a device. In each case the machine output is nearly free and the verification is the expensive part.

Dictation is where an ordinary reader meets this directly. Whisper scores 2.7% word error on clean benchmark audio and 8% to 12% on real speech, so the working error rate is three to 4.4 times the number quoted in marketing. On a fifteen-minute recording that is 180 to 270 wrong words. Separately, hallucinated phrases appear in about 1% of samples and 38% of those carry explicit harm, with the trigger conditions being pauses longer than 30 seconds, disfluencies and interruptions. That is a description of an ordinary conversation, and the arithmetic is worth knowing before you trust a transcript.

A rule you can apply tomorrow

Three questions, in order, before any AI deployment. If this task were done badly, would a number tell me? If yes, how long until that number exists? And who pays in the interval. Where all three answers are good, as at NVIDIA, hand over the work. Where the first is yes but the second is months, as at Klarna, keep a human on the exceptions and expect the reversal if you do not. Where the first answer is no, as with a Christmas film or an arrest, you are not deploying a tool. You are transferring judgment, and the bill arrives where you cannot see it.

There is a reading of this that matters specifically to a crypto audience, and it is not a slogan. A blockchain is, stripped of everything else, a machine that makes one class of claim cheap to verify and impossible to assert without proof. This week's evidence says adoption of AI binds on exactly that constraint: not on how capable the model is, but on how cheaply anyone can establish that it was wrong. Verification has quietly become the scarce input. The firms that built cheap checks are handing machines real work. The ones that did not are rehiring humans and explaining sentiment charts.

Informational material, not investment advice. Every figure here comes from a company disclosure, a peer-reviewed study or a published audit, linked where it first appeared; the comparisons between them are ours.

Published: 12:30 · 23.09.2026
Intokened.com

Author

Intokened.com

Crypto and AI news, market analysis and reference guides

Comments (0)

No comments yet — be the first!

The market talks all day. We write when it says something

Short, and it tells you why it came