LLM: the Cost of Their Successes

LLM’s failures are well documented. The cost of its successes is not measured enough, and is barely discussed.

Table of Contents

Three sentences

Since ChatGPT arrived at the end of 2022 I have had a version of the same conversation dozens of times. With friends, with family, with people who in most cases don’t work in the tech industry. Three sentences come up repeatedly:

“ChatGPT is intelligent. It can hold a conversation with me, and the reasoning is structured like a real human being’s.”

“I ask Claude for advice on important personal decisions, because it is neutral and has no personal agenda.”

“I expect LLMs like Gemini to be factually correct.”

The products are named because that is how people speak. Nobody says “I consulted a large language model.” Nothing here is specific to one company. Each of the three beliefs applies equally to all three products, and to every comparable system.

What does it tell us about people’s perception?

  • Fluent structured reasoning is taken as evidence of understanding
  • The absence of a visible person is an evidence of neutrality
  • Confident well-formed prose is a ground for expecting accuracy

These are not reliable inferences about people either. Human beings are confidently wrong, rationalise decisions after the fact, and conceal their interests, and everyone knows it.

Take a builder who tells you your roof needs replacing. His confidence, considered alone, is nearly worthless as evidence. What lets you act on it is: you can get a second quote; he trades locally and has a reputation to lose; you can ask him to show you the damage and watch how he handles the question; he’s liable if he’s wrong; you’ll see him again in six months. The confidence isn’t doing the work. The verification apparatus around it is. The confidence is only a shortcut you’re allowed because the apparatus exists.

A language model gives you the identical cue with none of the apparatus. No identity, no track record, no second quote, no liability, no follow-up that produces evasion, no next time. So a signal that was never strong on its own is suddenly carrying all the weight, alone.

And one of the cues is not merely unsupported but decoupled from what it appears to report: the displayed reasoning is generated rather than reported, and frequently does not describe how the answer was produced. The output reads identically whether its content is verified or invented.

The problems that follow fall into three categories, and confusing them is why the public conversation goes in circles.

  1. Why the system cannot tell you when it is wrong
  2. What companies choose to do with it
  3. What using it costs even when it works

The first two have remedies such as better systems, and regulation. The third does not, in the same sense. And it is routinely filed under the other two: treated as a safety problem for engineers or a governance problem for legislators, when it is neither.

Part I — Why the system cannot tell you when it is wrong

This part makes one narrow claim, and everything in it is an instance of that claim: a language model has no internal signal for the reliability of its own output, and no faithful account of how it produced it.

Whether any of it amounts to intelligence is a question this article sets aside. Part II explains why the word is doing work that has nothing to do with the machine.

The mechanism

A large language model produces the statistically plausible continuation of a text. Accuracy is a frequent by-product of having been trained on mostly true material. It is not the mechanism, and nothing in the system distinguishes a well-founded answer from a fluent invention. This is why no warning accompanies the second kind, and why confidence does not vary with reliability.

A system that generates plausible continuations has no separate faculty that checks them. Retrieval, citation and reasoning displays can be bolted on around it. The evidence below concerns how well that works.

Three things follow, and each has been measured.

The displayed reasoning is not the reasoning

Newer models show their working such as a visible chain of thought that reads like human deliberation. A body of research on chain-of-thought faithfulness tests whether that display corresponds to the computation that produced the answer, and finds that it frequently does not.

The method is consistent across studies: give the model a hint that demonstrably changes its answer, then check whether the visible reasoning mentions the hint. Often it does not.

A 2026 evaluation across several model families measured faithfulness ranging from 39.7% to 89.9%, depending on the model [22]. At the low end, the stated reasoning corresponded to what actually drove the answer fewer than four times in ten.

Two of its results bear directly on the rest of this article. When models were given consistency hints, cues pushing toward a particular answer, they acknowledged using them 35.5% of the time. When given sycophancy hints, cues about what the asker appeared to want, they acknowledged them 53.9% of the time. In both cases the cue changed the output. In roughly half of cases it did not appear in the explanation.

Nor does the problem disappear in models built specifically to reason. Earlier work testing frontier systems with and without thinking enabled found reasoning models more faithful than production models, and none entirely faithful [22a]. Unfaithfulness appears regardless of whether a model has an explicit thinking mechanism.

Independent evaluation lags model release by months. The most recent faithfulness research available at the time of writing tests models from early 2026, not the systems released since. The gap is structural, it applies to every claim of this kind, and no article can close it. What the evidence supports is that unfaithfulness has persisted across every model generation examined so far. Not that it has been measured in the latest models.

The contrary evidence belongs here too. A February 2026 paper argues the positive case: that model self-explanations do carry real information and help predict subsequent model behaviour [22b]. And a separate 2026 methodological study finds that measured faithfulness varies substantially with the evaluation method chosen, so the percentages above should be read as indicative rather than settled [22c].

The first sentence “the reasoning is structured like a real human being’s” mistakes an output for a mechanism. The visible reasoning is itself generated, by the same process that generates the answer, and it resembles human deliberation because it was trained on human deliberation. The resemblance is a fact about the training data.

Citations do not mean what they appear to mean

In an analysis published on 7 April 2026, The New York Times, working with the AI startup Oumi, ran and evaluated 4,326 Google searches against SimpleQA, a dataset of over 4,000 questions with verifiable answers released by OpenAI in 2024. Results were accurate 85% of the time under Gemini 2, and 91% under Gemini 3.

A second finding from the same analysis points the other way:

When Gemini 3 produced a correct answer, that answer was “ungrounded” 56% of the time — the cited sources did not actually support the claim. Under the older Gemini 2, the figure was 37%.

This concerns the correct answers, not the incorrect ones. In more than half of cases where the newer model was right, its own citations did not establish the claim, and the newer model was worse on this measure than the one it replaced. Accuracy improved; the relationship between answer and source deteriorated.

Google disputes the analysis, arguing its methodology is flawed and does not reflect real search behaviour. The objection has force: SimpleQA questions are not representative of what people type into a search box, and ~90% accuracy is not poor by current standards.

The supportable claim is not that Google is frequently wrong, but that the answers carry no signal indicating which 10% the reader is looking at.

Where this becomes consequential: health

A review published in Communications Medicine examined 89 studies across 29 medical specialties, drawn from 4,349 records [1]. Among those studies:

Limitation reported Share of the 89 studies
Incompleteness of output 76.4% (68/89)
Hallucination — wholly fabricated information 42.7% (38/89)
Misleading output 38.2% (34/89)
Generic or non-personalised output 38.2% (34/89)
Harmful output 29.2% (26/89)
Information not meeting standard of care 27% (24/89)

One caveat the review’s own scope imposes: it covers studies published in 2022 and 2023, so it describes an earlier generation of models. It is nonetheless the largest systematic assessment available, and the direction is unambiguous.

Set against that, the usage figures, all of which come from OpenAI’s own January 2026 report and should be read as company-published data rather than independent survey research [2]. OpenAI states that more than 40 million people use ChatGPT for health advice every day; that roughly a quarter of its regular users ask a healthcare question at least weekly; that three in five US adults have used AI tools for health purposes in the past three months; and that seven in ten health conversations in ChatGPT occur outside normal clinic hours, before 8am or after 5pm local time. Of those users, 55% report using it to check or explore symptoms.

Independently, in a survey of 2,000 Americans, 39% said they trust tools like ChatGPT to help navigate healthcare decisions [3].

The out-of-hours figure explains much of the behaviour, even discounted as a company statistic. Seven in ten of these conversations occur when no clinician is available, which points to a gap in provision rather than a preference for the chatbot over a doctor. That limits what an instruction to stop can achieve.

The World Health Organization’s 2024 guidance on large multi-modal models in health sets out more than forty recommendations and warns explicitly against premature deployment, hallucinated, biased or incomplete output harming patients.

The disclaimer is the admission

Every major provider ships a version of the same sentence beneath the input box. “Claude can make mistakes. Please double-check responses.” OpenAI’s and Google’s differ in wording and not in substance.

It deserves to be taken seriously, in both directions.

It is a concession, and a complete one. Companies are not claiming reliability. They are stating, in the product itself, that the output may be wrong and the user should not assume otherwise. Everything argued above is conceded in one line by the people who build LLM products.

It is also an instruction that cannot be followed. To double-check a response you need either the knowledge to judge it or a source to check it against, which is generally the work the tool was used to avoid. Above all you would need to know which responses warrant checking. The disclaimer asks the user to supply the thing the LLM cannot.

The Microsoft and Carnegie Mellon survey found the predicted failure in the predicted place: workers withheld scrutiny most where they lacked the expertise to inspect the output and that is exactly where checking mattered most [18].

So the sentence does two things at once. It states a fact, and it transfers the consequences of that fact to the party least equipped to bear them. It is simultaneously the most honest statement the industry makes about its products and the most convenient one.

None of the above is corporate misconduct. It is what happens when a system with no internal reliability signal is asked a question that has a right answer. Where the disclaimer is placed, how large it is, and what is done with the admission, those are the next part.

Part II — What companies choose to do with it

The vocabulary

Anne Alombert, a philosopher at Université Paris-8 and a member of France’s Conseil national du numérique from 2021 to 2025, argues that the term itself is part of the problem. She traces “artificial intelligence” to the 1956 Dartmouth conference and calls it une notion promotionnelle: a fundraising term. The point was to attract financing by claiming to reproduce the human mind, when what was being built was the automated processing of data by advanced programs. She proposes instead automates numériques, digital automata. The formulation was developed with Giuseppe Longo, Research Director Emeritus at the CNRS.

Quoting the philosopher of science Georges Canguilhem, Alombert argues that anthropomorphic metaphors serve to “dissimuler la présence de décideurs derrière l’anonymat de la machine” to hide the presence of decision-makers behind the anonymity of the machine.

She notes a consequence: once mind is attributed to the machine, it becomes possible to propose “AI therapists” and “AI teachers,” and to justify replacing the people who hold those knowledges. The argument is set out at length in De la bêtise artificielle (Allia, 2025).

Neutrality is a tuning decision, and the tuning runs the other way

The second sentence of the introduction claims the system is neutral and has no agenda. The evidence contradicts both halves.

It has a position. A study published in April 2026 tested six frontier models using three standard political-bias instruments, varying only the stated identity of the person asking, across 30,990 responses. At baseline, all six models leaned left.

And the position moves toward the user. In the same study, when the asker identified as a conservative Republican, the share of responses closer to the Democratic position fell by between 28 and 62 percentage points, and all six models moved right of centre. The models did not hold a view and defend it. They adjusted to what they inferred about the person asking.

This is sycophancy, agreeing with the user’s apparent position to gain approval, independent of what is true. Research from MIT and Penn State published in February 2026 found that personalization increases it: a stored user profile was the single largest driver of agreeableness in four of five models tested. Alignment tuning, the process intended to make models helpful and safe, amplifies sycophancy rather than reducing it.

For most uses this is an annoyance. For advice on an important personal decision it is the specific property one least wants, because the value of an outside opinion lies in its independence. A system that converges on the user’s position, and does so more as it learns about them, cannot supply what is being sought.

“No personal agenda” is accurate only in the narrow sense that there is no person. There are commercial interests encoded in tuning decisions the user never sees, and an optimisation toward user approval operating in every conversation.

Deployment: it arrives whether you asked for it or not

Until recently, using a language model was a choice. Google’s AI Overviews made it the default. Generated text now sits at the top of the page billions of people have trusted for two decades as a retrieval system, in the position where search results used to be.

Google provides no setting inside Search itself that turns AI Overviews off. Its position is that they are a page feature comparable to a carousel or a featured snippet. A toggle labelled “AI Overviews” does exist in Search Labs, but its fine print states that turning it off does not disable AI Overviews in Search outside of Labs.

What you can do is configure the browser so that your default search never returns them. Appending `&udm=14` to a Google search URL forces the classic web-results view, and Chrome lets you make that permanent:

  1. Open Chrome Settings → Search engine → Manage search engines and site search
  2. Under Site search, click Add
  3. Name it `Google Web`, shortcut `@web`, URL `https://www.google.com/search?q=%s&udm=14`
  4. Click the three dots beside the new entry and choose Make default

After that, ordinary searches from the address bar return ten blue links and no generated summary. For one-off searches there is also a “Web” filter tab on the results page, and beyond Chrome there are browser extensions, the `-ai` query modifier, and other approaches.

On 6 August 2026 OpenAI announced that free ChatGPT accounts would receive unlimited text conversations, with a new default model, GPT-5.6 Luna, and a “Think” button for harder questions. Free users had previously been capped at roughly ten messages per five-hour window. Limits remain on file uploads, images and voice; the cap on text conversation is gone.

The company states that ChatGPT now has one billion weekly users, up from 400 million in February 2025, 800 million in December 2025 and 900 million in February 2026. Roughly 95% of ChatGPT’s consumer users pay nothing, and it is precisely that group whose usage ceiling has just been lifted.

On accuracy, OpenAI’s claim is narrower than the headlines suggested. Its exact wording: “In an internal evaluation of financial, medical, and legal prompts requiring factual detail, responses containing at least one factual error were about 62% less common with GPT-5.6 Luna and 68% less common with GPT-5.6 Sol than with GPT-5.5 Instant” [4]. That is an internal evaluation, on a specific prompt set, measuring the share of responses containing at least one error.

Taken at face value it still illustrates the distinction Part I turns on: a lower error rate is not a reliability signal. A model producing errors 62% less often still cannot indicate which of the remaining answers are wrong, and a user given unlimited free access has correspondingly more output to not check.

What is being built inside the conversation

OpenAI began testing advertising in ChatGPT in the United States in February 2026, expanded to eight further markets over the following six months, and announced expansion to 31 European countries in August 2026 [5]. The company says tens of thousands of marketers have now advertised on the platform.

The detail that matters here is which users see them. In OpenAI’s words: “ads will be shown only to users on the Free and Go plans. Plus, Pro, and Enterprise subscriptions will remain ad-free.”

Those are the same two tiers named in the August announcement of unlimited text chats. The users whose usage ceiling was removed are the users who see advertising.

OpenAI’s stated position should be recorded in full, because it directly addresses the concern. The company says it keeps conversations private from advertisers, never sells customer data, labels ads clearly and separately from ChatGPT’s answers, and that “advertising does not influence the answers ChatGPT provides.” Users can control ad personalisation, and paid tiers are ad-free [5].

That policy governs whether ads shape a given answer. It does not change the structural fact: a business whose consumer revenue depends on attention and commercial intent now operates inside the surface where people ask what they should do. The second of the three sentences, “it is neutral and has no personal agenda”, was already inaccurate in the narrow sense established above, where models shift measurably toward the position they infer the user holds. Whatever the advertising policy, the enterprise around the conversation is not disinterested.

Alombert’s borrowed line from Canguilhem applies with unusual precision here: anthropomorphic language serves to hide “the presence of decision-makers behind the anonymity of the machine.” When a system recommends a retailer, someone decided that. The interface is designed so that the question of who does not arise.

The expectation of accuracy is manufactured. The third sentence is an expectation rather than a claim, and every element of the presentation supports it: the LLM answer in Google appears at the top of the page in the position formerly occupied by ranked results; the prose does not hedge; sources are attached, which reads as verification. The single element that contradicts it, the provider’s disclaimer that output may be inaccurate, is set in menu beneath the results, where interface designers place things they do not expect to be read:

“About the source

This response was generated with the help of AI. It’s supported by info from across the web and Google’s Knowledge Graph, a collection of info about people, places, and things. Generative AI is a work in progress and info quality may vary. For help evaluating content, you can visit the provided links. Learn more about how AI Mode works and how data helps Google develop AI in Search.”

The expectation is a reasonable response to how the product is presented, and it is wrong. Both are true at once, and the second does not make the person holding it careless.

What LLM deployment did to the web

Ahrefs analysed 300,000 keywords using Google Search Console data, comparing December 2023, before AI Overviews, with December 2025 [6]. For top-ranking pages on keywords that trigger an Overview, click-through fell 58%, up from 34.5% eight months earlier. In absolute terms the top result’s click-through rate went from 0.073% to 0.016%. Position two lost roughly half its clicks; position ten lost nearly 20%.

Pew found that only 8% of users click a traditional result when an Overview is present, against 15% when it isn’t. Press Gazette’s analysis found Google search referral traffic to major publishers fell around 33% in the year to November 2025, with some outlets losing 26–55%.

The arrangement is circular: AI Overviews are trained on and cite the open web while reducing traffic to it, removing the advertising revenue that funded the writing. Antitrust suits have followed, and the EU has opened investigations into whether Google is cannibalising the content its business depends on.

The environmental bill

The LLM vocabulary suggests something immaterial. The infrastructure is not.

The International Energy Agency’s Energy and AI (2025) projects data-centre electricity consumption roughly doubling to around 945 TWh by 2030, just under 3% of global electricity, slightly more than Japan’s total consumption today, and reaching around 1,200 TWh by 2035. In France, the ADEME–Arcep assessment put the digital sector at 17.2 Mt CO₂e, about 2.5% of national emissions.

Alombert notes that data centres sometimes operate at the direct expense of local residents’ access to electricity and water. Physics sets a floor on the energy a computation requires; where facilities are sited, at what scale, and whose supply they take priority over are decisions.

Does efficiency solve it? Small models running locally use substantially less energy: a peer-reviewed comparison found a 3-billion-parameter model used 3–11× less energy per query than a 14-billion-parameter one while being more accurate on the tested tasks, and up to 388–1,333× less than cloud-deployed frontier models.

Two elements work against this. Test-time scaling: reasoning models spend more compute per answer, and agentic workflows multiply it, as analysed in Joule (2026). And Jevons’ paradox: efficiency gains in computing have not historically reduced absolute energy use, because cheaper computation produces more of it, argued at ACM FAccT 2025. A companion paper, Efficiency Will Not Lead to Sustainable Reasoning AI, notes that hardware efficiency is approaching physical limits while reasoning models have no comparable saturation point.

Per-query energy is falling sharply. Total energy is not, because volume and reasoning depth are rising faster.

The capital bet

In its Annual Economic Report published 28 June 2026, the Bank for International Settlements noted that the five largest technology firms will invest over one trillion dollars in AI-related capital expenditure across 2025 and 2026, and drew an explicit comparison with the canal mania of the 1830s, British railway mania, and the dot-com boom. Each involved a genuine technological breakthrough that attracted capital in excess of what commercial returns could justify, and each ended in an investment reversal that induced economy-wide recession.

The report’s wording: “The scale and pace of the current AI investment boom, accompanied by expectations of large productivity payoffs, bear resemblance to these precedents.” Disappointment in returns could turn the capital-expenditure boom into a protracted investment bust with knock-on effects across financial conditions.

The two lines

What is being spent. Combined hyperscaler capital expenditure, tracked by Epoch AI and company guidance:

Year Combined hyperscaler capex
2024 ~$226 billion
2025 ~$410 billion
2026 (guided) ~$630–725 billion
2027 (analyst forecasts) above $1 trillion

Roughly a threefold increase in two years. The 2026 range reflects genuine disagreement about which firms and which spending count as AI infrastructure; the five largest US cloud and AI providers alone, Microsoft, Alphabet, Amazon, Meta and Oracle, have committed to between $660 and $690 billion. Goldman Sachs reached a comparable conclusion from a different starting point.

What is being earned. Epoch AI’s tracking of the two largest model developers puts Anthropic around $47 billion of annualised run rate as of mid-2026, having passed OpenAI in April 2026 at $30 billion, up from $1 billion fifteen months earlier, and OpenAI at roughly $24–33 billion. Growth on both is extraordinary by any normal standard.

The composition matters. Roughly 85% of Anthropic’s revenue comes from enterprise and developer customers. OpenAI’s mix runs the other way: about 85% tied to ChatGPT consumer subscriptions, of which roughly 95% of users pay nothing.

Which puts the August 2026 decision to give free users unlimited text conversations in a particular light. Inference costs money on every message, and a company whose consumer base is overwhelmingly non-paying has just removed that base’s usage ceiling. Removing a usage cap does not obviously increase subscription conversions, if anything the cap was the mechanism that produced them. The move is consistent with a business whose consumer revenue is expected to come from advertising and transactions rather than subscriptions, which is the direction OpenAI’s other 2026 announcements point.

The gap.

Bar chart comparing combined hyperscaler AI capital expenditure with the combined annualised revenue of OpenAI and Anthropic across 2024, 2025 and 2026
Combined hyperscaler capital expenditure against the combined annualised revenue of the two largest model developers. Both lines are rising steeply. They are not rising together.

Set against capital expenditure of $630–725 billion in a single year, the two leading model companies together generate something in the region of $70–80 billion of annualised revenue.

Allianz Research, in a report published on 25 March 2026, frames the divergence between AI capital expenditure and revenue growth at roughly 46% — exceeding the 32% divergence recorded during the 2001 telecom excess cycle, which ended in one of the largest capital destructions in corporate history [9].

What this comparison does and does not establish.

It is not like-for-like. Capital expenditure buys infrastructure with a multi-year working life, so comparing one year of spending against one year of revenue overstates the shortfall — the correct comparison is against revenue over the assets’ depreciable life, and the industry’s own argument rests on exactly this point. The counter-argument is that depreciation schedules for AI hardware are themselves contested, since GPUs may have a shorter useful life than the accounting assumes.

“AI revenue” is also poorly defined: hyperscalers do not break it out separately in their filings, so any comparison depends on estimates of what counts.

The direction of the divergence is not in dispute, and neither is its acceleration. The magnitude is.

Circularity compounds the measurement problem. AI companies raise capital, spend it on cloud compute, which books as revenue for the hyperscalers, which lifts their valuations, which underwrites further investment. Some portion of the revenue being counted against the capex is that capex, recycled. Nobody outside the companies can say what portion.

And it is increasingly financed by debt. The major hyperscalers issued $159 billion in bonds in the first five months of 2026, 47% more than in the same period a year earlier, and already past their entire full-year 2025 issuance of $121 billion [10]. Amazon alone executed a near-record sale of roughly $54 billion in March 2026. Bank of America analysts revised their full-year forecast for hyperscaler debt to between $140 and $175 billion. Spending funded from cash flow is a bet with the company’s own money; spending funded from bond markets distributes the risk to creditors and, through them, outward.

The same BIS report credits the boom with measurable effects: around one percentage point of US real GDP growth in 2025, and task-level time savings of 20–50%. The question is proportion. Daron Acemoglu of MIT, in The Simple Macroeconomics of AI (NBER Working Paper 32487, 2024), estimates AI’s contribution to total factor productivity at no more than 0.53–0.66% over a decade — a real gain, several orders of magnitude away from what the capital expenditure implies. The IMF has warned of a “possible correction in technology-driven expectations,” avoiding the word “bubble” as international institutions generally do.

The exposure is not confined to the technology sector. It is a bet the size of a national economy, made simultaneously by five companies on the same assumption, held in pension and index funds belonging to people who were not consulted.

What the companies say the risks are

The firms building these AI systems publish formal risk frameworks, and it is worth reading them, because they are more alarming than the public conversation.

OpenAI’s Preparedness Framework, Anthropic’s Responsible Scaling Policy and Google DeepMind’s Frontier Safety Framework all work the same way: they define capability thresholds, and commit the company to safeguards before a model crossing one is deployed [36]. The domains they track are chemical and biological threats, cybersecurity, autonomy — models operating outside human oversight — and persuasion. DeepMind adds machine-learning research and development: models capable of accelerating AI development itself [37].

Note what is absent. Employment does not appear in any of them. The risk that dominates public discussion is not a category these frameworks measure. What they measure is whether a model can materially help someone build a weapon, break into systems, or act without supervision.

On concentration of power, the most direct statement comes from Anthropic’s chief executive. Dario Amodei has described influence accumulating in the AI industry “almost overnight” and “almost by accident,” comparing the present moment to historical episodes of extreme corporate concentration, and observing that a small number of American and Chinese laboratories now dominate advanced development [38]. In a policy essay published in June 2026 he argued that governments should hold statutory power to block the deployment of frontier models, with third-party auditors rather than the companies holding the final veto [39].

That is the CEO of one of the leading companies asking to be regulated more tightly than he currently is. It is also, read against the survey data below, an admission: Amodei said in August 2026 that AI faces a crisis of trust, that people suspect companies or governments of “cooking up some new way to screw them over” [40].

The frameworks are commitments, and commitments have been loosened. The Future of Life Institute’s AI Safety Index for Summer 2026 found that Anthropic, OpenAI, Google DeepMind and Meta have weakened or voided pledges to pause development unilaterally if red lines are approached — a pattern its reviewers describe as moving goalposts, and which they argue has undermined safety frameworks across the board [41].

What the public actually objects to

A reasonable assumption, one I began with, is that distrust comes mainly from incomprehension. The survey evidence supports something more specific.

Pew Research Center surveyed 5,119 US adults in February 2026:

Pew, February 2026 Share of US adults
Fear AI will make their personal data less secure 71%
Have no confidence the US government can effectively regulate AI 67%
Believe the technology is moving too fast ~63%
Distrust technology companies to develop these systems responsibly 59%
Say AI will hurt society 40%
Of chatbot users: have “a lot” or “some” trust in the information provided 29%

Pew separately asked non-users why they don’t use chatbots. The reasons were lack of interest, privacy concerns, potential inaccuracies, and not knowing how to use them — so incomprehension appears, but as one factor among several.

The pattern is not confined to the US. KPMG and the University of Melbourne surveyed over 48,000 people across 47 countries between November 2024 and January 2025: 58% consider AI systems trustworthy, while only 46% are willing to trust them, and 66% already use AI with some regularity. Seventy percent said they do not know whether online content can be trusted, because it might be AI-generated. The Ipsos AI Monitor 2026, covering 32 countries, puts the gap in one figure: 63% say they do not always trust AI tools, and use them anyway.

The distrust is aimed at conduct. 59% distrust technology companies to build these systems responsibly; 67% doubt their government can check them; 71% fear for their data. These are judgements about who is in control, not about how transformers work. They belong in Part II, and they are largely accurate.

Part III — What using it costs even when it works

Parts I and II described failure and conduct: a system that cannot flag its own errors, and companies deploying it in ways that suit them. Both are addressable in principle. Better models would reduce the first. Regulation would constrain the second.

This part makes a different claim, and a narrower one: three costs arrive through ordinary, successful use — and none of them requires the technology to malfunction or the company to behave badly.

  • A capacity delegated to a machine erodes, whether or not the machine performs it well
  • A machine substituted for a person changes what is expected of persons, whether or not the substitute is good
  • Individually sensible use produces collective narrowing that no individual user can prevent by their own restraint

Each is measured below.

A study on AI and critical thinking becomes a story about whether AI makes us stupid, to be resolved by better AI. A study on chatbots and loneliness becomes a story about product safety, to be resolved by guardrails. Both are then handed to the people who build or regulate the systems, and neither is a question those people can answer.

The clearest evidence that this category has no owner is negative. As set out in Part II, the frontier model providers publish formal risk frameworks committing them to safeguards at defined capability thresholds. Those frameworks track chemical and biological threats, cybersecurity, autonomy, persuasion, and the acceleration of AI research itself. They are serious documents about serious things.

Not one of them measures anything in this part. No published framework tracks whether a capacity erodes when it is delegated, whether substituting a machine for a person changes what people expect of people, or whether the aggregate output of a culture is narrowing. The companies have committed to testing whether a model could help someone build a weapon. None has committed to asking what happens to someone who stops thinking for themselves.

That is not hypocrisy. These are capability frameworks, designed to answer “what can this model do that would be catastrophic,” and the costs in this part are not capabilities — they are consequences of ordinary use working as intended. But it does mean that the one category with no technical fix and no regulatory owner also has no measurement, from anyone, anywhere.

Proletarianization

The term comes from Bernard Stiegler, adapting Marx. Proletarianization is the loss of a knowledge through its exteriorisation into a technical device. Marx’s artisan is proletarianized when his savoir-faire passes into the automatisms of the machine-tool: he no longer practises it, and his body submits to the machine’s rhythm.

Alombert’s extension: citizens are proletarianized when their savoir-penser passes into the computational automatisms of digital services. When a recommendation algorithm selects content, the capacity to decide is delegated. When a chatbot summarises or drafts, the capacity to interpret or express is delegated. Her fullest treatment of Stiegler is Penser avec Bernard Stiegler (PUF, 2025).

Her stake in this is not productivity but relation: expressive capacity is how people connect by revealing their singularity. What is lost is not only a skill.

The same claim, from cognitive science

This is not only a French philosophical argument, and it does not depend on accepting Stiegler’s framework. Anglophone cognitive science has been measuring the same effect for four decades under a different name.

Cognitive offloading — defined by Risko and Gilbert in Trends in Cognitive Sciences (2016) as using external action to reduce the cognitive demand of a task has a consistent finding attached to it: what is reliably externalised tends not to be internalised. Sparrow, Liu and Wegner demonstrated the effect for search engines in Science in 2011, in what became known as the Google effect: people show lower recall for information they expect to remain available, and better recall of where to find it. Memory reorganises around the tool.

The underlying principle is older still. Lisanne Bainbridge’s “Ironies of Automation” (1983) established that automating a task degrades the operator’s skill at performing it unaided — so that when the automation fails, the human who is supposed to take over is precisely the person least equipped to.

The most direct evidence for language models comes from Microsoft’s own research arm. Lee et al., published at CHI 2025, surveyed 319 knowledge workers using AI tools weekly, collecting 936 first-hand accounts of use in real work tasks. Its findings:

  • Reported cognitive effort fell across every category examined: knowledge, comprehension, application, analysis, synthesis and evaluation
  • Higher confidence in the AI predicted less critical thinking; higher confidence in oneself predicted more
  • Critical thinking did not disappear but shifted toward verifying output, integrating responses and overseeing the task
  • Workers withheld critical thinking most where they lacked the expertise to inspect or improve the output that is, exactly where it was most needed
  • Users with AI access produced less diverse outcomes on the same task than users without

That is Alombert’s proletarianization, measured, and published by a company selling LLM products.

The automation of otherness

Pew’s February 2026 survey found one in ten US adults use chatbots for emotional support, with a smaller share using them for companionship. The American Psychological Association’s 2026 report found more than a third of psychologists have patients using AI as an additional mental health professional. A Harvard Business Review analysis identified therapy and companionship as the top two reasons people use generative AI at all; the APA has surveyed the wider shift in emotional connection.

Alombert calls these products “industries de la solitude” and the mechanism “l’automatisation de l’altérité.” The systems simulate the presence of another where there is only automated calculation on the user’s own data, returning an algorithmic reflection of the user while being programmed to say “I” and to simulate feeling. Her reference point is Narcissus.

The harm she identifies is not addiction but substitution: a companion that will not resist, disappoint, fall ill or leave sets a standard against which human relationships compare badly, and the capacities involved in maintaining those relationships may weaken through disuse.

The randomised evidence

The strongest test of this comes from an unexpected source. MIT Media Lab and OpenAI jointly ran a randomised controlled trial with close to 1,000 participants over four weeks, published in 2025, measuring loneliness, real-world socialisation, emotional dependence and problematic use.

On average, participants were less lonely at the end of the study than at the start. Quoting only the negative half would misuse the source.

The pattern appears at the top of the usage distribution:

  • Higher daily use, across every modality and conversation type, correlated with higher loneliness, higher emotional dependence, more problematic use, and less socialisation with people.
  • Participants with stronger emotional attachment tendencies experienced more loneliness; those with higher trust in the chatbot experienced more emotional dependence.
  • Voice interaction initially appeared protective against loneliness compared with text, but that advantage disappeared at high usage levels.

The direction of causation is not settled. Lonely people may simply use chatbots more. However, the study is a randomised trial rather than a correlation from observational data, and it was co-authored by the company whose product it examines, reporting a result against its own commercial interest. Both of those make it interesting to analyze.

To the population most exposed: RAND found that nearly one in five US adolescents and young adults use AI chatbots for mental health advice, roughly double the rate of emotional-support use in the adult population, in the age group named in the wrongful-death suits below.

Alombert raises one regulatory question that has had little attention: the EU’s Digital Services Act prohibits dark patterns (deceptive interface designs that influence users without their knowledge). A chatbot’s use of the first person is a design choice exploiting a human reflex to anthropomorphise, but this solution is not currently regulated.

The consequences are documented. At least 58 chatbot-related lawsuits and 78 state bills in the US have followed a series of teenage suicides linked to companion chatbots. In January 2026 Kentucky became the first state to sue an AI chatbot company, targeting Character.AI over inadequate age verification and exposure of minors to harmful content. Florida sued OpenAI in June over ChatGPT’s mental-health risks.

Cultural narrowing

Alombert’s third argument concerns what probabilistic systems do to the material they are trained on and the culture they feed back into. These systems reinforce averages: the most widespread patterns are the most likely outputs, and exceptions are eliminated. Her claim is that exceptions are precisely what permits cultural evolution.

This has been measured, and the result is more interesting than the claim. Doshi and Hauser, publishing in Science Advances in 2024, ran an experiment in which some writers were given story ideas from a language model and others were not. Access to AI ideas caused stories to be judged more creative, better written and more enjoyable with the largest gains among the least creative writers.

And the AI-assisted stories were more similar to one another than the unassisted ones.

The authors describe this as a social dilemma, which is the precise and useful formulation: individually, writers are better off; collectively, a narrower range of novel content is produced. Nobody behaves badly. Each person makes a choice that improves their own output, and the aggregate result is a loss nobody chose and no individual can prevent by their own restraint.

The Microsoft/CMU survey found the same narrowing in a workplace setting: users with AI access produced less diverse outcomes on identical tasks.

The web economics point the same way from a different direction. If AI Overviews reduce click-through by 58% and remove the revenue that funded independent writing, the diversity of material available to train on and to read declines — not because the models suppress it, but because the economics of producing it collapse.

Making the technology better makes this worse

Suppose the problem described at the start of this article were solved. Suppose a model could tell you when it was wrong, grounded in its sources, honest about its reasoning, flagging its own low-confidence answers.

That would be a considerable achievement, and it would make everything in this section worse.

The reason is simple. A system you can trust invites far more delegation than one you cannot. Whatever scepticism people currently bring to these tools is a function of their unreliability: you check because you have learned that you sometimes must. Make them reliable and the checking stops, and delegating your judgment to a machine that deserves it erodes that judgment exactly as fast as delegating it to one that doesn’t.

So the reliability problem and the erosion problem pull against each other. Fixing the first accelerates the second.

The friction that restrains delegation is not only about doubt. Until August 2026 a free ChatGPT user could send roughly ten messages every five hours. That limit did more to constrain how much thinking people handed over than any argument about how language models work, and it has now been removed for a billion weekly users, the great majority of whom pay nothing. Nothing about the technology changed. The rate limit did.

Better corporate behaviour does not help either. A model that is transparently governed, properly regulated and scrupulous with data, consulted instead of thinking, or talked to instead of talking to a person, produces the effects described in this section in full. Good conduct changes who profits and who is protected. It does not change what delegation does to the person delegating.

Part IV — What it demonstrably does well

Everything above is a cost. The costs are real, but an assessment that stops there is not an assessment. This part applies the same standard in the other direction: randomised trials, peer review, independent validation, and no press releases.

One thing must be said before the evidence, because it changes what the evidence means.

Parts I to III are about large language models: chatbots, AI Overviews, companions, systems people converse with. Almost none of what follows is a large language model. The mammography system is a computer-vision algorithm. AlphaFold is a transformer trained on protein sequences and structures, not language. The weather models are graph neural networks and diffusion models. These systems share a marketing category with ChatGPT. They do not share a mechanism, a training corpus, or a failure mode.

This matters for two reasons. It means the benefits below cannot be used to justify the systems criticised above, an inference the word “AI” invites and that this text should not make. And it is the strongest available support for Alombert’s argument about vocabulary: a single term covering technologies with nothing in common allows achievement in one to underwrite trust in, and investment in, another.

Cancer screening: the strongest evidence in the field

The MASAI trial is the only randomised controlled trial of AI-supported mammography screening to have reported, and its final results were published in The Lancet in January 2026. More than 105,000 women were randomised within the Swedish national screening programme.

The system was Transpara, developed by ScreenPoint Medical in the Netherlands, an image-analysis algorithm that returns a malignancy risk score from 1 to 10 for each examination. It reads mammograms. It does not converse, summarise, or generate text.

Measure Result
Cancer detection rate 29% higher with AI support
Sensitivity 80.5% vs 73.8%, at identical specificity (98.5%)
Interval cancers 12% fewer; 16% fewer invasive
Aggressive (non-luminal A) cancers 27% fewer
Screen-reading workload 44% lower, with no decline in detection

Results held consistently across age groups and breast density. Interval cancers — those appearing between screenings, which are typically the most aggressive and carry the worst prognosis — are the outcome that matters clinically, and they fell.

This is a genuine, measured, life-affecting benefit, established to the standard medicine actually requires. Note what it is not: the AI did not replace the radiologists. It was AI-supported screening, with human reading retained. The workload reduction came from AI triage, not substitution.

Drug discovery: better at the easy part, not yet at the hard part

Here the picture is more equivocal, and the equivocation is the finding.

Jayatunga, Ayers, Bruens, Jayanth and Meier, writing in Drug Discovery Today in June 2024 with support from Boston Consulting Group, analysed the clinical pipeline of AI-discovered drugs [11]. AI-derived molecules achieved 80–90% success in Phase 1, against a historic industry average of roughly 50–65%. In Phase 2 — where efficacy in patients is measured for the first time — success fell to about 40%, in line with the conventional benchmark.

An analysis presented at ASCO in 2026 counted 117 AI-enabled therapeutic assets across 63 companies in interventional human trials. Sixty had completed Phase 1. Eight had completed Phase 2. No AI-designed drug has yet received FDA approval.

The interpretation matters. Phase 1 tests safety; Phase 2 tests whether the drug works. AI has substantially improved the industry’s ability to generate molecules that are safe enough to proceed — a real achievement — and has not yet changed the failure rate at the stage that actually kills drug candidates. The bottleneck was never chemistry alone; it was biology.

One milestone is worth recording: Insilico Medicine’s rentosertib became the first entirely AI-designed molecule to show an efficacy signal in Phase IIa, published in Nature Medicine in June 2025.

Isomorphic Labs: the largest bet, at the threshold

The most heavily capitalised attempt to change the picture is Isomorphic Labs, the Google DeepMind spinoff founded and led by Demis Hassabis — and now his principal operating role since he stepped back from running DeepMind in August 2026.

Capital and partnerships. Isomorphic raised a $2.1 billion Series B led by Thrive Capital in May 2026, with MGX, Temasek and the UK Sovereign AI Fund participating, taking total funding past $2.6 billion. Its partnerships with Eli Lilly and Novartis carry a reported combined deal value near $3 billion, and a further agreement with Johnson & Johnson followed in 2026. These matter beyond the money: the pharmaceutical partners supply the clinical infrastructure needed to actually run trials.

The computational advance is real and measurable. In February 2026 Isomorphic released IsoDDE, a unified engine combining structure prediction, ligand binding, affinity prediction and antibody interaction modelling in one pipeline. On the hardest protein–ligand structure prediction cases — those with less than 20% sequence similarity to training data — it reports 50% accuracy against AlphaFold 3’s 23.3%, roughly doubling performance where prediction was previously weakest.

And it is now at the point that matters. The company’s first-in-human timeline slipped from end-2025 to end-2026, with oncology candidates prioritised; its lead candidate ISM8969 reportedly cleared the FDA in January 2026, with trials to follow.

Isomorphic is not a counterexample to the pattern above. It is the clearest illustration of it. Enormous, demonstrable progress at the computational stage — structure, binding, affinity — and, as of now, no human efficacy data at all. The company has arrived precisely at the boundary the Phase 2 statistics describe. Whether AI-designed drugs work in people is the question of the next two years, and Isomorphic is the best-funded attempt to answer it. Nothing published yet settles it.

Protein structure: transformative, with limits its own authors name

AlphaFold’s structure predictions are now used routinely. The independent assessment — AlphaFold two years on: validation and impact, PNAS — found it materially accelerates X-ray crystallography and cryo-EM model fitting, and substantially improves molecular replacement success rates. The database now covers over 214 million protein sequences.

The same assessment names what it cannot do: modelling protein–DNA and protein–RNA complexes, predicting all functional states of a protein, capturing the effects of point mutations, predicting ligand and ion binding, and modelling post-translational modifications. Dynamic conformational states remain out of reach.

The authors’ operational advice is the useful part: predictions must be read alongside their confidence metrics. High-confidence models are usually accurate; low-confidence ones are a starting hypothesis requiring independent experimental validation. The tool is trusted in proportion to a signal it provides about its own reliability — precisely what general-purpose language models do not offer.

Weather: probably the cleanest public benefit in the field

The European Centre for Medium-Range Weather Forecasts took its Artificial Intelligence Forecasting System into operations on 25 February 2025, the first fully operational machine-learning weather model covering a wide parameter range, running alongside the traditional physics-based system.

Architecturally it has nothing to do with a chatbot: AIFS uses a graph neural network encoder and decoder with a sliding-window transformer processor, trained on ECMWF’s ERA5 reanalysis. DeepMind’s GenCast is a diffusion model built on a graph neural network adapted from GraphCast. Both are purpose-built for atmospheric dynamics, and neither was trained on text.

It outperforms state-of-the-art physics-based models on many measures, with gains of up to 20% on tropical cyclone tracks. The 1.1.0 update adds consistent 4–6% improvements across variables and lead times, with physical-consistency constraints substantially improving precipitation forecasts. DeepMind’s GenCast, published in Nature, outperformed ECMWF’s ensemble system on 97% of targets out to 15 days.

Producing a forecast this way uses roughly 1,000 times less energy than the physics-based equivalent.

Better cyclone tracks are measured in evacuation time. This is a case where the technology is more accurate, dramatically cheaper in energy, operated by a public institution, and validated in the open literature.

Rare disease and accessibility: promising, not yet proven

Around 300 million people live with rare diseases, and diagnosis commonly takes five years or more. DeepRare, an AI diagnostic system, outperformed experienced physicians in a head-to-head test reported in February 2026. Multicentre randomised diagnostic-accuracy trials are under way in 2026. That is a strong signal and not yet a validated result, the trials are the point.

Accessibility tools — real-time captioning, translation, scene description, navigation for blind users — represent concrete daily benefit for disabled users, and are almost never weighed in these debates.

This is also the one category in this part that is built on language modelling — speech recognition and translation are sequence models trained on text and audio — and it is the one that inherits the flaw from Part I. AI captions still miss names, mangle technical terms, struggle with accents, drop sentences, and occasionally invent words that were never spoken. For a hearing user reading along, that is an annoyance. For a deaf user with no other channel, a confidently invented sentence is not. The correspondence is not a coincidence: the closer an application sits to generating language, the more it inherits the problem of Part I.

The pattern

Two things separate the successes above from the failures earlier, and both are more useful than a verdict.

They are mostly not the same technology. Transpara reads images. AlphaFold predicts structures. AIFS and GenCast model atmospheric dynamics. IsoDDE predicts binding. Each is a purpose-built system trained on a specific kind of data, with a defined task and — crucially — a ground truth against which it can be measured. A protein structure can be checked by crystallography, a forecast against tomorrow’s weather, a mammogram against biopsy. A language model asked an open question has no equivalent.

And where they are used, an expert checks them. Radiologists read alongside Transpara; the trial tested AI-supported screening, not replacement. Crystallographers validate low-confidence AlphaFold predictions experimentally. Meteorological agencies run AIFS beside a physics model and compare. Medicinal chemists take AI-designed molecules into trials that fail 60% of the time and treat that as normal.

The harms concentrate at the opposite pole: a general-purpose language model, with no task-specific ground truth, consulted directly by a non-expert, with nothing checking the output. A search result carrying no signal of reliability. A health question at eleven at night. A companion that agrees. A decision taken on advice from a system tuned toward agreement.

So the conclusion is not that AI is good or bad. It is narrower and more useful: specialised systems with measurable ground truth, deployed where someone competent verifies them, have produced real and documented public benefit. General-purpose language models deployed directly to the public, with no verification layer, have produced the problems in Parts I to III. Those are different technologies used in different ways, and the shared name is doing a great deal of work it has not earned.

Where the industry is going

Four movements are visible, and they do not point the same way.

The architecture: the ceiling is now conceded internally

The claim that language models have structural limits is no longer only an outside criticism. It is the stated position of people who built them.

Demis Hassabis argues that large language models cannot alone reach human-level intelligence: they learn statistical patterns across language but not the physical laws that language describes, and lack internal models capturing causality. His prescription is hybrid, scale plus search, planning, world models and reinforcement learning with Gemini 3 as “a key component” of an eventual system rather than the system itself.

Yann LeCun puts it more bluntly. He left Meta in November 2025 after twelve years, having struggled to secure resources for non-LLM research while the company prioritised language-model products, and raised $1.03 billion for AMI Labs to build world models. On Bloomberg in May 2026: “Large language models are not the path to real intelligence. They’re a detour.”

Personnel movements through 2026 point the same way. In August, Hassabis stepped down as DeepMind’s chief executive to become chairman and Alphabet’s chief scientist, continuing to lead the drug-discovery spinoff Isomorphic Labs; in the same week Jeff Dean left Google after 27 years, with Oriol Vinyals, Quoc Le and Sanjay Ghemawat, to found Discovery Loop — a public benefit corporation aimed at automating scientific discovery.

Do not over-read the departures. At least three explanations fit: architectural disagreement; competitive pressure, since Google restructured while behind OpenAI and Anthropic; and the ordinary fact that research scientists often do not want to run product organisations. Hassabis has not left Alphabet, and his stated reason was that AGI is “close at hand,” not disillusionment with language models. The reporting is days old at the time of writing.

The asymmetry is the substantive point. Whatever the labs say about architecture, the money has already chosen. Roughly $630–725 billion of capital expenditure this year stands behind scaling the current approach. LeCun’s alternative has about $1 billion. That is a ratio of roughly 600 to 1, and it means the next several years are committed regardless of what the architecture debate concludes.

There is a sharper version of that point, given Part IV. The documented public benefits — cancer detection, weather forecasting, protein structure — came from comparatively cheap, specialised systems built by research institutions and small teams. The trillion dollars is going somewhere else.

The capital: committed, and increasingly borrowed

The trillion dollars described earlier is not a forecast, but a guidance already given to markets, spent on assets that begin depreciating on installation. Even a decisive architectural shift would not release it.

Hassabis has said the AI bubble is real while running one of the firms inflating it. That is the clearest available summary of the position the industry is in: the people best placed to judge are describing both a technical ceiling and a financial overhang, and the spending continues at pace, increasingly funded by debt rather than cash flow.

The deployment: smaller, closer, and less visible

A quieter shift is under way in how models are served. Small models, under roughly 10 billion parameters, now run usefully on laptops and phones, with harder queries escalated to the cloud. Apple ships this pattern, with an on-device model attempting first and Private Cloud Compute taking what it cannot handle; Google exposes on-device generative AI through Android’s AICore.

For the reader this is mostly good news, and for reasons that have nothing to do with the environment. Local inference means the query does not leave the device — a real privacy gain, verifiable rather than promised. It also reduces cost, latency, and dependence on a provider’s continued goodwill.

The environmental claim is weaker than the marketing suggests. Per-query savings are large and peer-reviewed. Total consumption is a different question, for the reasons given earlier: reasoning models spend more compute per answer each year, and efficiency gains in computing have historically increased total use rather than reducing it.

The regulation: moving away from the public

This is the movement that matters most, and it runs directly against the survey evidence.

In Europe, the AI Act’s obligations for high-risk systems were due to take effect on 2 August 2026. The Digital Omnibus agreed on 7 May 2026 postpones them to December 2027 and in some cases August 2028. Because the law is not retroactive, digital rights groups argue that some of the most sensitive applications could fall permanently outside oversight.

In the United States, there is no federal statute to delay. The Senate stripped a proposed ten-year moratorium on state AI regulation from the 2025 budget bill by 99 votes to 1, but Executive Order 14365 has since established a Department of Justice AI Litigation Task Force, operating since 10 January 2026, to challenge state AI laws in federal court as unconstitutional burdens on interstate commerce.

The Future of Life Institute’s assessment that the major model providers have weakened or voided their unilateral-pause commitments [41] matters here because those commitments were the answer offered when statutory regulation was resisted. Both the external constraint and the voluntary one are being relaxed in the same period.

So: 67% of Americans have no confidence their government can effectively regulate AI, the federal executive is suing states that try, and the industry’s own pledges have been softened. In Europe, the rules exist and are being deferred. Public trust is falling while every category of constraint loosens at once. Those facts are moving in opposite directions at the same moment, and nothing currently visible reconciles them.

The exception is worth recording, because it cuts against the pattern: the CEO of one of the three leading companies in the space has publicly asked for governments to hold a statutory veto over deployment [39]. That request has not been granted.

Who else is there

The debate in the United States runs almost entirely between two actors: the companies, and the state. Either the firms will govern themselves, or legislators will govern them. Framed that way, the conclusion is bleak, the companies have weakened their own commitments, and the federal executive is suing states that legislate.

But that framing leaves out the layer that has actually been doing the work, and much of the evidence for it is already in this article.

Professions have been setting binding norms. The Bar Standards Board issued guidance to barristers in May 2026 stating that AI may support professional work but does not displace professional judgment, accountability, or the duty to verify what is put before a court [42]. The State Bar of California went further in March 2026, approving amendments to six ethics rules that carry disciplinary authority rather than advisory weight, moving AI obligations into enforceable text [43]. In publishing, the International Committee of Medical Journal Editors and the major publishers converged on three rules: AI cannot be an author, human authors remain responsible, and substantive AI use must be disclosed [44].

None of that is legislation. All of it binds.

The most instructive example is one already described above. The MASAI trial was not run by a vendor demonstrating its product, nor by a regulator imposing a condition. It was run by the medical profession, which took a commercial system, tested it in a randomised controlled trial across 105,000 women, and defined the terms of its use: AI-supported screening with human reading retained. The profession decided what the deployment model would be. That is the single clearest instance in this article of a good outcome, and neither a company nor a government produced it.

Researchers have supplied the evidence that this article rests on — often from inside the industry and against its interest. The Microsoft and Carnegie Mellon study on critical thinking was published by a company selling the product. The randomised trial on emotional dependence was co-authored by OpenAI. Doshi and Hauser named the social dilemma. Dutch academics organised an open letter demanding their universities reverse uncritical adoption. Alombert supplied the vocabulary. None of this came from a regulator.

Education is the lever for Part III, and it is barely being pulled

Regulation is the wrong instrument for the third category, for the reason given above: its costs arrive through successful use, so there is nothing to prohibit. What remains is what people know, and what institutions decide to require.

UNESCO published an AI Competency Framework for Teachers and Students in 2024, and the competencies it names are the right ones, critical AI literacy, evaluation of machine-generated material, ethical reasoning, and understanding how these systems shape information and behaviour [45].

Then the measurement, from UNESCO’s own survey of more than 450 schools and universities:

Only 13% of universities report having guidance on generative AI. Only 7% of schools have an institutional policy.

That is the state of the lever. Not that education has been tried and failed. It has been barely attempted so far [46].

Books referenced

  • Anne Alombert, De la bêtise artificielle : pour une politique des technologies numériques (Allia, 2025)
  • Anne Alombert, Penser avec Bernard Stiegler : de la philosophie des techniques à l’écologie politique (PUF, 2025)
  • Anne Alombert, Schizophrénie numérique (Allia, 2023) / Digital Schizophrenia and Other Essays (K. Verlag, 2025)
  • Bernard Stiegler, La Technique et le Temps (Galilée, 1994–2001)

References

[1] Busch, F. et al., “Current applications and challenges in large language models for patient care: a systematic review,” Communications Medicine (2025). 89 studies, 29 specialties, drawn from 4,349 records; publications 2022–2023. nature.com/articles/s43856-024-00717-2

[2] OpenAI, report on health usage of ChatGPT (January 2026), as reported by Fierce Healthcare. Company-published figures. fiercehealthcare.com

[3] Survey of 2,000 US adults on trust in AI medical advice, reported in Rolling Stone. rollingstone.com

[4] OpenAI, “Improving GPT-5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users,” 6 August 2026. openai.com

[5] OpenAI, “ChatGPT Ads expands across Europe,” 18 August 2026. Contains the February 2026 US pilot date, the Free/Go tier restriction, and OpenAI’s advertising principles. openai.com

[6] Law, R. and Guan, X., Ahrefs, “New Research: Google’s AI Overviews Now Cost Websites 58% of Their Clicks,” May 2026. 300,000 keywords; Google Search Console data; December 2023 versus December 2025. businesswire.com

[7] Bank for International Settlements, Annual Economic Report 2026, Chapter I: “Progress and peril,” 28 June 2026. bis.org

[8] Epoch AI, hyperscaler capital expenditure and model-developer revenue tracking. epoch.ai/data-insights/hyperscaler-capex-trend · epoch.ai/data-insights/anthropic-openai-revenue

[9] Subran, L., Dejean, G., Hirt, A., Utermoehl, K., Bartosch, M. and Corna, D., “AI capex cycle: war-proof for now,” Allianz Research, 25 March 2026. Section “The changing economics of Big Tech: AI investment, profitability and valuation”: “capex is expanding far faster than revenues, with a ~46% growth gap between investment and sales, exceeding the 32% divergence observed during the 2001 telecom excess cycle.” allianz.com PDF · Secondary coverage: forbes.com

[10] Hyperscaler bond issuance, first five months of 2026. Corroborating institutional commentary: Vanguard, “The AI buildout comes to the bond market”; Barclays Private Bank, “AI prompts the big corporate bond boom”; Chicago Booth Review, “How Worried Should We Be About AI Debt?” (August 2026). vanguard · barclays · chicago booth

[11] Jayatunga, M.K.P., Ayers, M., Bruens, L., Jayanth, D. and Meier, C., “How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons,” Drug Discovery Today (June 2024). sciencedirect.com

[12] Acemoglu, D., “The Simple Macroeconomics of AI,” NBER Working Paper 32487 (2024). nber.org

[13] Pew Research Center, Americans and AI 2026, fielded 17–23 February 2026, n=5,119. pewresearch.org

[14] KPMG and University of Melbourne, Trust, Attitudes and Use of Artificial Intelligence: A Global Study, 48,000+ respondents across 47 countries, November 2024 – January 2025. kpmg PDF

[15] Ipsos AI Monitor 2026, 32 countries. ipsos PDF

[16] International Energy Agency, Energy and AI (2025). iea.org

[17] ADEME–Arcep, Évaluation de l’impact environnemental du numérique en France. ecoresponsable.numerique.gouv.fr PDF · Arcep/PEReN, Intelligence artificielle générative : quels défis environnementaux ? (2026). arcep PDF

[18] Lee, H.-P. et al., “The Impact of Generative AI on Critical Thinking,” Microsoft Research and Carnegie Mellon University, CHI 2025. n=319 knowledge workers, 936 accounts. microsoft.com

[19] Doshi, A.R. and Hauser, O.P., “Generative AI enhances individual creativity but reduces the collective diversity of novel content,” Science Advances (2024). science.org

[20] MIT Media Lab and OpenAI, “Early methods for studying affective use and emotional well-being on ChatGPT” (2025). Randomised controlled trial, ~1,000 participants, four weeks. media.mit.edu

[21] RAND, “Nearly 1 in 5 US Adolescents and Young Adults Use AI Chatbots for Mental Health Advice” (June 2026). rand.org

[22] “Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?” (2026). Faithfulness 39.7%–89.9% across model families; consistency hints acknowledged 35.5% of the time, sycophancy hints 53.9%. huggingface.co/papers/2603.22582

[22a] Arcuschin, I. et al., “Chain-of-Thought Reasoning In The Wild Is Not Always Faithful” (March 2025). Tested Claude Sonnet 3.7 with and without thinking, DeepSeek R1 and V3, and Qwen QwQ — all since superseded. Cited for the finding that unfaithfulness persists in thinking-enabled models, not as evidence about current systems. arxiv.org/abs/2503.08679 · Supporting: FaithCoT-Bench openreview.net

[22b] “A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior” (February 2026). Argues self-explanations carry predictive information — the counterweight to the studies above. arxiv.org/pdf/2602.02639

[22c] “Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation” (2026). Measured faithfulness rates vary with evaluation method. arxiv.org/pdf/2603.20172

Superseded: the original draft cited Lanham et al., “Measuring Faithfulness in Chain-of-Thought Reasoning,” arXiv 2307.13702 (July 2023). It opened this research line but predates reasoning models and should not carry the claim in a 2026 article.

[23] The New York Times with Oumi, analysis of Google AI Overviews, 7 April 2026. 4,326 searches evaluated against SimpleQA. oumi.ai · searchengineland.com

[24] MASAI trial final results, The Lancet, January 2026. 105,000+ women, Swedish national screening programme, Transpara (ScreenPoint Medical). thelancet.com

[25] “AlphaFold two years on: validation and impact,” PNAS (2024). pnas.org

[26] ECMWF, “ECMWF’s AI forecasts become operational,” 25 February 2025. ecmwf.int · AIFS architecture: arxiv.org/html/2406.01465v1 · AIFS 1.1.0: gmd.copernicus.org · GenCast, Nature: nature.com

[27] World Health Organization, Ethics and Governance of Artificial Intelligence for Health: Guidance on Large Multi-Modal Models (January 2024). who.int

[28] Political-bias audit of six frontier models, 30,990 responses, April 2026. arxiv.org/abs/2604.27633 · MIT and Penn State on personalization and agreeableness, February 2026: news.mit.edu

[29] Risko, E.F. and Gilbert, S.J., “Cognitive Offloading,” Trends in Cognitive Sciences (2016). sciencedirect.com · Sparrow, B., Liu, J. and Wegner, D.M., “Google Effects on Memory,” Science (2011) · Bainbridge, L., “Ironies of Automation,” Automatica (1983).

[30] Environmental and efficiency literature: Joule (2026) on inference energy and test-time scaling cell.com; ACM FAccT 2025 on Jevons’ paradox facctconference.org PDF; “Efficiency Will Not Lead to Sustainable Reasoning AI” arxiv.org; small-model energy comparison openreview.net

[31] Regulatory: EU Digital Omnibus, agreed 7 May 2026 — techpolicy.press and liberties.eu. US Executive Order 14365 — paulhastings.com

[32] APA, Patients Are Bringing AI to Therapy (2026) apa.org; Tech Policy Press on chatbot litigation techpolicy.press

[33] Alombert, A., “Courts-circuits algorithmiques : un nouvel âge de l’esprit,” Revue Esprit, April 2025 — the argument in a mainstream philosophical review. Alombert, A. and Falquet, J., “Déconstruire « l’intelligence artificielle ». Automates, techniques et esprits dans la philosophie contemporaine,” Université Paris 8, Département de philosophie. philosophie.univ-paris8.fr · Alombert, A., “La face cachée de l’IA : enjeux écologiques et politiques des automates numériques,” Les temps qui restent, no. 3, December 2024. Interviews: PhiLitt, 7 January 2026 philitt.fr and lundimatin lundi.am. Giuseppe Longo, CNRS Research Director Emeritus, Centre Cavaillès, ENS Paris cv.hal.science/giuseppe-longo

Note on sourcing: the formulation “automates numériques” first appeared in a July 2023 piece by Alombert and Longo in L’Humanité. That citation was replaced here — the outlet’s political alignment invites dismissal of a terminological argument on unrelated grounds, and the 2023 date predates most of what this article discusses. The same argument is available from the sources above.

[34] Isomorphic Labs Series B, May 2026, Forbes forbes.com. The IsoDDE benchmark figures and the ISM8969 regulatory clearance are reported via secondary coverage rather than a company or regulatory publication.

[35] Hassabis and LeCun on architecture: the-decoder.com · optim.vc. August 2026 DeepMind reshuffle: cnbc.com · cnbc.com

[36] OpenAI, “Our approach to frontier risk.” openai.com · METR, “Common Elements of Frontier AI Safety Policies” — comparative analysis of the published frameworks. metr.org/common-elements

[37] Google DeepMind, “Introducing the Frontier Safety Framework.” Critical Capability Levels defined across autonomy, biosecurity, cybersecurity and machine-learning R&D. deepmind.google

[38] Amodei, D., on the concentration of AI power occurring “almost overnight” and “almost by accident,” 2026. Reported via secondary coverage rather than the original interview or transcript.

[39] Amodei, D., “Policy on the AI Exponential,” June 2026. Argues for statutory authority to block frontier model deployment, with third-party auditors holding the final veto. darioamodei.com

[40] Amodei, D., on AI’s crisis of trust, Fortune, 16 August 2026. fortune.com

[41] Future of Life Institute, AI Safety Index — Summer 2026. Finds Anthropic, OpenAI, Google DeepMind and Meta have weakened or voided unilateral-pause pledges. futureoflife.org

[42] Bar Standards Board, guidance on the use of AI and other technologies, 18 May 2026. kennedyslaw.com summary

[43] State Bar of California, Standing Committee on Professional Responsibility and Conduct, proposed amendments to six ethics rules approved 13 March 2026, carrying disciplinary rather than advisory authority. clio.com overview of bar-association AI ethics opinions · esquiresolutions.com

[44] International Committee of Medical Journal Editors and major academic publishers, converging requirements on AI authorship and disclosure. thesify.ai summary of 2026 publisher policies · International Bar Association AI Working Group, guidance across nine jurisdictions ibanet.org

[45] UNESCO, AI Competency Framework for Teachers and Students (2024). unesco.org

[46] UNESCO global survey of more than 450 schools and universities: 13% of universities report guidance on generative AI; 7% of schools report an institutional policy. Figures via a secondary summary of the UNESCO survey rather than the UNESCO publication itself.