Nash equilibrium and the AI economy, 2026–2031. The best-measured result in the field says that automating your side of a transaction earns you nothing until your counterparty automates too. That is not an arms race. It is a threshold, and it changes which questions are worth asking.
The arms-race frame predicts continuous escalation and one winner. The evidence describes something else: markets that sit still until both sides are automated, then move at once.
Primary sources verified at arXiv, JPE, AER, NBER, Google PatentsEvery video checked for speaker, venue and runtime2026-08-01
The marks
Five grades, used consistently below
A citation is not a unit of evidence; a claim is. One paper below carries two claims at different grades, split rather than levelled up to the stronger one. The grade is not a quality score — a good simulation and a good field study are both good work. It records what the number is made of.
Measured
Observed in a real market, with real transactions and real money. Stated n, identified variation, reproducible in principle.
Benchmarked
Measured with the same rigour, but on a constructed task rather than a market. Internally valid; external validity is the open question. Kept separate from measured on purpose, because collapsing the two is the most common over-claim in this literature.
Modelled
Output of a model — a simulation, a cost reconstruction, a theorem. True of the model. Whether it is true of the world is a separate claim requiring separate evidence.
Self-reported
Stated by a party about its own operations or strategy. Usable, and often the only thing available, but testimony.
Not checkable
The artifact could not be obtained, or accounts of it disagree. Distinct from false.
Where a source has a stake in its own conclusion, that is noted inline rather than folded into the grade. Interested evidence is admitted evidence; it is not disqualified, but the reader is told.
The result everyone quotes, cut one sentence too early
Stephanie Assad, Robert Clark, Daniel Ershov and Lei Xu studied Germany's retail gasoline market, where algorithmic-pricing software became widely available to station operators in 2017. Station-level adoption is endogenous — the stations that adopt are not a random sample — so they instrument it with brand-headquarter-level adoption decisions, and detect adoption by testing for structural breaks in pricing behaviour. The data is high-frequency prices for every retail station in the country, January 2016 to December 2018 — roughly 1,300 of the ZIP-code markets are duopolies. Published in the Journal of Political Economy, volume 132, issue 3, pages 723–771.
The sentence that travels is: algorithmic pricing raises margins by 9% in non-monopoly markets. That is accurate — adopting stations gain about 0.8 cents per litre against an average non-adopter margin of 8.3 cents — and it is the wrong stopping point. The paper continues into duopolies, and the duopoly result is the one that carries the argument.
Adoption increases margins but only for nonmonopoly stations. In duopoly and triopoly markets, margins increase only if all stations adopt.
Published abstract — Journal of Political Economy 132(3), 2024, pp. 723–771
Unilateral adoption does nothing measurable. Mutual adoption moves the market. Everything downstream in this article is an attempt to work out which other markets have that shape.
The magnitude needs a version number
The published abstract, quoted above in full on this point, contains no percentage at all — the result it states is qualitative. The figure most often attached to it in secondary literature is +28% for duopolies where both stations adopt, drawn from the published paper's body. The January 2021 working paper, which is freely readable, reports the same experiment as a 3.2 cent per litre rise, “a substantial increase of nearly 38% relative to the baseline.”
So the estimate moved by roughly ten percentage points between the working paper and the journal — which is peer review doing its job, and a warning that quoting “the” number without a version is unsafe. This article leans on the qualitative claim, which is identical across both versions and is what the argument actually requires. Magnitudes are labelled with the version they come from, and the +28% is marked Not checkable here because the published body sits behind a paywall this research could not open.
One more sentence past the good part, from the working paper, because it closes the obvious escape hatch. If one-sided adoption looks like nothing because the adopter's gain cancels the rival's loss, the threshold reading collapses. The authors tested exactly that: comparing non-adopters in duopoly markets before and after their rival adopts, they “do not see any statistically significant changes in margins and prices following a rival's AI adoption, ruling out this explanation.” One-sided automation does not quietly redistribute. It does nothing.
Try it
Set each station's pricing method. The highlighted cell is the observed outcome for that combination.
Observed market-level margin change, German retail gasoline duopolies
Station A ↓ / B →
Manual pricing
Algorithmic pricing
Manual pricing
baselinePre-adoption margin, the reference point
no changeOne adopter. No statistically detectable market-level effect
Algorithmic pricing
no changeOne adopter. No statistically detectable market-level effect
margins riseBoth adopt. +28% published body (paywalled); +38% in the 2021 working paper
Station A
Station B
What this table is not
It is an outcome table, not a payoff matrix. The distinction matters and is routinely lost. Assad and co-authors measured market-level margins; a payoff matrix requires each player's own profit under each combination, which the paper does not report. Reading the both-adopt cell as a Nash equilibrium requires an additional assumption — that an individual station's profit moves with the market margin, and that no station has a profitable unilateral deviation out of that cell. That assumption is plausible. It is not tested here.
Nor is it evidence of an agreement, and the authors say so themselves in a footnote on their first page: “To our knowledge, there is no direct evidence of anticompetitive behavior on the part of any algorithmic-software firms or gasoline brands mentioned in this paper.” They analyse the effect strictly as economics. Anyone citing this study as proof of collusion is citing something the authors explicitly declined to claim.
So the honest statement is narrower than the one that gets repeated, and still striking: the returns to automation in this market were threshold-shaped, appearing only when both sides automated. Whether that constitutes equilibrium play, and whether it constitutes collusion, are the two live questions — and the field does not agree on either. Section 09 leaves them open.
Section 02
What a Nash equilibrium buys you here, and what it does not
A Nash equilibrium is a profile of strategies where no player gains by unilaterally changing theirs. That is a definition, not a prediction. It buys you three specific things in this domain, and it is worth being blunt about the fourth thing it does not buy.
It tells you where to look for stability
If a configuration is not an equilibrium, someone is leaving money on the table and the configuration should decay. Persistent non-equilibrium arrangements are evidence that the model has the payoffs wrong — which is usually more interesting than the model being right.
It separates “nobody wants this” from “nobody can stop”
Prisoner's-dilemma structures produce outcomes every participant dislikes and none can unilaterally escape. That is a different problem from one where participants simply want different things, and it takes a different intervention — changing payoffs rather than changing minds.
It makes threshold effects visible
Coordination games have multiple equilibria and the interesting question is which one gets selected, and when a market tips between them. The gasoline result is unintelligible as a race and legible as a tipping problem.
It does not tell you what will happen
Equilibrium selection is unsolved in general. Real players are boundedly rational, the payoff matrices in this article are reconstructions rather than observations, and economics is a field of contested models. Every prediction below is conditional on a model being right about the world, which is a claim requiring its own evidence.
Ben Cottier, Robi Rahman, Loredana Fattorini, Nestor Maslej, Tamay Besiroglu and David Owen built a cost model for frontier training runs and found growth of 2.4× per year since 2016, with a 90% confidence interval of 2.0× to 2.9×.
The scope is narrower than the headline. This is the amortised cost of hardware and energy for the final training run — not total AI capital expenditure, not research staff across a programme, not inference infrastructure. It is a reconstruction from external observables by researchers without access to the labs' internal accounts, which is why it carries a confidence interval. Treating it as “AI spending grows 2.4× a year” imports a precision the method does not deliver.
The paper's forward statement — that on current trends the largest runs pass a billion dollars by 2027 — is a projection of a fitted trend, and inherits every assumption in the fit.
The equilibrium question, and why it is harder than it looks
An arms race is a specific structure: defection dominates, the cooperative outcome is unreachable without commitment, and every player is worse off than under an agreement none can sign. Whether frontier capex has that structure depends on the payoff to winning — and the payoff to winning is falling fast, because the capability you spent a billion dollars to reach becomes purchasable for a fraction of that within roughly a year. Section 04 documents the decline.
That combination — escalating entry cost, depreciating prize — is not the classic arms race. It is closer to a war of attrition, where the equilibrium prediction concerns not who spends most but who exits first, and where the answer turns on each player's outside option rather than its balance sheet. Firms whose core business is not model sales have a different outside option from firms whose only asset is a frontier model, and the theory says they should behave differently at the exit margin.
Nobody can instrument this. A lab's stopping rule is an internal deliberation with no external log. There is no measured evidence about who stops first and there will not be until someone stops. What exists is theory and testimony, and this article does not dress either as observation.
Section 04MeasuredModelled
Pricing: a Bertrand problem with no interior solution
Inference has near-zero marginal cost and very high fixed cost. Textbook Bertrand competition on a homogeneous good under those conditions drives price to marginal cost, which does not cover fixed cost, so no firm can survive — the classic result that a pure-strategy equilibrium fails to exist in any form that sustains the industry. That is why the observed behaviour is not a straight price war but relentless differentiation: capability tiers, context windows, latency classes, rate limits, bundling into subscriptions, and enterprise agreements that make per-token comparison difficult on purpose. Differentiation is not a marketing artifact here; it is the only thing standing between the market and an equilibrium in which nobody recovers fixed costs.
How fast prices actually fall, stated carefully
Epoch AI's price data is a genuine measurement — posted API prices are public and observable. But its own headline is that prices have fallen “rapidly but unequally across tasks”, and the inequality is the finding.
Rate of decline in price to reach a fixed capability level
Measure
Value
What it is
Range across six benchmarks
9× – 900× / yr
The spread, not an error bar. A 100-fold difference between the slowest and fastest task.
Median across all trends
50× / yr
A median over a range this wide describes no particular task.
Median, trends starting after Jan 2024
200× / yr
Faster recently — and a shorter window, so a less settled estimate.
GPT‑4 performance on PhD-level science questions
40× / yr
A single, concrete, well-specified milestone. The most quotable number here.
A widely circulated example needs correcting, including in the brief that commissioned this piece. On FrontierMath, reaching roughly 27% accuracy took about 43 million output tokens with o4‑mini at high reasoning effort in April 2025, and about 5 million tokens with GPT‑5.2 at low reasoning effort that December. The token ratio is about 8.6×. The cost ratio is not, because the two models are priced differently per token — and Epoch says so in the next sentence: “Even accounting for the difference in price per output token, that's roughly a 3x cost reduction over eight months.” Quoting the token counts and stopping is a threefold over-claim produced without a single false statement.
Declared interest: Epoch AI is a research organisation with no model to sell, which is why its price series is unusually usable. The framing is nonetheless directional — the Gradient Updates piece by Jean-Stanislas Denain (16 February 2026) argues inference costs are not a persistent burden, and the 5–10×-per-year figure it offers is an author's summary judgement across mechanisms, not a fitted trend.
The equilibrium implication is uncomfortable for the seller and excellent for the buyer: with capability commoditising on this schedule, durable margin cannot come from the model. It has to come from distribution, switching costs, proprietary data, or the workflow the model sits inside — which is precisely where the strategic action in sections 05 through 07 turns out to be.
Section 05
Agents against agents: the market that already exists
The literature on autonomous agents negotiating with autonomous agents is usually described as speculative. It is not. Online advertising has run agent-versus-agent auctions at enormous scale for years, and it has a serious equilibrium literature. The mistake is looking for the future in chatbots when it is already operating in ad exchanges.
What follows is deliberately built as a ladder, strongest evidence first, because the grades diverge sharply and the conclusions people draw usually borrow the confidence of the top rung for claims that belong on the bottom one.
The German gasoline study in section 01. Note what the agents were: commercial pricing software configured by humans, not learning systems pursuing open-ended objectives. This is the strongest evidence available about automated agents changing a market outcome, and it is about a technology considerably simpler than the one everyone is now deploying.
Auto-bidding: the largest agent-versus-agent economy running today
“Auto-bidding and Auctions in Online Advertising: A Survey” (14 August 2024) is a 25-author survey covering bidding algorithms, equilibrium analysis and efficiency of common auction formats, and optimal auction design for markets that have, in the authors' framing, embraced autobidding — advertisers state high-level goals and budgets, and an automated agent bids on every query on their behalf. This is agents negotiating with agents, billions of times a day, with real money settling.
Declared interest: the author list — Aggarwal, Mehta, Mirrokni, Paes Leme, Sivan and others — is substantially Google Research, surveying the mechanism design of a market Google operates and profits from. That does not make the mathematics wrong; the theory is checkable independently. It does mean the survey's characterisation of how well these markets perform is an interested account, and the adoption framing is self-reported rather than independently measured.
Learning algorithms find collusion without being told to
Emilio Calvano, Giacomo Calzolari, Vincenzo Denicolò and Sergio Pastorello, American Economic Review 110(10), October 2020, pages 3267–97. Q‑learning agents in a repeated Bertrand oligopoly “consistently learn to charge supracompetitive prices, without communicating with one another,” sustained by strategies with “a finite phase of punishment followed by a gradual return to cooperation.” Robust to cost and demand asymmetries, to the number of players, and to several forms of uncertainty.
The punishment structure is the striking part: these are recognisably the trigger strategies that sustain tacit collusion in repeated-game theory, arrived at by agents with no model of each other and no channel to communicate. But this happens in a simulated market with a specified demand system and a small discrete action space. It is a result about Q‑learning in that environment. Whether production pricing systems in real markets do the same thing is exactly the question the German gasoline study addresses empirically — and the two papers are complementary, not confirmatory.
Wenyue Hua, Ollie Liu, Lingyao Li and co-authors (8 November 2024) put LLMs into classical negotiation games. The finding: models “frequently deviate from rational strategies, particularly as the complexity of the game increases” — larger payoff matrices, deeper game trees — and struggle to compute Nash equilibria reliably under uncertainty and incomplete information.
Stopping there would misrepresent the paper, because that is its motivation rather than its conclusion. The contribution is a structured game-theoretic workflow that substantially improves performance and reduces the models' “susceptibility to exploitation during negotiations.” So the correct reading is not models cannot negotiate but strategic competence here comes from scaffolding, not from the model — which relocates the strategic question from model capability to who writes the workflow. These are constructed games with known payoff structures, not markets.
“Multi-Agent Risks from Advanced AI” (Cooperative AI Foundation Technical Report #1, 19 February 2025), led by Lewis Hammond, Alan Chan and Jesse Clifton with some 50 contributors. Three failure modes — miscoordination, conflict, collusion — across seven risk factors including information asymmetries, commitment problems, selection pressures and emergent agency. Its value here is conceptual: it supplies the vocabulary for distinguishing failure types that get lumped together, and it identifies commitment as the pivot, which is the right call. Classical results on repeated games turn on the ability to commit credibly, and an agent whose policy can be inspected or whose source can be shown is a fundamentally different bargaining counterparty from a human who can only promise.
Declared interest: an organisation constituted around cooperative AI, reporting that AI systems will have cooperation problems. The taxonomy is a framework rather than a measurement; it is offered as such and is not evidence that these failures occur at any particular rate.
Patents: what firms bothered to defend, and what they let go
Patents are a distinct evidence class. They do not show what a company built — most are never shipped — but they show what it thought was worth the cost of defending, and when. The filing date is a dated record of corporate belief, which is unusual and useful.
Microsoft described 2026's agentic commerce in 2016, then abandoned it
“Artificial Intelligence Negotiation Agent.” Assignee Microsoft Technology Licensing LLC, inventor Georgios Krasadakis, filed and priority-dated 31 March 2016, published 5 October 2017. The application describes buyer-side and seller-side AI agents that discover each other, exchange offers across multi-round negotiations within user-defined parameters and elasticity, and close transactions autonomously, informed by real-time competitor pricing, inventory and sentiment data.
That is a functional description of what the industry announced as new roughly a decade later. Its legal status on Google Patents is abandoned. A company with every resource needed to pursue this filed it early and let it lapse — and the most economical explanation is not that the idea was bad but that a negotiation mechanism was never going to be the defensible asset. Which is exactly what the threshold argument predicts: a mechanism only pays once both sides adopt it, and you cannot drive both sides to adopt something you are charging rent on.
A negative result: the “agentic” payment patents are not about agents
An industry-analytics post from February 2026 presents a portfolio of Mastercard “agentic AI patents” as evidence that the payment networks are staking out agent-to-agent commerce. It cites three specific numbers, which is commendable and also what makes it checkable. All three were pulled at Google Patents. None of them is about agents negotiating.
Cited as “agentic AI”
What the primary record says
Negotiation?
US11250461B2
“Deep learning systems and methods in artificial intelligence.” Granted 15 Feb 2022, inventor Suqiang Song, status active. Claim 1 is a propensity engine, a model-serving engine and a fulfilment engine using neural collaborative filtering to match cardholders to merchant offers.
None. It is a recommender.
US2024/0119517
Predicting creditworthiness of merchants from invoice and transaction-network data. Filed Dec 2022, published Apr 2024.
None. It is credit scoring.
US2023/0385701
An engine for entity resolution and standardisation of transaction records using natural-language processing. Filed May 2023.
None. It is data cleaning.
Three conventional machine-learning filings, two of them predating the current agent wave, relabelled downstream as “agentic.” The word was applied by the commentary, not by the patents. Nothing here was fabricated — the numbers are real, the assignee is right, the patents exist — and the claim built on top of them still does not hold. That is the more common failure than invention, and it costs one lookup per number to catch.
Reporting this matters more than the positive finding above. The patent record does not show the payment networks fencing off agent-to-agent negotiation; the search for that fence came back empty. What the networks are doing instead is publishing open protocols. A firm that believed the negotiation mechanism were the durable asset would be filing on it and licensing it, and the absence of those filings is evidence about what they believe — which is the whole reason to read patents as a source class.
Self-reported
The protocols are being given away, which is the tell
The agentic-commerce plumbing is arriving as open specifications rather than proprietary moats — an Agentic Commerce Protocol from OpenAI and Stripe, Google's Agent Payments Protocol and Universal Commerce Protocol, Anthropic's Model Context Protocol, alongside agent-payment products from Visa and Mastercard. Announcement details, launch-partner counts and transaction volumes here are vendor statements; this article does not treat them as measured, and deliberately does not lean on the widely quoted multi-trillion-dollar market projections, which are models produced by parties selling into the market.
The structural point survives without those numbers. Every serious participant is trying to give the protocol away. Under the threshold logic that is the rational move rather than a generous one: the value is unlocked at mutual adoption, so the binding constraint is getting the other side automated, and a licensing fee on the mechanism is a tax on reaching the threshold you need.
Open weights: commoditise your complement, partially confirmed
Releasing weights is costly and irreversible, which makes it a clean strategic move to analyse. The standard story is Joel Spolsky's: commoditise your complement, so that demand shifts to the thing you still sell. Mahyar Habibi's “Open Sourcing GPTs: Economics of Open Sourcing Advanced AI Models” (20 January 2025) tests it, and the story survives in part.
Finding
Direction
Reading
Likelihood of open-sourcing vs. the model's performance edge over the best existing open model
− 10–11 pp
Per 10-point quality gain on a 100-point scale. Firms release what does not lead. Consistent with protecting rents on the frontier, and with commoditising the tier below.
Big Tech vs. other for-profit organisations
+ 20%
Ceteris paribus. The firms with the largest complementary businesses — cloud, devices, distribution — are the most willing to give the model away.
Owner's share of compatible applications
inverted U
A theoretical prediction: propensity to open-source rises then falls with the owner's size. Own too little and you capture none of the spillover; own too much and you cannibalise yourself.
The first two rows are empirical; the third is the model's prediction and is marked as such in the paper. The inverted‑U is the part worth holding lightly and watching, because it is the row that would explain why the same firm releases weights in one year and withholds them the next without either decision being a change of philosophy.
What the commoditise-your-complement story does not explain on its own is release as a recruiting and standard-setting move — getting an architecture into the tooling, the papers and the graduate curriculum. Those payoffs are real and are not in the model, which is a reason to treat the fit as partial rather than a confirmation.
Section 08Modelled
Safety and evaluation: undersupplied unless someone is paid
Safety research and shared evaluation infrastructure have the two defining properties of a public good: one firm's use does not diminish another's, and non-contributors cannot practically be excluded from the benefit. The standard result follows immediately — in equilibrium it is undersupplied relative to the social optimum, because each firm's private return on the marginal safety dollar is a fraction of the social return.
This is theory, and it is unusually robust theory, but it remains a claim about a model. The measured version — how much is actually spent on safety across the industry, against what a welfare calculation would prescribe — requires internal accounting nobody publishes and a social welfare function nobody agrees on. Anyone stating an underinvestment ratio as a fact is reporting an estimate built on assumptions they should be showing you.
Two structural notes that follow from the sections above rather than from first principles. Evaluation is the clearer case than safety research: a shared benchmark is close to a pure public good, and the observed pattern — benchmarks built by academics, non-profits and consortia rather than by the firms with the most to gain from them — is what undersupply looks like from outside. And the collusion failure mode in the Cooperative AI taxonomy is precisely the one that the gasoline evidence suggests appears at a mutual-adoption threshold, which means the monitoring problem gets materially harder at exactly the moment it starts to matter. Detecting tacit coordination between two automated systems that never communicate is not a solved problem in either economics or computer science.
Section 09
Where the field genuinely disagrees
These are presented unresolved because they are unresolved. A survey that adjudicates them is over-claiming.
Does collusion require intent?
Calvano and co-authors show algorithms reaching supracompetitive prices with no communication and no instruction to collude. Competition law in most jurisdictions turns on agreement or concerted practice — a mental state that a Q‑learning agent does not have. Either the economics of harm and the law of liability come apart, or the legal concept has to be rebuilt around outcomes rather than intent. Both positions are seriously argued, and the gasoline result does not settle it: the authors identify a margin effect, not an agreement.
Does AI concentrate markets?
Hal Varian's “Artificial Intelligence, Economics, and Industrial Organization” (NBER Working Paper 24839, July 2018) examines how machine learning availability affects the industrial organisation of firms supplying and adopting it, and is notably unwilling to conclude that AI implies monopoly — emphasising limits to returns from data and the availability of AI capability through cloud services to firms that could never build it. Declared interest: written by Google's Chief Economist. Against that, the cost trend in section 03 points toward a shrinking set of organisations able to train at the frontier, which is a claim about the supply side that Varian's argument about the adoption side does not contradict. Much apparent disagreement here is two questions wearing one name.
Who should hold the property right in training data?
Joshua Gans, “Copyright Policy Options for Generative Artificial Intelligence” (NBER Working Paper 32106), frames this as a bargaining problem and reaches a conditional answer: for small models, where content providers can feasibly negotiate with AI providers, copyright protection produces superior welfare outcomes; for large models, where negotiation is prohibitive, the comparison is ambiguous and turns on the harm to content providers and the importance of content to training quality. The honest summary is that the welfare ranking depends on transaction costs, and transaction costs are exactly what the agent infrastructure in section 06 is built to reduce — so the answer is not stable over the period this article covers.
Section 10
Four predictions, each with a date and a disproof
The point of an equilibrium frame is that it commits to something. Every row below names a date, an observable a reader can actually go and check, and the specific finding that would kill it. Nothing here rests on a vendor's market-size model or on internal spending that no one publishes — if a prediction could not be checked from public records, it was cut rather than softened. Two were cut on that rule; they are named underneath.
Prediction
Check by
Where to look
Disproved if
The major agent-payment and agent-commerce protocols remain published specifications that can be implemented without a per-transaction licence fee to the specification's author.
31 Dec 2029
The published specification and its licence terms for each protocol named in section 06.
Any one of them moving to terms that charge the implementer a fee per transaction for use of the negotiation or mandate mechanism itself, while competitors implement it anyway.
At least one competition authority in the EU, UK or US opens a formal investigation in which the alleged coordination is between pricing or bidding algorithms with no communication between the firms.
31 Dec 2030
Case registers and press releases of DG COMP, the CMA, the FTC and the DOJ Antitrust Division.
No such investigation opened by that date, or every opened case pleading a conventional communicated agreement instead.
No lab open-sources model weights that are, on the day of release, the top-ranked model on a major public leaderboard — and if one does, it does not do it twice.
31 Dec 2029
Release announcements and licence files, against leaderboard standings on the release date.
Two separate open-weight releases that each hold the top public ranking at release. One is a plausible outlier; two is a pattern, and Habibi's result would not survive it.
At least one peer-reviewed empirical paper measures a price or margin effect from agent-mediated transactions using an adoption-threshold design — comparing markets by how many sides are automated, as Assad and co-authors did — and finds the effect concentrated where both sides are automated.
Such a study finding the effect present with one-sided automation, or spread evenly regardless of counterparty automation. Also disproved, more weakly, if nobody runs the study at all — an untested thesis is not a confirmed one.
Cut, not softened. Two predictions from an earlier draft failed the rule and were removed rather than hedged. “Frontier capex separates, and model-only firms exit first” depends on exit order tracking outside options rather than balance sheets, and there is no public record that distinguishes those motives — a firm that sells or shuts down does not publish which constraint bound. “Shared evaluation stays underfunded relative to capability spending” requires both an industry safety-spend figure and a training-spend figure, and section 08 already concedes that neither is published. A prediction whose evidence does not exist is a slogan with a date on it.
Source classes
What each kind of source was actually good for
Papers — the only place the measured claims came from
Every load-bearing empirical number in this article is from a peer-reviewed paper or a working paper verified at its primary record: JPE 132(3) for the gasoline result, AER 110(10) for the simulation, arXiv:2405.21015 for costs, arXiv:2501.11581 for release behaviour, NBER 24839 and 32106 for the disagreements. Papers state their n, their identification strategy and their limits. Nothing else in the source list does.
Video — good for framing, checked for existence
Four talks verified through the Simons Institute's own records rather than a search summary, from the Algorithmic Game Theory and Practice workshop, Calvin Lab Auditorium, 20 November 2015: Paul Milgrom, “Adverse Selection and Auction Design for Internet Display Advertising” (43:37); Susan Athey, “Designing Online Advertising Markets” (48:54); R. Preston McAfee, then at Microsoft, “Machine Learning in an Exchange Environment” (49:38); Jon Kleinberg, “On-Line Systems with Long-Range Goals” (54:52). Full-length seminar talks by the people who built the field — Milgrom shared the 2020 Nobel for auction theory. Their date is the point: the equilibrium analysis of automated advertising markets was mature a decade before “agentic commerce” was a phrase.
Also verified: Chi Jin (Princeton), “Multi-Agent Reinforcement Learning” Parts I and II, Learning and Games Boot Camp, 28 January 2022, scheduled 2–3pm and 3:30–4:30pm PT. Speaker, venue, date and session confirmed on the Institute's talk pages; recorded runtimes could not be retrieved, so no duration is asserted for these two.
Google Tech Talks — this source class is empty
Stated plainly rather than papered over: this article cites no Google Tech Talk, because no relevant one could be verified to exist. The channel's current feed carries no economics, auction or game-theory talk — the fifteen most recent entries are differential privacy, membership inference, machine unlearning and model poisoning. Secondary summaries assert that a relevant talk exists; it is not in the feed, and a title that survives only in summaries of itself is precisely what this class invites you to cite.
The cost of filling the slot anyway is not hypothetical. A “conference talk” checked during this project turned out to run sixty-seven seconds. A named void is a finding; an unverified citation is a liability that looks like a finding.
Patents — dated evidence of intent, including intent withdrawn
The most valuable finding in this article came from a patent, and specifically from its legal status: an abandoned 2016 Microsoft application describing autonomous buyer and seller agents. Patents also produced the most useful negative result, when the Mastercard “agentic” portfolio turned out on inspection to be recommender and credit-scoring work. Both checks took one primary lookup each, which is the argument for the class.
Research blogs — timely, directional, always affiliated
Epoch AI supplied the inference price series, which no paper covers at comparable recency, and Epoch has no model to sell. But the blog format argues a thesis, and the piece cited here argues one — that inference cost is not a persistent burden. The measurement and the argument were separated above, and the widely repeated version of its own example was corrected against its next sentence.
Corrections made during research
The FrontierMath token example is a threefold cost improvement, not eightfold. The 43M → 5M output-token comparison implies 8.6×; Epoch's following sentence puts the cost reduction at roughly 3× once per-token prices are accounted for. Corrected in section 04.
The 2.4×/year figure is not “AI training spend.” It is amortised hardware and energy for the final training run, modelled from outside with a 90% interval of 2.0×–2.9×. Scope stated in section 03.
The gasoline duopoly table is an outcome table, not a payoff matrix. The paper reports market-level margins, not per-station profits; calling the both-adopt cell a Nash equilibrium requires an untested assumption, which is now stated rather than assumed in section 01.
The LLM negotiation paper's headline is its motivation, not its finding. Citing only the failure to compute equilibria inverts the paper, whose contribution is a workflow that largely repairs the failure. Both halves are in section 05.
“Mastercard agentic AI patents” did not survive a primary check. Reported as a negative result in section 06 rather than dropped.
Papers verified at arXiv, AEA, University of Chicago Press and NBER primary records; patents at Google Patents, including legal status; talks at the Simons Institute's own event pages, with runtimes where retrievable. Where a runtime, a market size or an adoption figure could not be confirmed at primary source, it is either marked or absent. Interested sources are cited and labelled, not excluded.