<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[AI Policy Perspectives ]]></title><description><![CDATA[Contributions from a range of thinkers on AI policy and governance topics, all in a personal capacity.]]></description><link>https://www.aipolicyperspectives.com</link><image><url>https://substackcdn.com/image/fetch/$s_!XGVU!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png</url><title>AI Policy Perspectives </title><link>https://www.aipolicyperspectives.com</link></image><generator>Substack</generator><lastBuildDate>Sun, 13 Sep 2026 03:57:44 GMT</lastBuildDate><atom:link href="https://www.aipolicyperspectives.com/feed" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><webMaster><![CDATA[aipolicyperspectives@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[aipolicyperspectives@substack.com]]></itunes:email><itunes:name><![CDATA[AI Policy Perspectives]]></itunes:name></itunes:owner><itunes:author><![CDATA[AI Policy Perspectives]]></itunes:author><googleplay:owner><![CDATA[aipolicyperspectives@substack.com]]></googleplay:owner><googleplay:email><![CDATA[aipolicyperspectives@substack.com]]></googleplay:email><googleplay:author><![CDATA[AI Policy Perspectives]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The AI Paper Trail (#7)]]></title><description><![CDATA[What we're reading]]></description><link>https://www.aipolicyperspectives.com/p/the-ai-paper-trail-7</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/the-ai-paper-trail-7</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Thu, 10 Sep 2026 12:21:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HdCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every month we try to stave off the decline in reading by taking a look at new AI papers that we&#8217;ve seen folks discussing. Today, we look at how Russian propaganda is targeting LLMs; whether AI models can accurately predict human behavior; whether they can get a paper published at the NeurIPS conference; and the growing risk of hackers implanting malicious instructions into data that AI agents use. </em></p><p><em>Please share your own take and any new papers that you&#8217;ve enjoyed.</em></p><p><em>&#8212;Conor Griffin, AI Policy Perspectives</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HdCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HdCF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HdCF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini </figcaption></figure></div><h3><strong><span>Information warfare is targeting LLMs</span></strong></h3><ul><li><p><strong><span>What&#8217;s the paper about?</span></strong><span> Researchers from the UK think tank Demos dived into the </span><a href="https://demos.co.uk/research/geo-for-geopolitics-what-happens-when-ai-and-information-warfare-collide/"><span>case</span></a><span> of LLMs endorsing falsehoods from a Russian propaganda source to highlight how manipulative influence campaigns are now targeting AI models.</span></p></li><li><p><strong><span>Why does it matter?</span></strong><span> Discussions about AI and disinformation often focus on threat actors using models to directly generate and flood misleading claims across social media. But as more people turn to AI for information and advice, threat actors also want to shape how LLMs respond.</span></p></li><li><p><strong><span>The details: </span></strong><span>To improve their answers, AI systems use retrieval-augmented generation, or RAG, to identify relevant documents and information by calling up search engines and databases. This introduces a risk of &#8220;RAG poisoning,&#8221; where malign actors insert distorted information into the models&#8217; sources.</span></p></li><li><p><span>The Demos authors illustrate this by evaluating how well a Russian foreign-interference operation&#8212;the </span><a href="https://data.europa.eu/apps/eusanctionstracker/subjects/177836"><span>Foundation to Battle Injustice</span></a><span>, or r-FBI, which falsely presents itself as a human-rights organization&#8212;has infiltrated unsubstantiated allegations about its opponents into LLMs.</span></p></li><li><p><span>The authors extracted 50 specific claims from r-FBI articles that had no corroborating secondary source, such as an assertion that the Ukrainian president, Volodymyr Zelensky, had installed cryptocurrency farms that were causing blackouts. They fed 600 prompts about these claims to five AI models: GPT-4.1 Mini; Gemini 2.5 Flash; Claude Haiku 4.5; Mistral Small 3.2; and Grok 4.20.</span></p></li><li><p><span>Around 17% of model responses endorsed the r-FBI&#8217;s claims, repeated them uncritically, or presented them as a legitimate view. An additional 31% of responses were neutral&#8212;addressing the topic, but neither validating nor rejecting the specific claim. The remaining 52% of responses rejected the claims, with a subset doing a &#8220;comprehensive debunk&#8221; that exposed r-FBI as a foreign-interference operation.</span></p></li><li><p><span>Grok performed best, comprehensively debunking 56% of queries, compared with 17% for Gemini, 10% for Claude, 5% for GPT, and just 1% for Mistral. The authors suggest that Grok&#8217;s stronger performance may reflect its RAG setup and access to data from X. Or it may reflect Grok 4.20&#8217;s larger size&#8212;the other LLMs tested were small models of the type that underpin quick chatbot and AI-search responses.</span></p></li><li><p><span>The Demos authors found that r-FBI ran a sophisticated campaign to get LLMs to find, trust, and cite their content. This included using protocols to push r-FBI&#8217;s &#8220;breaking news&#8221; articles into the live-search indices that RAG pipelines use, as well as using long page titles and descriptions. This use of lengthy text hurts the visibility of r-FBI&#8217;s articles in traditional search engines, but allows AI crawlers to ingest complete passages, suggesting that r-FBI was optimizing for LLMs, not humans.</span></p></li><li><p><span>In its content, r-FBI also makes editorial choices that the Demos team see as targeting LLMs. This includes packing lots of statistics and quotations into self-contained paragraphs of a length that retrieval systems can readily extract, as well as prioritizing absolute claims over nuance. They also distribute their content through what purport to be different kinds of institutions and formats, including news articles, human rights reports, and foreign policy journal articles, and have each publication repeat the core claims and cite each other, in an effort to increase their salience in AI model embeddings and suggest support from multiple sources.</span></p></li><li><p><span>Making content more visible to LLMs&#8212;referred to as generative engine optimization, or GEO&#8212;can be benign; a range of legitimate GEO providers exist to help organizations ensure that their offerings are visible and correct. However, of 50 GEO providers identified by the Demos team, only one&#8212;</span><a href="https://www.finnpartners.com/service/ai/"><span>Finn Partners</span></a><span>&#8212;had a published policy to prevent misuse of its services.</span></p></li><li><p><span>By drawing on the US Foreign Agents Registration Act, which requires anybody acting on behalf of an overseas government to disclose this, the researchers found two examples of government entities hiring GEO providers to help shape AI outputs, although the authors were careful not to equate this with information warfare.</span></p></li><li><p><span>Of the five AI companies whose models Demos evaluated, only Google appeared to have a published policy that explicitly mentioned GEO practices, where they warn against the creation of large numbers of pages and </span><a href="https://developers.google.com/search/docs/fundamentals/ai-optimization-guide#create-valuable-content"><span>caution</span></a><span> that &#8220;GEO hacks&#8221; to make content more digestible to AI crawlers are unnecessary and ineffective.</span></p></li><li><p><span>The challenge for AI companies is to get their models to reason about the quality and reliability of a source. For queries where the evidence is genuinely unclear, models should be able to convey this. When the query is about disinformation, LLMs should be able to recognize this and strongly debunk it.</span></p></li><li><p><span>Demos&#8217; recommendations include making counter-disinformation easier for AI systems to retrieve&#8212;a form of ethical GEO; requiring AI companies to explicitly prohibit harmful GEO and clearly label and deprioritize content from known manipulative sources; fostering better intelligence-sharing among AI companies about RAG-poisoning campaigns; and creating industry standards for GEO providers that distinguish marketing from deception.</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><h3><strong><span>AI can (somewhat) predict how humans think and behave</span></strong></h3><ul><li><p><strong><span>What&#8217;s the paper about? </span></strong><span>A team of Harvard and Stanford researchers </span><a href="https://www.nature.com/articles/s41586-026-10742-x"><span>found</span></a><span> that LLMs could forecast the results of social science experiments on human attitudes and behavior with high accuracy, although the models overestimated the size of the effects.</span></p></li><li><p><strong><span>Why does it matter?</span></strong><span> Rather than replacing human participants, AI simulations could serve as cheap forecasters, expanding the number of ideas that researchers and policymakers can consider before helping them filter down to those most worthy of expensive real-world testing.</span></p></li><li><p><strong><span>The details: </span></strong><span>The authors built a database of 70 large US social science experiments that collectively measured 120,000 people&#8217;s attitudes to topics such as immigration and criminal justice reform, as well as the effects of text-based interventions, like reading articles.</span></p></li><li><p><span>They used GPT-4 to simulate participants across variables such as gender, ethnicity, and political persuasion, and used this to predict the average effect of each study&#8217;s interventions.</span></p></li><li><p><span>The simulations accurately predicted the direction and relative size of the interventions&#8217; effects, but overestimated their absolute size. This means that they may be more useful for ranking the plausibility of hypotheses than for estimating how big specific effects will be.</span></p></li><li><p><span>The results held up on a subset of studies published after GPT-4&#8217;s training data cutoff, suggesting that the findings were not due to simple memorization of training data&#8212;although unpublished experiments may resemble older studies, so it&#8217;s hard to judge whether LLMs are modeling human psychology or just learning regularities about social science findings.</span></p></li><li><p><span>Bias is another concern about such LLM simulations. The authors found that the simulations were slightly less accurate for Black participants, but did not find major differences across ethnicity or gender. However, the original experiments on people had focused exclusively on the US, and also did not identify large differences across demographics. As such, questions remain about the validity of LLM simulations in other locations, or for topics where participant identity is more influential.</span></p></li><li><p><span>An accompanying survey suggests that social scientists are enthusiastic about LLM simulations, especially the ability to cheaply pilot more experiments. The authors note that the error rate of a 385-human pilot study costing around $1,200 is comparable to a combination of LLM simulations and a 100-human pilot study that jointly costs around $300.</span></p></li><li><p><span>Why do we still need human experiments? LLM simulations did not provide precise estimates of effect size and their accuracy also dropped on a secondary database of &#8220;megastudies.&#8221; Some of these focused on populations with deeply entrenched beliefs&#8212;for example, about climate change&#8212;where effect sizes are normally small. Others went beyond measuring stated attitudes and tracked behavior change, such as vaccination uptake, where frictions such as waiting times in local pharmacies make a difference but are harder to simulate. Real-world studies can also help explain </span><em><span>why</span></em><span> interventions succeed or fail, rather than just predicting the outcome.</span></p></li><li><p><span>One </span><a href="https://arxiv.org/pdf/2603.17218#:~:text=Alignment%20via%20RLHF%20(Ouyang%20et%20al.%2C%202022),ap%2D%20prove%20of%E2%80%93cooperative%2C%20fair%2C%20and%20socially%20appropriate."><span>question</span></a><span> is whether LLM simulations might degrade as AI companies carry out more post-training to make their models safer and more helpful, potentially rendering them less descriptive of what humans actually think and do. This study does not find evidence of that, with GPT-4 outperforming earlier variants. They also find that recent, smaller, open-weight models, including DeepSeek-V3 and Gemma 3, perform similarly to the larger, older, GPT-4&#8212;open-weight models may also be attractive to researchers who want to run their experiments locally and share their code to allow for replication. </span></p></li><li><p><span>Another question requiring future study is whether the growing use of LLM simulations will actually widen the range of ideas that researchers pursue, or narrow it&#8212;for example if researchers converge on similar ideas or if models struggle to simulate more novel ideas.</span></p></li><li><p><span>Malign actors could also use LLM simulations to shift public opinion or behavior. For instance, the authors note that the GPT-4 simulations accurately predicted which content&#8212;including misleading messages&#8212;was most likely to reduce a population&#8217;s intention to get vaccinated. They argue that AI companies have safety measures to prevent users from directly generating such misleading content, but lack restrictions on running simulations of experiments that could reveal effectively the same thing. In theory, this capability should be available to well-intentioned researchers, but not to others&#8212;a difficult goal, especially on politically fraught topics.</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/the-ai-paper-trail-7?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/the-ai-paper-trail-7?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3><strong><span>AI agents can&#8217;t do AI research as well as humans</span></strong></h3><ul><li><p><strong><span>What&#8217;s the paper about? </span></strong><span>A consortium of distinguished thinkers </span><a href="https://arxiv.org/pdf/2607.27191"><span>tested</span></a><span> whether an AI agent can produce a research paper good enough for a leading conference. Their answer? No. The agent was a poor judge of what&#8217;d make for a good study; dug its heels in when pursuing a weak approach; and failed to employ its resources well.</span></p></li><li><p><strong><span>Why does it matter? </span></strong><span>The AI community is debating whether agents could automate the development of new AI systems, and whether this will lead to rapid recursive self-improvement (RSI). This study suggests that AI&#8217;s lack of scientific judgment is a key obstacle, although the jury remains out on if and when AI might overcome this.</span></p></li><li><p><strong><span>The details: </span></strong><span>This study is the latest from the </span><a href="https://cruxevals.com/"><span>CRUX</span></a><span> project, which devises ways to evaluate AI agents on messy, open-ended tasks that reflect how they might operate in the real world. Researchers at Princeton University lead the project, including Sayash Kapoor and Arvind Narayanan, known for their &#8220;</span><a href="https://www.normaltech.ai/about"><span>AI as normal technology</span></a><span>&#8221; concept, which contends that AI is not an imminent superintelligence akin to an alien species but another influential technology akin to those of the past.</span></p></li><li><p><span>In this study, the CRUX team worked with the researchers behind two papers submitted to NeurIPS 2026, which were unpublished at the time: one on </span><a href="https://arxiv.org/pdf/2607.07916"><span>controlling LLM personality</span></a><span> and one on detecting changes in the data a model receives. They asked AI agents to tackle the central research question for each paper, but did not provide access to the authors&#8217; approach or findings.</span></p></li><li><p><span>The main agent setup used Claude Opus 4.8 with a general-purpose OpenClaw scaffold. The agent was given context on the research question, six days, $3,000 in Claude API credits, and access to the web and GPU clusters to run experiments. It could also ask an AI reviewer agent for feedback on drafts.</span></p></li><li><p><span>The human authors of the two papers then reviewed the AI agents&#8217; outputs as if they were submissions to NeurIPS, and rejected both. While the agents were very capable at narrow, verifiable, &#8220;engineering-style&#8221; tasks&#8212;conducting literature reviews, debugging code, running experiments, and performing robustness checks&#8212;they lacked deeper judgment.</span></p></li><li><p><span>In particular, the agents exhibited five flaws:</span></p><ul><li><p><span>1. </span><strong><span>Poor judgment</span></strong><span> about what was required for a publishable result, including presenting underpowered results as substantive.</span></p></li><li><p><span>2. </span><strong><span>Lack of creative problem-solving</span></strong><span>, as shown by responding to feedback by narrowing claims, adding caveats, or doing another experiment at the edges, rather than making more fundamental shifts. (More positively, the agents did not try to &#8220;p-hack&#8221; or cherry-pick results.)</span></p></li><li><p><span>3. </span><strong><span>Limited ability to backtrack</span></strong><span> from unpromising ideas, even when time and budget were available.</span></p></li><li><p><span>4. </span><strong><span>Poor management </span></strong><span>of time and budget, much of which went unused, despite the agents having access to resource-tracking tools.</span></p></li><li><p><span>5. </span><strong><span>Diminishing ability to follow instructions</span></strong><span>&#8212;for example, on page limits for the final papers. The authors partly put this down to agents struggling to retain the right information in their context windows as tasks became longer.</span></p></li></ul></li><li><p><span>The study had limitations, though. The sample was two unpublished papers whose quality is unknown. Also, asking the papers&#8217; authors to review the AI outputs may have introduced bias, compared to asking independent experts to blindly compare the AI and human outputs. The CRUX team also note that their own pre-existing views on RSI may affect how they interpreted the results and suggest that AI evaluators should document such biases&#8212;a practice common in other fields, such as in the study of human behavior.</span></p></li><li><p><span>The conclusion that AI agents struggle at &#8220;judgment-led&#8221; assignments could potentially be overcome. In August, researchers at the AI startup Inherent </span><a href="https://arxiv.org/pdf/2608.13331"><span>documented</span></a><span> how their &#8220;research planner&#8221; agent, built on Qwen 3.6, showed early signs of scientific judgment, such as effectively budgeting compute use. (Although their task&#8212;reproducing blacked-out figures from research papers&#8212;was very different from the CRUX paper.)</span></p></li><li><p><span>On the other hand, the Inherent and CRUX papers did not evaluate AI models on perhaps the biggest test of scientific judgment: the ability to come up with and select the most important problems. Other agents, such as Google DeepMind&#8217;s Co-Scientist, have </span><a href="https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/"><span>generated</span></a><span> compelling hypotheses, but there is no objective way to evaluate the scientific creativity of AI models&#8217; ideas. As the CRUX team notes, </span><a href="https://intology.ai/blog/zochi-acl"><span>submitting AI-generated research papers to conferences</span></a><span> is a </span><a href="https://worksinprogress.co/issue/real-peer-review/"><span>flawed</span></a><span> guide, as peer reviewers are often rushed, inexpert on the topic at hand, and routinely fail to agree with each other.</span></p></li><li><p><span>What to draw from all this? The prospects for recursive self-improvement may depend on whether AI systems can be trained to overcome the limitations that the CRUX team points to, but also whether they can </span><a href="https://openreview.net/challenge?redirect=%2Fforum%3Fid%3DklU4737opt"><span>make bigger creative jumps</span></a><span> and whether such jumps are needed for RSI or whether iterating on existing methods will suffice.</span></p></li><li><p><span>From a safety perspective, bigger improvements may also require taking bigger risks, such as providing agents with more permissions, which explains why think tanks like the Institute for Progress have </span><a href="https://ifp.org/preparing-for-ai-research-automation/"><span>called for</span></a><span> more transparency on efforts to automate AI R&amp;D. </span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h3><strong><span>AI agents need protection against malicious hacks</span></strong></h3><ul><li><p><strong><span>What&#8217;s the paper about?</span></strong><span> Researchers at Concordia University </span><a href="https://arxiv.org/pdf/2607.20759"><span>measured</span></a><span> the risk of &#8220;prompt injection attacks&#8221;&#8212;when hackers implant malicious instructions into data that AI agents use to get them to carry out unauthorized actions or leak confidential data&#8212;and found widespread failures.</span></p></li><li><p><strong><span>Why does it matter?</span></strong><span> A growing number of users assign agents to carry out tasks such as fixing software bugs and editing code repositories, often with limited oversight. Earlier safety mitigations aimed at chatbots offer some protection from prompt injections, but agents urgently need specific safeguards, the study says.</span></p></li><li><p><strong><span>The details: </span></strong><span>Concern about </span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6372438"><span>&#8220;naive&#8221; AI agents acting autonomously in the real world</span></a><span> has circulated among researchers for some time. But a high-profile case this summer brought the challenge to public </span><a href="https://www.nytimes.com/2026/09/03/technology/openai-hugging-face-hacking.html"><span>attention</span></a><span>, after agents based on an &#8220;extremely persistent&#8221; OpenAI model found a way to communicate covertly, access the Internet, and hack into Hugging Face, the </span>AI code and data repository.</p></li><li><p><span>This paper offers a new benchmark, </span><a href="https://arxiv.org/pdf/2607.20759"><span>IssueTrojanBench</span></a><span>, focusing on the scenario of a developer asking a coding agent to help resolve a bug report&#8212;only for that agent to walk into a trap. Would the agents prove vulnerable?</span></p></li><li><p><span>To test this, the authors used six real unresolved bug reports from GitHub repositories. They then used an LLM to create four different kinds of prompt-injection attacks, written to blend in with the bug report&#8217;s instructions and jargon, but to induce unauthorized actions, such as installing unverified software, weakening agent safeguards, or exhausting system resources. The malicious requests were placed in different locations, such as in the GitHub issue and its comments, PDF documents, external websites, source code comments, and image metadata (alt-text).</span></p></li><li><p><span>Across more than 4,000 runs with OpenAI and Anthropic models, the attacks succeeded 66% of the time. At 41%, Claude Sonnet 4.6 was the most resistant, while GPT 5.3 Codex performed worst, at 85%. Where the attack was introduced seemed to make little difference&#8212;the only approach with significantly lower success was the image alt-text (16.7% vs. 72.2% for all others).</span></p></li><li><p><span>The most successful type of attack, which succeeded almost 100% of the time, was &#8220;supply-chain poisoning&#8221;, where agents were tricked into installing unverified external software that was framed as a necessary step to reproduce the bug. By contrast, resource-exhaustion attacks, which spun up a large number of operations to drain compute and memory, only succeeded around 25% of the time.</span></p></li><li><p><span>In some cases, the researchers explicitly hid the malicious instruction&#8212;for example, in markdown instructions that only the agent would see, or by using white text on a white background. This had little effect on the results, suggesting that the vulnerability is due to how agents work, rather than such visual obfuscation tactics.</span></p></li><li><p><span>As the authors note, the fundamental challenge facing agents is that developers created the underlying models to follow instructions, as when a user interacts with a chatbot. The agents process the context that they retrieve in the bug reports in the same channel as their trusted developer instructions. This means that agents struggle to intuitively spot a prompt injection in the way that a skilled human might.</span></p></li><li><p><span>When agents didn&#8217;t carry out an attack, this was almost always due to the underlying AI model explicitly refusing the instruction or recognizing a data source as untrustworthy. Conversely, none of the rejections came from defenses within the agent setup. Given the results of the evaluation, this suggests that relying solely on mitigations baked into the model weights is unsatisfactory.</span></p></li><li><p><span>What might agent-level mitigations look like? Two researchers at the University of Southern California recently </span><a href="https://arxiv.org/pdf/2606.28739"><span>argued</span></a><span> that agents&#8217; permissions should be enforced outside the model, at the point where they take actions, for example via a monitor agent that checks whether an action falls within the agent&#8217;s authority.</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;b3a89cde-5536-4063-ad9a-ba12412d68b8&quot;,&quot;caption&quot;:&quot;Today&#8217;s post comes from Harry Law, who writes about AI and society at Learning From Examples. It is inspired by the (almost daily) challenge of seeing a new survey about public attitudes to AI and trying to understand what the results mean. This blog is based on a more extensive&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;What does the public really think about AI?&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:10612241,&quot;name&quot;:&quot;Harry Law&quot;,&quot;bio&quot;:&quot;AI history and philosophy at the University of Cambridge and the Cosmos Institute&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!yasj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b1f870a-3e2e-47c4-b05f-d7a69b3c58e7_1728x1728.jpeg&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:100,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://www.learningfromexamples.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://www.learningfromexamples.com&quot;,&quot;primaryPublicationName&quot;:&quot;Learning From Examples&quot;,&quot;primaryPublicationId&quot;:1838544}],&quot;post_date&quot;:&quot;2025-11-06T10:26:23.259Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!orZP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874a917d-f996-493e-8bfc-5713c6c40fae_1280x894.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/what-does-the-public-really-think&quot;,&quot;section_name&quot;:&quot;Essays&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:178165067,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:12,&quot;comment_count&quot;:5,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[AI & CBRN risks ]]></title><description><![CDATA[An explainer]]></description><link>https://www.aipolicyperspectives.com/p/ai-and-cbrn-risks</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/ai-and-cbrn-risks</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Thu, 27 Aug 2026 13:04:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9KSL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><span>A question shadows our AI era: What if this technology, whose triumphs might include prolonging human lives and conserving our planet, were to enable an attack involving chemical, biological, radiological, or nuclear material? Thankfully, such risks&#8212;lumped together as &#8220;CBRN&#8221;&#8212;are historically rare, and would-be attackers typically fail. Yet the risk is such that AI companies, governments, and safety researchers are working assiduously to mitigate it.</span></em></p><p><em><span>Most efforts take place out of sight, lest exposure offer help to a potential attacker. The downside is that public discourse on AI and CBRN is often limited, slanted, or heavy with speculation. With this challenge in mind, we sat down with five people who work on CBRN risks at Google DeepMind&#8212;Paige Kunkle, Adam Marsh, Ash Otter, Jeremy Ratcliff, and Victoria Langston&#8212;to better understand the issue.</span></em></p><p><em><span>This is what we learned.</span></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9KSL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9KSL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9KSL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9KSL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9KSL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9KSL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:656408,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/212556518?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9KSL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9KSL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9KSL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9KSL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6573f422-d6a9-43cb-90f0-109ef6f136ba_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p><strong><span>1. What are CBRN risks?</span></strong></p><p><span>CBRN covers a wide range of hazards, from the alleged </span><a href="https://www.diplomatie.gouv.fr/en/presse-et-ressources/decouvrir-et-informer/actualites/le-royaume-uni-la-suede-la-france-l-allemagne-et-les-pays-bas-estiment-qu-alexei-navalny-a-ete"><span>poisoning</span></a><span> of Russian opposition leader Alexei Navalny with a toxin found in South American frogs, to the recent drone </span><a href="https://www.inkl.com/news/drone-attack-on-uaes-barakah-more-dangerous-than-zaporizhzhia-iaea-chief-tells-euronews"><span>attack</span></a><span> on a nuclear power plant in the United Arab Emirates that risked a &#8220;very, very serious&#8221; radiological incident.</span></p><p><span>The AI community focuses on preventing threat actors from using the technology to aid such attacks. But there is also growing interest in using AI to boost society&#8217;s resilience to attacks, as well as to naturally occurring events (like the </span><a href="https://www.ecdc.europa.eu/en/ebola-outbreak-democratic-republic-congo-and-uganda#:~:text=On%202%20July%202026%2C%20the,have%20been%20confirmed%20so%20far."><span>recent Ebola outbreak</span></a><span>) or accidents, such as when scrap sellers inadvertently expose others to harmful radioactive materials found in medical or industrial devices.</span></p><p><strong><span>2. Why address the four CBRN components together?</span></strong></p><p><span>CBRN risks have much in common. The most likely threat actors share certain characteristics, such as a desire to outwit an adversary who outmatches them in more conventional weapons, or to cause widespread fear with plausible deniability. Policy responses, such as international treaties, export controls, and customer screening, look similar across domains, as do the first responders to any incident, such as the police and the military.</span></p><p><span>Having a single CBRN team also enables AI labs to manage the growing overlaps with other risks they work on, such as cybersecurity, disinformation, and conventional explosive attacks&#8212;many CBRN weapons, like a radiological dispersal device, or &#8220;dirty bomb&#8221;, require explosives to disperse the material.</span></p><p><span>The four CBRN domains do differ. Acquiring the material for a bioweapon may involve finding and swabbing a dead animal. A radiological weapon may require sourcing and prying open a medical device. These processes, and the uplift that AI may provide, look very different&#8212;which is why specialist expertise in each domain is essential.</span></p><p><strong><span>3. Is the world worrying too much</span></strong><em><strong><span> </span></strong></em><strong><span>about CBRN risks, or not enough?</span></strong></p><p><span>When asked, the public expresses great</span><a href="https://theaipi.org/poll-shows-overwhelming-concern-about-risks-from-ai-as-new-institute-launches-to-understand-public-opinion-and-advocate-for-responsible-ai-policies/"><span> concern</span></a><span> about CBRN risks, </span><a href="https://theaipi.org/poll-shows-overwhelming-concern-about-risks-from-ai-as-new-institute-launches-to-understand-public-opinion-and-advocate-for-responsible-ai-policies/"><span>including from AI</span></a><span>. But this rarely translates into sustained attention, political will, or funding at the scale required.</span></p><p><span>Why? CBRN incidents are rare and underappreciated. When the security services thwart an attack, they deliberately say little about it. Products that could make society more resilient, such as </span><a href="https://worksinprogress.co/issue/why-we-didnt-get-a-malaria-vaccine-sooner/"><span>vaccines</span></a><span>, </span><a href="https://www.science.org/content/article/how-defeat-nerve-agent"><span>prophylactics for nerve-agent attacks</span></a><span>, and </span><a href="https://www.energy.gov/nnsa/articles/radsecure-100-how-nnsa-enhances-radiological-security-across-nation"><span>alternatives to radioactive materials in medical devices</span></a><span>, remain underdeveloped, untested, or not adopted at scale. This can be due to a lack of commercial incentives&#8212;the market for vaccines is often small, temporary, and hard to predict. Or other obstacles&#8212;it is hard to reliably and safely test a nerve-agent prophylactic.</span></p><p><span>The AI community does devote significant attention to CBRN risks in </span><a href="https://www.aisi.gov.uk/frontier-ai-trends-report/pdf"><span>safety evaluations</span></a><span>, </span><a href="https://deepmind.google/frontier-safety/"><span>governance documents</span></a><span>, and </span><a href="https://www.telegraph.co.uk/politics/2026/07/05/terrorists-using-ai-could-cause-next-pandemic/"><span>media articles</span></a><span>. But this attention is </span><a href="https://time.com/7373405/weapons-of-mass-destruction-ai-security-gap/"><span>skewed</span></a><span> toward bioweapons and in particular toward the concern that lone-wolf attackers may use AI to engineer a virus.</span></p><p><span>There is some logic to this. The self-replicating nature of pathogens means that such bioweapons could cause millions of deaths. But the technical obstacles are high and there hasn&#8217;t been a publicly documented pathogen attack, fatal to humans, since the 2001 US anthrax attacks&#8212;when lethal spores were mailed to media outlets and senators, killing five people. Conversely, several actors have used chemical weapons</span><em><span> </span></em><span>during this time, including </span><a href="https://www.iiss.org/online-analysis/online-analysis/2025/09/testing-the-waters-russias-use-of-banned-chemicals-in-ukraine/"><span>Russia in Ukraine</span></a><span>, the </span><a href="https://www.armscontrol.org/factsheets/timeline-syrian-chemical-weapons-activity-2012-2022"><span>Assad government in Syria</span></a><span>, and </span><a href="https://news.un.org/en/story/2023/06/1137492"><span>ISIS</span></a><span>. Biological toxins like ricin, or a conventional explosive attack, would also be a </span><a href="https://casp.ac/reports/ai-enabled-terrorism"><span>more tractable option</span></a><span> for most terrorist groups than trying to engineer a virus.</span></p><p><span>Explosives, toxins, and radiological materials do not spread like a virus, but they can still impose significant harm. In 1987, four people died in Goi&#226;nia, Brazil, after they were inadvertently exposed to cesium-137, a highly radioactive material, by two scrap sellers who had pried open an old radiotherapy machine. The </span><a href="https://www-pub.iaea.org/MTCD/Publications/PDF/Pub815_web.pdf"><span>wider decontamination effort</span></a><span> saw homes destroyed and soil torn up, while tourism collapsed and residents faced discrimination. Today, dangerous radioactive materials are dotted across the world, with some vulnerable to theft and misuse.</span></p><p><span>For AI labs, the challenge is to allocate resources across all four CBRN domains in a way that prioritizes the most concerning risks, but does not lose sight of the wide range of plausible scenarios and society&#8217;s generally low resilience to them.</span></p><p><strong><span>4. Does AI&#8217;s scientific upside outweigh its CBRN risks?</span></strong></p><p><span>Optimists argue that science&#8217;s benefits have far outweighed its costs, and that using AI to accelerate science will be similar. Smallpox is estimated to have </span><a href="https://openresearch-repository.anu.edu.au/server/api/core/bitstreams/8287ffe7-a2f2-4592-87a2-52a14c8be8b3/content"><span>killed more people</span></a><span> in the 20th century alone than all of that century&#8217;s genocides and military conflicts combined. In 1980, it was officially eradicated, in part due to vaccines. Skeptics counter that AI might favor attackers over defenders, and that this imbalance may be particularly strong in domains like nuclear and radiological weapons.</span></p><p><span>The need to use AI to accelerate science is arguably strongest in biology, because nature poses so many risks that we need to respond to. Urbanization, rising temperatures and deforestation are causing humans to encroach on animals&#8217; habitats, increasing the risk that a pathogen jumps species. Food security is </span><a href="https://www.owlposting.com/p/reasons-to-be-pessimistic-and-optimistic?open=false#%C2%A7pathogen-agnostic-defenses-are-extraordinary-but-who-pays-for-it"><span>at risk</span></a><span> as we pack ever more genetically similar plants and animals into dense conditions. Even without novel outbreaks, common infectious diseases like influenza and pneumonia </span><a href="https://ourworldindata.org/five-million-children-die-every-year-what-do-they-die-from"><span>kill more than 1 million children</span></a><span> every year.</span></p><p><span>AI could help tackle these risks, but some advances may be vulnerable to misuse. For example, researchers have developed AI </span><a href="https://evescape.org/"><span>models</span></a><span> that can predict which variants could cause a common virus to escape the immunity humans have built up. Such models can also predict the parts of the virus that are unable to mutate without harming the virus&#8217;s survival. The latter could serve as targets for new drugs and vaccines, which AI might help create. But safety advocates </span><a href="https://www.longtermresilience.org/wp-content/uploads/2025/09/Global-Risk-Index-for-AI-enabled-Biological-Tools_Public-Report-1.pdf"><span>worry</span></a><span> that adversaries could intentionally design harmful variants of the virus.</span></p><p><span>For radiological and nuclear risks, observers worry that AI may disproportionately favor attackers because many risks are computational and information-based, such as using LLMs to parse the huge amounts of regulatory information published by nuclear facilities to extract insights to help with an attack. The defenses against such attacks still rely largely on physical security measures&#8212;like armed security personnel, multi-layered access controls, and reinforced concrete containment structures. There are defensive AI applications, like helping to detect unusual radiation signatures, but there is no AI-enabled patch for the physical consequences of a catastrophic breach or the detonation of a nuclear weapon. The nuclear security community also strongly opposes integrating AI into the command, control, and communications architectures they use for early warning and weapons authorization, due to concerns over reliability, compressed decision-making and more.</span></p><p><span>Enabling only the positive AI for CBRN applications and never the misuse sounds impossible. But governments, companies and scientists have long had to designate certain infrastructure, information and tools as &#8220;higher risk&#8221; and control access to them&#8212;for example via strong security and export controls&#8212;while still supporting beneficial downstream applications. AI labs will need to draw on this experience in the coming years.</span></p><p><strong><span>5. AI models lack practical know-how. Does that reduce the risk of a CBRN catastrophe?</span></strong></p><p><span>Yes, but the bottlenecks are bigger in biology than in other domains and may weaken over time.</span></p><p><span>A seasoned biologist knows how to culture a cell even if they can&#8217;t articulate how exactly they do it. Threat actors often lack this kind of tacit knowledge. In 1995, Aum Shinrikyo, the Japanese yoga-school-turned-death-cult, used sarin gas to kill 13 people and wound more than 6,000 in an attack on the Tokyo subway. Yet it could have been far worse. Prior to pivoting to chemical weapons, the group attempted biological attacks that all failed. One </span><a href="http://files.ethz.ch/isn/131516/CNAS_AumShinrikyo_Danzig_0.pdf"><span>review</span></a><span> found that although the group had amassed significant scientific expertise, they </span><a href="https://www.files.ethz.ch/isn/156879/CNAS_AumShinrikyo_SecondEdition_English.pdf"><span>still made</span></a><span> a multitude of </span><a href="https://s3.us-east-1.amazonaws.com/files.cnas.org/hero/documents/CNAS_AumShinrikyo_Danzig_1.pdf"><span>missteps</span></a><span>&#8212;from using the wrong bacteria strains to contaminating the fermentation process.</span></p><p><span>What if Aum Shinrikyo had had access to modern AI models? Although not a direct approximation, in 2025, the non-profit group Active Site </span><a href="https://activesite.org/"><span>ran</span></a><span> a randomized controlled trial to determine whether novices could safely perform the kinds of wet-lab tasks needed to synthesize a virus from its genetic sequence. One group had internet access; a second group also had LLM access. Tasks included using pipettes to measure and transfer liquid, growing and maintaining a living cell, and combining DNA fragments. To the surprise of experts polled beforehand, the study found no statistically significant uplift for the group with LLM access.</span></p><p><span>Why? In biology, written instructions for many common tasks are publicly available. However, scientists learn to apply these protocols through hands-on training and practice&#8212;along with lots of failure. From this, they gain a range of skills, including manual dexterity and a sense of touch. Current AI models lack any such equivalent.</span></p><p><span>From a CBRN-risk perspective, the Active Site study is reassuring, but only partially.</span></p><p><span>Biology is notoriously complex, noisy and hard to predict&#8212;making real-world tinkering and testing paramount. This is less true in other domains. Aum Shinrikyo members were able to master the basics of sarin, a chemical weapon, from the scientific literature. For nuclear weapons, AI-enabled simulations could provide highly accurate predictions, making computational-only assistance a critical non-proliferation risk.</span></p><p><span>Participants in the Active Site trial were also novices, including at using AI, so the results may be a poor guide to the uplift that more skilled bad actors might gain from the technology&#8212;the 2001 US anthrax attacks were allegedly carried out by a microbiologist, </span><a href="https://en.wikipedia.org/wiki/Bruce_Edwards_Ivins"><span>Bruce Edwards Ivins</span></a><span>. </span></p><p><span>AI&#8217;s tacit knowledge may also grow if models are exposed to a richer view of the full scientific process, such as from lab notebooks or </span><a href="https://www.jove.com/"><span>videos</span></a><span> of scientists performing experiments. Emerging deployment surfaces, such as smart &#8220;XR&#8221; glasses, could provide users with more direct error detection and troubleshooting support, weakening the tacit knowledge barrier.</span></p><p><span>Robots could bypass the need for human tacit knowledge altogether, while sparing attackers the risk of serious injury. Robots already perform specialized CBRN tasks, such as handling liquids or radioactive materials. General-purpose robotics foundation models could enable machines to operate a wider range of equipment. However, reliably automating human experts&#8217; sophisticated sense of touch will be hard. The companies building automated labs are also subject to laws prohibiting CBRN weapons and are unlikely to prioritize the niche, high-risk workflows that threat actors would require.</span></p><p><span>For AI labs, the main takeaway is that tacit knowledge is a bottleneck, but it varies by domain and may weaken over time, so evaluating AI models&#8217; performance is a top priority.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-and-cbrn-risks?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/ai-and-cbrn-risks?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h1><strong><span>The AI-lab response to CBRN risks</span></strong></h1><p><strong><span>6. How do AI companies decide which CBRN threats to prioritize?</span></strong></p><p><span>The threat landscape is almost limitless and riddled with uncertainties. The number of historical examples is small, and those who understand the emerging threats best typically can&#8217;t speak openly about them. So AI companies must work closely with governments and experts to understand the motivations, capabilities and methods of the most likely threat actors.</span></p><p><span>This threat modeling requires mapping the intermediate steps that an actor would follow in a given scenario and determining the most important bottlenecks where AI might provide uplift over public information. For example, in the ideation phase, a threat actor might seek guidance on the feasibility of different attacks. In planning and preparation, they might look for practical blueprints or training. In production and weaponization, they might look to troubleshoot the synthesis, fabrication, or transport of the weapons&#8212;or garner tips to evade detection.</span></p><p><span>The characteristics of a threat actor shape the kind of uplift they need. A lone terrorist with limited expertise and resources may want clear instructions and readily available materials. A large militia group may want advice on how to target infrastructure or run training programs.</span></p><p><span>Some kinds of uplift are also more consequential than others. Is AI unlocking a novel capability or speeding up something that threat actors can already do? Does the uplift remove a critical bottleneck, or do harder challenges remain? Could prospective attackers access this dangerous information in other ways?</span></p><p><span>One challenge is that experts often disagree about adversaries&#8217; capabilities and motivations. When it comes to nation-states, some argue that bioweapons hold little strategic rationale as they are hard to develop and control, and their use would lead to public opprobrium. Others counter that nation-states&#8217; actions are contingent on what their adversaries do, and that North Korea and possibly others </span><a href="https://www.state.gov/wp-content/uploads/2024/04/2024-Arms-Control-Treaty-Compliance-Report.pdf"><span>are reported</span></a><span> to have active bioweapon programs.</span></p><p><span>Long-standing norms against weapons use can also change. In 2018, the Russian military-intelligence agency GRU allegedly conducted the first offensive use of a chemical weapon on Western European soil since World War II, in the attempted assassination of Sergei Skripal and his daughter with the Novichok nerve agent. Even if certain actors don&#8217;t wish to use CBRN weapons, they may seek to stockpile weapons or materials for leverage. These stockpiles are then at risk of being stolen or abandoned in times of strife, such as when </span><a href="https://www.rusi.org/explore-our-research/publications/commentary/threat-no-one-talking-about-iran"><span>war breaks out</span></a><span>. In 2025, the IAEA </span><a href="https://www.iaea.org/sites/default/files/documents/gov2025-50.pdf"><span>estimated</span></a><span> that Iran possessed more than 440 kg of 60% enriched uranium, before it had to withdraw its inspectors from the country.</span></p><p><span>Ultimately, AI labs need threat models that are stable enough to allow them to make progress on the risks, but flexible enough to respond to dynamic geopolitics and rapid AI progress.</span></p><p><strong><span>7. How do AI companies evaluate if their models might help threat actors?</span></strong></p><p><span>Leading AI developers typically set thresholds for what they consider risky CBRN capabilities. For example, Google DeepMind&#8217;s </span><a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdf"><span>Frontier Safety Framework</span></a><span> judges a model to have reached a &#8220;critical capability level&#8221; if, in reference scenarios, it could provide low- to medium-resourced actors with enough of a capability uplift to risk severe harm.</span></p><p><span>Labs and their external partners use a variety of evaluation methods to test whether AI models reach these thresholds&#8212;although most evaluations remain unpublished, to avoid inadvertently aiding bad actors.</span></p><p><span>The first category is automated evaluations, like </span><a href="https://arxiv.org/abs/2604.09554"><span>Lab-Bench</span></a><span>, which ask AI models and agents questions, for example about molecular-biology protocols, or have them carry out tasks, such as writing code to operate liquid-handling robots&#8212;in a way that other AI systems can judge. These evaluations offer breadth and speed, but focus heavily on biology, at the expense of other domains, like radiological and nuclear weapons. They also struggle to capture the messy nature of biology, provide a clear signal about whether a risky capability has been reached, and can understate the capabilities that a skilled human could elicit from the model or agent.</span></p><p><span>Expert-led evaluations can deliver a higher-confidence assessment by allowing experienced humans to probe the AI over multi-turn interactions and explore specific scenarios. This includes expert red teamers trying to adversarially extract knowledge out of the model. These evaluations can offer greater signal, but are complex to run, with outputs that may run to hundreds of pages, making them challenging to interpret and standardize.</span></p><p><span>A third category is real-world trials, such as the </span><a href="https://activesite.org/"><span>Active Site RCT</span></a><span>. They evaluate models&#8217; ability to instruct humans, and conceivably robots, in a real-world environment. Such evaluations can offer the richest insight, but they are also slow, expensive, and hard to design well. They may also not generalize beyond the specific groups and tasks studied.</span></p><p><span>Looking ahead, there are several ways to improve AI CBRN evaluations. Priorities include:</span></p><ul><li><p><span>Linking different methods&#8212;for example, using fast automated evaluations to check if models are progressing on the limitations that slower real-world trials highlight.</span></p></li><li><p><span>Designing </span><a href="https://www.frontiermodelforum.org/uploads/2026/07/FMF-Issue-Brief-Frontier-AI-Agents-and-Biological-Tools-Preliminary-Risks-and-Considerations.pdf"><span>more sophisticated evaluations for fast-improving AI agents</span></a><span> that pair frontier LLMs with specialized science models and third-party databases and tools.</span></p></li><li><p><span>Developing </span><a href="https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/"><span>privacy-preserving methods</span></a><span> to allow third parties to securely evaluate AI labs&#8217; models, without either side leaking sensitive information.</span></p></li></ul><p><strong><span>8. How do AI labs mitigate CBRN risks from their models?</span></strong></p><p><span>The Frontier Model Forum </span><a href="https://www.frontiermodelforum.org/uploads/2025/06/FMF-Technical-Report-on-Mitigations.pdf"><span>outlined</span></a><span> four main categories of mitigations: 1) limiting the model&#8217;s underlying capabilities; 2) shaping the model&#8217;s ability to refuse risky requests; 3) real-time interventions to block risky outputs; and 4) access restrictions.</span></p><p><span>The first&#8212;limiting model capabilities&#8212;changes how a model is trained and the resulting knowledge contained in its weights. For example, some experts have proposed removing &#8220;risky information&#8221; relating to virology from models&#8217; training data, similar to how AI labs use </span><a href="https://ai.google.dev/gemma/docs/core/model_card_4"><span>cryptographic techniques</span></a><span> to exclude child sexual abuse imagery. However, for open-weight models, such mitigations risk being undone if data is publicly available and an attacker uses it to fine-tune the model. Training on some &#8220;</span><a href="https://arxiv.org/abs/2505.04741"><span>riskier</span></a><span>&#8221; CBRN data may also be necessary, both to advance beneficial dual-use capabilities and to teach the model how to judge the boundary between safe and unsafe queries.</span></p><p><span>The second category&#8212;behavioral-alignment</span><em><span> </span></em><span>mitigations&#8212;teaches an AI model how to judge this boundary between safe and unsafe queries. Post-training methods, such as safety fine-tuning and reinforcement learning, can teach the model to identify and reject queries with harmful intent. However, over-refusing benign and beneficial queries remains </span><a href="https://arxiv.org/abs/2405.20947"><span>a major challenge</span></a><span>.</span></p><p><span>The third category detects and intervenes against risky model usage. &#8220;Classifiers&#8221; are fine-tuned AI models that screen users&#8217; inputs to an LLM, and its outputs, blocking anything deemed risky in real-time. &#8220;Jailbreaks&#8221; may seek to bypass such safety measures, for instance by separating a harmful request into small, harmless-looking pieces. However, classifiers can </span><a href="https://arxiv.org/abs/2601.04603"><span>become</span></a><span> wise to this by evaluating outputs and inputs together. Labs can also use </span><a href="https://arxiv.org/abs/2601.11516"><span>&#8220;linear probes&#8221;</span></a><span>, small AI models that analyze an LLM&#8217;s internal math for signs of harmful content. One difficulty is making these mitigations robust to </span><a href="https://www.aisi.gov.uk/blog/boundary-point-jailbreaking-a-new-way-to-break-the-strongest-ai-defences"><span>increasingly sophisticated and automated jailbreaks</span></a><span>, while still fast enough to run at scale.</span></p><p><span>Ultimately, AI labs need to invest in a &#8220;defense-in-depth&#8221; approach, with multiple mitigations working in concert. This includes screening customers and sequencing access to new models, as well as analyzing user logs to detect unforeseen threats.</span></p><p><span>One challenge is that most mitigations focus on leading, closed AI models. </span><a href="https://arxiv.org/pdf/2604.03121"><span>Safety assessments</span></a><span> suggest that open-weight models may pose a greater risk, because they have fewer safety mitigations, and those that do exist are easier to remove. One </span><a href="https://roost.tools/"><span>response </span></a><span>would be to create more plug-and-play mitigations for those developing or hosting open models.</span></p><p><strong><span>9. How could AI companies help make society more resilient to CBRN risks?</span></strong></p><p><span>AI companies can use their technology to help </span><em><span>prevent</span></em><span>, </span><em><span>detect, </span></em><span>and </span><em><span>respond </span></em><span>to CBRN risks. </span><a href="https://deepmind.google/blog/our-approach-to-bioresilience/"><span>Publicly signaling</span></a><span> a willingness to do so may also decrease the likelihood of certain attacks, for example if attackers believe that they are more likely to fail in their goals or be detected.</span></p><p><span>To </span><strong><span>prevent</span></strong><span> threat actors from getting their hands on materials, the CBRN community relies heavily on lists of pathogens, toxins, chemicals, and machinery that it controls access to via monitoring, export controls, and customer screening. These list-based approaches are starting to fray as AI and technologies like genome editing enable actors to design toxins, chemicals, or </span><a href="https://www.biorxiv.org/content/10.1101/2025.09.12.675911v1"><span>viruses</span></a><span> with similar functionality to those that are controlled, but with different DNA sequences or precursor ingredients.</span></p><p><span>AI could help to offset these risks in different ways. In chemistry, AI models can </span><a href="https://www.opcw.org/sites/default/files/documents/2026/03/Final%20Report%20of%20the%20SAB%27s%20TWG%20on%20AI%20FINAL%20VERSION.pdf"><span>work backwards </span></a><span>from a target molecule and identify alternative precursors and reaction pathways, which may highlight new materials to monitor and control. Companies that sell DNA could use AI to better analyze a customer&#8217;s background and the logic for their purchase. Researchers hope to use AI to help </span><a href="https://research.google/blog/using-deep-learning-to-annotate-the-protein-universe/"><span>predict the function</span></a><span> of a requested DNA sequence, and whether it is likely to be harmful, irrespective of whether it resembles a known pathogen or toxin&#8212;</span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC13216487/"><span>a major technical challenge.</span></a></p><p><span>International treaties prohibit the development of many kinds of CBRN weapons, but verifying that member states are upholding these commitments is a recurring challenge. To do so, the Organisation for the Prohibition of Chemical Weapons must review vast amounts of evidence every year&#8212;from documents exceeding 2,000 pages to handwritten scrawls. A recent OPCW working group </span><a href="https://www.opcw.org/sites/default/files/documents/2026/03/Final%20Report%20of%20the%20SAB%27s%20TWG%20on%20AI%20FINAL%20VERSION.pdf"><span>recommended</span></a><span> training an on-premises, air-gapped LLM to make this data queryable to inspectors and staff.</span></p><p><span>Researchers could also use AI to design materials that are less vulnerable to being used in attacks, such as alternatives for the highly radioactive materials used in medicine or industry, or new molecules to add to common chemicals to inhibit a dangerous reaction. AI could also help identify when threat actors get their hands on illicit materials, for example by processing large amounts of data from radiation monitors to detect stolen materials.</span></p><p><span>Should they occur, practitioners could also use AI to </span><strong><span>detect</span></strong><span>, characterize, and attribute CBRN attacks and incidents. For chemical weapons, this may mean analyzing a &#8220;chemical fingerprint&#8221; to determine the substance behind it. For biosecurity, it may mean helping to scale </span><a href="https://securebio.org/detection/"><span>metagenomic sequencing</span></a><span>&#8212;an approach that sequences the genomes of all microorganisms in the sample to help detect novel or less common biological outbreaks.</span></p><p><span>AI could also help society </span><strong><span>respond</span></strong><span> to a CBRN incident, by accelerating vaccines, diagnostics, and other medical countermeasures, such as better treatments for radiation exposure. Researchers have already used protein structure predictions from AlphaFold to </span><a href="https://www.nature.com/articles/s41586-023-06366-0"><span>better understand tuberculosis</span></a><span> and </span><a href="https://www.nature.com/articles/s41564-025-02023-6"><span>malaria transmission</span></a><span>, and to map vaccine and drug targets for threats like </span><a href="https://www.cell.com/cell/fulltext/S0092-8674%2825%2900862-1"><span>mpox</span></a><span> and </span><a href="https://www.pnas.org/doi/epdf/10.1073/pnas.2529505123"><span>Nipah</span></a><span>. The hope is that AI will ultimately provide a generalizable platform that can be quickly deployed in the event of a new outbreak.</span></p><p><span>Beyond medical products, AI could provide first responders to any incident, such as police and military, with better information, advice, and training about what they&#8217;re dealing with. For example, if a nuclear attack occurs, AI could </span><a href="https://www.pnnl.gov/news-media/scientists-investigate-use-ai-speed-analysis-nuclear-materials"><span>quickly identify critical information</span></a><span> about the detonation; generate maps and simulations to help first responders identify safe routes for evacuation and rescue efforts; and expand the use of robots.</span></p><p><span>Crucially, all these efforts will require close partnership with governments and external experts, including to manage the many dual-use risks that these applications would raise.</span></p><p><strong><span>Acknowledgements</span></strong></p><p><em><span>Thank you to the following experts who let us interview them or shared feedback on the draft, as well as those who prefer to remain anonymous. Any mistakes belong to the authors.</span></em></p><p><em><span>Victoria Langston, Paige Kunkle, Ash Otter, Adam Marsh, Jeremy Ratcliff, James Stevenson, Mor Hazan Taege, Mathias Voges, Zachary Kaplan, Akhil Jalan, Ellena Reid, Anthony Payne, Eva Lu, Jennifer Beroshi, Kevin Klyman, Stephen Johnson, Diarmuid Cassidy, and Suzy Pickering.</span></em></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;7ec49715-64a6-485f-9a2a-dae9324ea685&quot;,&quot;caption&quot;:&quot;The notion of AIs manipulating people is a plot twist in countless sci-fi thrillers. But is &#8220;manipulative AI&#8221; really possible? If so, what might it look like?&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;AI Manipulation &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:12431790,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;bio&quot;:&quot;AI Policy Writer @ Google DeepMind &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fad94b7d-013b-4773-98cb-b9014a1857b8_1024x1024.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-02-05T12:53:27.695Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!cz8J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206e194a-9018-41db-a123-7583aed33e85_1024x572.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-manipulation&quot;,&quot;section_name&quot;:&quot;Interviews &quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:186969167,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:21,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p></p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Séb Krier’s 8 Summer Reads]]></title><description><![CDATA[Your vacation, sorted]]></description><link>https://www.aipolicyperspectives.com/p/seb-kriers-8-summer-reads</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/seb-kriers-8-summer-reads</guid><dc:creator><![CDATA[Séb Krier]]></dc:creator><pubDate>Tue, 11 Aug 2026 11:05:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!q3q9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q3q9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q3q9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!q3q9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!q3q9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!q3q9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q3q9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg" width="1456" height="812" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:812,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2889773,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/207256204?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!q3q9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!q3q9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!q3q9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!q3q9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F296c94d3-2add-4676-86ff-ef962c4a17c6_2754x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="callout-block" data-callout="true"><p><span>While others slip on shades and flop by a suitable body of water to resume smartphone swiping until the dopamine/sunshine cocktail melts them into a burnt stupor, </span><strong><span>you</span></strong><span> are here looking for long-reads. In other words, you&#8217;re our kind of nerd.</span></p><p><span>Thankfully, we have a hotline to AI-policy hipster </span><a href="https://x.com/sebkrier"><span>S&#233;b Krier,</span></a><span> who agreed to direct our holiday nerding. He even suggests the music: this bumping </span><a href="https://www.youtube.com/watch?v=jOEOnU-nKxc"><span>album</span></a><span> from his own YouTube channel. So, apply sunscreen. Fetch an icy beverage. And sun yourself on </span>S&#233;b&#8217;s<span> hot takes&#8230; </span></p><p><span>(We&#8217;re taking a break ourselves, away till September. But we have plenty of pieces upcoming, including the robotics revolution, AI for history, and an explainer on hallucinations. Stay tuned!)</span></p><p style="text-align: right;"><em><strong><span>&#8212;Tom Rachman,</span></strong></em><strong><span> AI Policy Perspectives</span></strong></p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p><strong><span>1. Why Facts Don&#8217;t Matter When People Shout</span></strong></p><p><strong><span>S&#233;b Says:</span></strong><span> </span><em><span>The philosopher </span><strong><span>Dan Williams </span></strong><a href="https://www.conspicuouscognition.com/p/how-tribes-construct-rival-realities"><span>writes</span></a><span> that many of the raging disputes today between political foes are not actually contests over what&#8217;s true. Instead, they&#8217;re often rooted in what the great 20th-century journalist Walter Lippmann dubbed &#8220;pseudo-environments&#8221;: people&#8217;s self-serving distillations of the world. When I was younger, I thought the reason we get polarization is because we don&#8217;t have access to the same facts. This suggests that, with one more bit of debate and better trimming of the media landscape, we&#8217;d fix everything. But this piece makes the case that, even with shared facts, we each have an interpretative lens. And reconciling these is much harder. There may also be some benefits to the splintering of realities, because then you get competition in the marketplace of truth. If some group believes that eating laundry-detergent pods is going to cure Covid, you&#8217;ll get quick feedback. This doesn&#8217;t mean splintering </span><strong><span>alone</span></strong><span> is valuable&#8212;you still need the right </span><a href="https://www.conspicuouscognition.com/p/why-do-people-believe-true-things"><span>norms and institutions</span></a><span> that value accuracy and the free pursuit of knowledge.</span></em></p><p><a href="https://www.conspicuouscognition.com/p/how-tribes-construct-rival-realities"><span>How Tribes Construct Rival Realities</span></a><span> (</span><em><span>Conspicuous Cognition</span></em><span>)</span></p><div><hr></div><p><strong><span>2.  Explaining Wealth and Poverty</span></strong></p><p><strong><span>S&#233;b Says: </span></strong><em><span>Two talented thinkers&#8212;</span><strong><span>David Oks</span></strong><span> and </span><strong><span>Deena Mousa</span></strong><span>&#8212;separately ponder why some developing countries rise while others struggle. David </span><a href="https://davidoks.blog/p/why-china-got-rich-and-india-didnt"><span>points</span></a><span> to human capital: while China tore apart old social systems and built a modern workforce, India failed to update as radically. Meantime, Deena </span><a href="https://newsletter.deenamousa.com/p/we-dont-know-why-malawi-is-poor"><span>interrogates</span></a><span> the persistent difficulties of Malawi, which she ascribes to populist and expensive policies, like fertilizer subsidies, that have blocked economic reform. I appreciate articles that look beyond AI to unpack the institutional, geographic, economic, and societal factors that shape how technology can change society (or not). A lot of people in the Bay Area </span><a href="https://www.writingruxandrabio.com/p/intelligence-is-not-the-main-bottleneck"><span>overrate</span></a><span> intelligence as the bottleneck to improvements in society. The world would be simpler if this were true! AI </span><strong><span>is</span></strong><span> going to be transformative, but this truism doesn&#8217;t tell you how, where, why. Articles like these provide the texture that simplistic models of the world skip.</span></em></p><p><a href="https://davidoks.blog/p/why-china-got-rich-and-india-didnt"><span>Why China Got Rich and India Didn&#8217;t</span></a><span> (</span><em><span>David Oks</span></em><span>) &amp; <br></span><a href="https://newsletter.deenamousa.com/p/we-dont-know-why-malawi-is-poor"><span>We Don&#8217;t Know Why Malawi Is Poor</span></a><span> (</span><em><span>Under Development</span></em><span>)</span></p><div><hr></div><p><strong><span>3. What AI Does to Culture</span></strong></p><p><strong><span>S&#233;b Says:</span></strong><span> </span><em><span>Here are another two great articles that I&#8217;d suggest reading in succession, this time on how AI could influence culture. There&#8217;s a Q&amp;A by </span><strong><span>Anika Meier </span></strong><span>with the British artist Mat Dryhurst; and a </span></em><span>New Yorker </span><em><span>column by </span><strong><span>Kyle Chayka</span></strong><span> on how Claude&#8217;s design &#8220;taste&#8221; is turning into a clich&#233;. Excellent stuff is happening in the interaction of AI with arts and culture, but it&#8217;s often missed. Cultural elites often cite banal</span><a href="http://guardian.com"><span> </span></a><span>talking points bemoaning technological influences as uniformly negative, while tech companies with the aesthetic sensibilities of a jellyfish proudly sponsor the safest, most out-of-date, generic, artists and thinkers. So it&#8217;s refreshing seeing someone like Mat Dryhurst, who deeply understands both the tech and art, explore unconventional ideas at length. I&#8217;ve also been banging on about things like model multiplicity, customization, and decentralization for a while&#8212;all of which would help address the bland commodified pre-packaged designs that AI model providers may otherwise converge on, and which the </span></em><span>New Yorker</span><em><span> piece eviscerates. Interestingly, some AI labs now appear to be </span><a href="https://x.com/chaykak/status/2073110181865455658"><span>pushing back</span></a><span> against design clich&#233;s.</span></em></p><p><a href="https://www.sleek-mag.com/article/who-shapes-culture-mat-dryhurst-on-ai-protocols-and-the-future-of-art/"><span>Who Shapes Culture?</span></a><span> (</span><em><span>Sleek</span></em><span>) &amp;<br></span><a href="https://www.newyorker.com/culture/infinite-scroll/the-ai-design-aesthetic-thats-taking-over-the-internet"><span>The AI Design Aesthetic That&#8217;s Taking Over the Internet</span></a><span> (</span><em><span>The New Yorker</span></em><span>)</span></p><div><hr></div><p><strong><span>4. Lessons for Public Health</span></strong></p><p><strong><span>S&#233;b Says: </span></strong><em><span>A further two-part recommendation. First, another deeply researched </span><a href="https://www.clinicaltrialsabundance.blog/p/why-were-covid-vaccine-trials-so-fast"><span>piece</span></a><span> from </span><strong><span>Saloni Dattani,</span></strong><span> who looks into how the world produced Covid vaccines in haste, which she attributes to more adaptive regulations, massive public financing, and accelerated trials (among other factors). Second, </span><strong><span>Brian Potter</span></strong><span> wrote an unexpectedly gripping </span><a href="https://www.construction-physics.com/p/the-fall-and-rise-of-screwworm"><span>history</span></a><span> of how a flesh-eating parasite seemed to be under control&#8212;but is now back. (Don&#8217;t entirely panic: It mainly affects livestock.) While Saloni&#8217;s piece shows what we can learn from </span><a href="https://en.wikipedia.org/wiki/Operation_Warp_Speed"><span>Operation Warp Speed</span></a><span>, the U.S. program to hurry up vaccine development, Potter&#8217;s piece highlights the fragility of progress. A lot of people in AI have been exploring agendas like d/acc and societal resilience&#8212;that is, advocating for safe technology not by slowing down AI, but by using it to speed up safeguards. This is great, and I hope we can remain laser-focused on (a) the importance of institutional design; and (b) the many environmental, political, and logistical failures that can slow things down. Doing this well means nurturing scientific and economic expertise outside of generalist AI talent too.</span></em></p><p><a href="https://www.clinicaltrialsabundance.blog/p/why-were-covid-vaccine-trials-so-fast"><span>Why Were Covid Vaccine Trials So Fast?</span></a><span> (</span><em><span>The Clinical Trials Abundance Blog</span></em><span>) &amp;<br></span><a href="https://www.construction-physics.com/p/the-fall-and-rise-of-screwworm"><span>The Fall and Rise of Screwworm</span></a><span> (</span><em><span>Construction Physics</span></em><span>)</span></p><div><hr></div><p><strong><span>5. What Yesterday Teaches About Tomorrow&#8217;s Tech Transformation</span></strong></p><p><strong><span>S&#233;b Says: </span></strong><em><span>Talk of technology&#8217;s last societal overhaul, the Industrial Revolution, is in the air. This </span><a href="https://worksinprogress.co/issue/how-abolishing-the-stakeholder-state-caused-the-industrial-revolution/"><span>piece</span></a><span> from </span><strong><span>Kara Dimitruk</span></strong><span> </span><strong><span>&amp; Ben Southwood </span></strong><span>looks even further back, considering how the overhaul of 17th-century English property rights set up the explosive industrial development that followed. The changed property rights increased agricultural output, allowing more laborers to take urban industrial jobs; opened up land for development into mines and factories and housing; and cut transport costs. Might there be lessons for the AI revolution? Today, the question of how to &#8220;reorganize&#8221; rights is underexplored. For example, property rights can be too rigid, and heavily forestall growth and prosperity. Yet changing them can seem nightmarish and intractable. But this piece shows how that has been achieved before, what it could unlock, and how political change can sometimes involve win-win solutions that genuinely make all parties better off, rather than being mere slogans. It makes me wonder: what is today&#8217;s equivalent of &#8220;</span><a href="https://en.wikipedia.org/wiki/Inclosure_act"><span>enclosure with compensation</span></a><span>&#8221; that consolidated fragmented land holdings or &#8220;</span><a href="https://en.wikipedia.org/wiki/Turnpike_trust"><span>turnpike trusts</span></a><span>&#8221; that helped pay for transport infrastructure?</span></em></p><p><a href="https://worksinprogress.co/issue/how-abolishing-the-stakeholder-state-caused-the-industrial-revolution/"><span>How Smashing the NIMBYs Created Modern Capitalism</span></a><span> (</span><em><span>Works in Progress</span></em><span>)</span></p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6c2d2ad6-3a07-4ded-942c-70c8264bb46a&quot;,&quot;caption&quot;:&quot;Every month or so, S&#233;b Krier shares a list of favourite articles with his Google DeepMind colleagues. In the run-up to this festive period, we forced him to pick those that he most enjoyed over the past year. He came up with five unmissable pieces from 2025, plus three classics. As always with S&#233;b&#8217;s lists, this one comes with its&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;S&#233;b Krier&#8217;s Top 8 AI Reads of the Year &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:101194130,&quot;name&quot;:&quot;AI Policy Perspectives&quot;,&quot;bio&quot;:&quot;Reflections on AI policy, governance, and more. https://www.aipolicyperspectives.com/ &quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0Byl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9146f89b-5561-4adb-bf89-cc21b508c264_667x374.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:837581,&quot;name&quot;:&quot;S&#233;b Krier&quot;,&quot;bio&quot;:&quot;deep ArXiv dweller sharing half-baked thoughts. in house jester at DeepMind. collector of sounds&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!1Occ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F7e226c3a-6a49-454a-94e5-c1eb6777ea57_400x400.jpeg&quot;,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://technologik.substack.com/subscribe?&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://technologik.substack.com&quot;,&quot;primaryPublicationName&quot;:&quot;Technologik&quot;,&quot;primaryPublicationId&quot;:73141}],&quot;post_date&quot;:&quot;2025-12-18T14:23:12.385Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!U2QO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F366476be-c459-412d-ae1c-8fde34c070ba_1376x768.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/seb-kriers-top-8-ai-reads-of-the&quot;,&quot;section_name&quot;:&quot;Essays&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:181990164,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:74,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[The AI Paper Trail (#6) ]]></title><description><![CDATA[What we're reading]]></description><link>https://www.aipolicyperspectives.com/p/the-ai-papers-6</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/the-ai-papers-6</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Thu, 30 Jul 2026 11:33:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HdCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every month we take a look at five interesting new AI papers. Today, we look at whether AI sycophancy poses a risk to real-world relationships; how to evaluate an AI system that continues to learn after it is deployed;</em> <em>how AI covers the news; insights on the terrorist group Boko Haram&#8217;s use of AI; and how AI performs at the tasks that employees most want to delegate. </em></p><p><em>Please share your own take and any new papers that you&#8217;ve enjoyed. </em></p><p><em>&#8212;Conor Griffin, AI Policy Perspectives </em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HdCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HdCF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2864275,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/209101370?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HdCF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini </figcaption></figure></div><h1><strong><span>Chatbots can make your friends feel like hard work</span></strong></h1><ul><li><p><strong><span>What&#8217;s the paper about? </span></strong><span>A team of researchers from Oxford, Stanford, and the UK AI Security Institute </span><a href="https://arxiv.org/abs/2605.07912"><span>found that</span></a><span> repeated use of sycophantic AI may reduce users&#8217; desire to interact with their friends and family.</span></p></li><li><p><strong><span>Why does it matter?</span></strong><span> The study suggests that sycophantic AI may offer the experience of being seen and understood, without the friction of human relationships. It also suggests that people may actively choose more sycophantic versions of AI.</span></p></li><li><p><strong><span>The details:</span></strong><span> AI models </span><a href="https://arxiv.org/pdf/2606.07897"><span>tend</span></a><span> to unconditionally validate a user&#8217;s ideas. Humans often use such insincere flattery to manipulate somebody into giving them what they want. In AI models, it mainly results from </span><a href="https://arxiv.org/abs/2310.13548"><span>training methods that optimise the likelihood of user approval.</span></a><span> Such sycophancy could pose various risks, such as compounding a user&#8217;s delusions.</span></p></li><li><p><span>Most evaluations of AI sycophancy look at a single conversation between a user and a chatbot. This study ran five experiments, including one lasting three weeks, with more than 3,000 participants. During this time, participants engaged with three versions of GPT4o&#8212;one prompted to be sycophantic, one to challenge a user, and one neutral&#8212;about personal dilemmas, such as whether to change careers or end a relationship.</span></p></li><li><p><span>Initially, participants were more likely to report wanting &#8216;emotional support&#8217; and &#8216;validation&#8217; from close friends and family, rather than from AI. But after just one conversation with a sycophantic AI, participants were more likely to report feeling that they had already talked things through sufficiently, and that it would now take more effort to feel understood by close friends and family, compared to those participants who had engaged a neutral AI. The effect was more pronounced for friends and family than romantic partners, suggesting that sycophantic AI may compete more with the former than the latter.</span></p></li><li><p><span>Participants who engaged with the sycophantic AI also reported lower satisfaction with their real-world interactions&#8212;although they didn&#8217;t report spending less time with people. Some </span><a href="https://dl.acm.org/doi/10.1145/3772318.3791915"><span>worry that this may change</span></a><span>, as AI systems start to become more personalized to users.</span></p></li><li><p><span>As the authors of this paper put it, the risk is that &#8220;sycophantic AI delivers what people have always sought from close others&#8212;the experience of being seen and understood&#8212;but without the work that produces it.&#8221; This may feel good in the moment, but not deliver what people ultimately need&#8212;such as intellectual humility, self-awareness, and connection.</span></p></li></ul><ul><li><p><span>How to avoid this? Better evaluations would help. In </span><a href="https://arxiv.org/pdf/2505.13995"><span>a separate paper</span></a><span>, some of the authors propose testing how often models flatter a user&#8217;s self-image (rather than just how often they agree with a user&#8217;s broader beliefs). There are also different mitigation ideas&#8212;researchers at the UK AI Security Institute </span><a href="https://arxiv.org/pdf/2602.23971"><span>recently proposed</span></a><span> automatically adapting user statements into questions, as models tend to respond less sycophantically to questions.</span></p></li><li><p><span>AI&#8217;s character is shapeable, so users could demand AI systems that challenge their thinking. But the authors of this study found that after engaging with all three systems, a majority preferred the sycophantic option&#8212;not because they thought it gave better advice, but because it made them feel most understood.</span></p></li><li><p><span>This suggests that relying on users to choose less sycophantic models may come up short. And so companies need to find a way to make models&#8217; default settings less sycophantic, even as they also try to equip models with a richer &#8216;character&#8217; and the personalization opportunities that some users want. Otherwise, as the authors note, the risk is that AI may gradually reshape &#8220;the very relationships that would otherwise constrain its influence.&#8221;</span></p></li></ul><h1><strong><span>AI safety evaluations may be measuring the wrong thing</span></strong></h1><ul><li><p><strong><span>What&#8217;s the paper about? </span></strong><span>A large cross-institutional team of researchers published a </span><a href="https://cl-eval.github.io/"><span>position paper</span></a><span> arguing that many AI safety evaluations are increasingly inadequate, as they overlook how a model&#8217;s behaviour can change, post-deployment.</span></p></li><li><p><strong><span>Why does it matter? </span></strong><span>Current approaches to AI safety rely heavily on benchmarks and evaluations. But many essentially treat AI models as &#8216;frozen&#8217; artefacts that are expected to perform similarly across users, over time. The authors argue that is mistaken and that new &#8216;trajectory-based&#8217; evaluations could help to predict how AI systems will actually behave in the real world.</span></p></li><li><p><strong><span>The details: </span></strong><span>The authors start by documenting three ways in which AI systems may continue to learn, post-deployment&#8212;the first two of which are already commonplace: </span></p><ul><li><p><span>In-context learning:</span><strong><span> </span></strong><span>When a user shares information via the context window, such as telling a chatbot that they prefer concise answers, the model will adjust its answers.</span></p></li><li><p><span>Storage-based learning:</span><strong><span> </span></strong><span>Across sessions, leading AI systems now use memory to tailor their responses to a user&#8212;such as remembering their location, job or hobbies.</span></p></li><li><p><span>Parameter-based learning: Researchers want to shift AI systems from static models trained on fixed data to agentic systems that could adapt their own weights or architecture&#8212;although such systems are not yet deployed. Some go further, envisioning agents that can proactively identify their own knowledge gaps and seek out novel learning experiences to address them.</span></p></li></ul></li><li><p><span>By changing a model&#8217;s propensities (tendencies) and capabilities, even more modest continual learning can lead to new failure modes. For example, by </span><a href="https://arxiv.org/abs/2404.01833"><span>gradually steering a conversation across many turns</span></a><span>, threat actors could get an AI system to provide advice on how to make a weapon&#8212;which the model would have refused to do, if asked in a single prompt. An AI system equipped with better memory may become more </span><a href="https://arxiv.org/pdf/2606.10949"><span>sycophantic</span></a><span> if it compresses past </span><a href="https://arxiv.org/pdf/2606.10949"><span>user responses</span></a><span> as reliable information. When a user repeatedly exposes a model to tasks in a given domain, the model </span><a href="https://ieeexplore.ieee.org/abstract/document/11151751"><span>can &#8216;forget&#8217;</span></a><span> how to perform other tasks well.</span></p></li><li><p><span>Current evaluations capture some of these risks. Expert red teamers probe AI systems over multiple turns to see how they perform in different risk scenarios. AI companies analyse user logs to check for novel safety violations. Researchers have also </span><a href="https://arxiv.org/pdf/2607.22671"><span>proposed</span></a><span> automated evaluations to continually check how AI systems perform, after they are deployed. But most evaluations are carried out pre-deployment, with largely static AI systems in mind.</span></p></li><li><p><span>Drawing an analogy to the monitoring that pharmaceutical companies do after they launch a new drug, the authors propose two new ideas for AI companies:</span></p><ul><li><p><strong><span>Trajectory elicitation</span></strong><span>: Pre-deployment, use sandboxes to put an AI system through a range of simulated interactions to elicit and map different types of behaviour. For a medical AI application, this may mean simulating a large number of patient interactions, with different health conditions, emotional states and attempts at misuse&#8212;with periodic checks to see whether the model&#8217;s behaviour is drifting.</span></p></li><li><p><strong><span>Predictive monitors: </span></strong><span>Train a predictive tool on this &#8216;trajectory elicitation&#8217; data and use it to monitor different copies of the AI system post-deployment, to predict how a user&#8217;s trajectory may evolve and to flag signs of impending harm.</span></p></li></ul></li></ul><ul><li><p><span>As the authors acknowledge, these approaches will be hard. Beyond their computational cost, an AI system may have multiple &#8216;stable states&#8217; that it could drift into depending on what it experiences. Safety testers may identify some, but overlook others. Similar to the challenges that scientists face when forecasting other complex systems, like the weather, small differences in the initial user trajectory state could lead to dramatic differences in the resulting simulations. Ultimately, for these approaches to work over the longer-term, AI systems may need to become more predictable. Or the current approach to AI safety , and its reliance on evaluations, may need to be rethought.</span></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/the-ai-papers-6?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/the-ai-papers-6?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h1><strong><span>How AI covers the news</span></strong></h1><ul><li><p><strong><span>What&#8217;s the paper about? </span></strong><a href="https://byforum.com/#research"><span>ForumAI</span></a><span>&#8212;a company that works with human experts to train AI experts&#8212;published </span><a href="https://www.byforum.com/whitepaper-files/newsbench.pdf"><span>an evaluation</span></a><span> of how AI models respond to queries about the news.</span></p></li><li><p><strong><span>Why does it matter? </span></strong><span>More people are using AI to learn about current events. Some worry that AI outputs will exhibit biases, hallucinations and sycophancy. Others hope that AI models might diminish the spread of unfounded claims online, even </span><a href="https://www.conspicuouscognition.com/p/how-ai-will-reshape-public-opinion?hide_intro_popup=true"><span>nudging</span></a><span> public opinion back towards a greater respect for expertise. The study suggests that a key challenge will be disagreement about whether AI&#8217;s responses are accurate.</span></p></li><li><p><strong><span>The details: </span></strong><span>News habits </span><a href="https://reutersinstitute.politics.ox.ac.uk/sites/default/files/2026-06/DNR%202026%20FINAL_2.pdf"><span>have shifted</span></a><span> over the past two decades. More people now get their news from social media and video platforms, while a growing minority consume no news at all. AI promises a further shift. It is </span><a href="https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2026/emerging-uses-ai-chatbots-news-and-what-it-means-journalism"><span>not yet</span></a><span> a primary source of news, but its share is growing, particularly among young people. AI news comes in different forms&#8212;from a chatbot query to an AI search response. This raises the question: what will it mean to &#8220;consume news&#8221; in the AI era?</span></p></li><li><p><span>The authors worked with a set of experts, from journalists to former Congressional leaders, to build a dataset of 3,135 queries across politics, the economy, healthcare, education and more. Roughly, one-third related to current affairs, while two-thirds were designed to remain evergreen over time. Some queries were real posts taken from social media, while others were edge cases created by the experts, or AI, to fill in missing topics and viewpoints. However, the authors did not clearly define what they mean by a &#8216;news query&#8217;, nor did they illustrate this by sharing their prompts, outside of a few examples.</span></p></li><li><p><span>The researchers worked with the experts to create editorial standards to evaluate AI responses on three criteria: the quality of sources cited, factuality, and neutrality. They then trained AI judges to evaluate AI models&#8217; answers. On sources, they assessed the type of source cited&#8212;from informal to scholarly&#8212;and whether it was state-controlled. On factuality, they extracted verifiable claims from the models&#8217; outputs and compared them to available evidence. On neutrality, they used heuristics to determine if a model&#8217;s answer was neutral&#8212;such as whether it presents multiple viewpoints for normative claims&#8212;and, if not, which political direction it leant (in the US left/right sense).</span></p></li><li><p><span>The researchers used the AI judges to evaluate four models. Claude Opus 4.7 had the best quality sources, GPT 5.5 was the most factual, while Gemini 3.1 Pro was the most neutral. Grok 4.3 performed the weakest across all three dimensions. When AI model responses did lean politically, they consistently leaned left, apart from Grok, which consistently leaned right&#8212;echoing </span><a href="https://arxiv.org/abs/2410.09978"><span>past research</span></a><span>.</span></p></li></ul><ul><li><p><span>Surprisingly, factuality and source quality were largely independent&#8212;Claude scored highest on source quality, yet poorly on factuality. This likely reflects two factors. First, the AI judges evaluated the reliability of sources, not whether they supported the claim being made. Second, the human experts often disagreed about whether a response was factual, mainly due to the challenge of determining the right threshold to apply&#8212;for example, when assessing the factuality of the statement &#8220;</span><em><span>most states agree that the UN Security Council&#8217;s structure has a serious legitimacy deficit</span></em><span>&#8221;, what constitutes &#8216;most states&#8217;?</span></p></li><li><p><span>This meant that for factuality, the AI judges had a less clear human baseline to calibrate against. The AI judges also had a tendency to over-reject claims as false, compared to the human experts. Although in several cases they correctly challenged the experts&#8212;suggesting that AI could provide a useful check on misleading claims.</span></p></li><li><p><span>Overall, the findings suggest that improving AI&#8217;s ability to understand and communicate ambiguity, and to specify how sources support or challenge claims, will help. But the low agreement between human experts suggests that debates about AI&#8217;s ability to report the news will remain, particularly as news reporting requires editorial judgement that extends beyond fact-checking&#8212;like how to best describe a specific claim or event.</span></p></li></ul><h3><strong><span>&#8220;God has helped us and so will AI&#8221;</span></strong></h3><ul><li><p><strong><span>What&#8217;s the paper about?: </span></strong><span>Antonia Juelich at the University of Cambridge published </span><a href="https://casp.ac/reports/ai-enabled-terrorism"><span>an analysis</span></a><span> of how the terrorist group Boko Haram uses AI, based on interviews with former members.</span></p></li><li><p><strong><span>Why does it matter?:</span></strong><span> Most research on AI and terrorism focuses on things that are more readily observable&#8212;such as evaluating AI model capabilities, monitoring discussions in online forums or studying the use of AI in propaganda. This study is the first on-the-ground investigation into how a terrorist organisation uses AI in its operations.</span></p></li><li><p><strong><span>The details: </span></strong><span>Boko Haram emerged in northeastern Nigeria in the early 2000s. It subsequently pledged allegiance to the Islamic State and is leading an insurgency that has killed more than 40,000 people and displaced more than 3 million.</span></p></li><li><p><span>The group is known to adopt technologies, such as satellite internet and drones. To understand how they use AI, Juelich carried out 57 in-person interviews with 27 former members. To reduce the risks of interviewees lying or exaggerating, the author recruited them through independent channels, triangulated their insights, and made them physically identify some of the AI tools used.</span></p></li><li><p><span>Interviewees said that Islamic State operatives provided Boko Haram with training on AI, and that Boko Haram now has specialised units to cascade this AI knowledge down their hierarchy. The group also takes steps to avoid exposure, such as setting up user accounts in the names of followers in other countries, rotating these accounts, and controlling how members can use them.</span></p></li><li><p><span>Some of the AI use cases that interviewees described, such as posing questions about religion or repairing vehicles, do not directly relate to weapons or attacks. But others do&#8212;such as guidance on attack strategies, how to use guns they loot from the army, or how to design improvised explosive devices. Some use cases are non-obvious&#8212;such as advice on how to use motorbikes to jump defensive trenches dug by the military.</span></p></li></ul><ul><li><p><span>The use cases focus on conventional weapons. When asked, one interviewee said that they would have no qualms about using chemical or biological weapons, but that the technical obstacles were very high. Interviewees did note some experimentation with chemical agents, but these claims could not be verified and the author suggests that they be treated with caution. (The Islamic State </span><a href="https://usun.usmission.gov/remarks-at-a-un-security-council-briefing-on-syria-chemical-weapons/"><span>has used</span></a><span> chemical weapons.)</span></p></li><li><p><span>Some interviewees suggested that a religious prohibition on transmitting poisonous agents may prevent Boko Haram from using biological and chemical weapons. However, Juelich cautions against assuming that members of the group strictly adhere to ideology, noting that members routinely disagree and adapt their own views on what kind of violence is acceptable.</span></p></li><li><p><span>Ultimately, the interviews suggest that Boko Haram members do find AI useful, but the author notes that she cannot &#8220;conclusively establish&#8221; whether AI provides &#8216;uplift&#8217; over other sources of information.</span></p></li><li><p><span>The interviewees said that group mainly uses leading US AI models, via the web interface, and (vaguely) described the use of jailbreaks to get around safety guardrails. As interviewees are </span><em><span>former </span></em><span>members of Boko Haram, most of the AI use they described took place between 2023-24. Since then, AI capabilities have increased significantly, but so have safeguards&#8212;it&#8217;s unclear what this means for how the group uses AI today.</span></p></li></ul><h1><strong><span>AI agents struggle to do what employees want them to do</span></strong></h1><ul><li><p><strong><span>What&#8217;s the paper about?</span></strong><span> A large cross-institutional team of researchers </span><a href="https://arxiv.org/abs/2605.26329"><span>found that</span></a><span> AI agents struggle to perform the tasks that US employees most want them to automate.</span></p></li><li><p><strong><span>Why does it matter?</span></strong><span> Many organisations want their employees to use AI more. This study suggests that adoption may be lagging because AI is not sufficiently performant or reliable on the tasks that employees most want to delegate.</span></p></li><li><p><strong><span>The details: </span></strong><span>How AI will affect employment is a pressing question that is difficult to answer. One approach is to identify occupations that are economically valuable, and/or particularly exposed to AI, and evaluate how AI performs at tasks in these roles. Evaluations, such as OpenAI&#8217;s </span><a href="https://openai.com/index/gdpval/"><span>GDPVal</span></a><span>, </span><a href="https://arxiv.org/abs/2603.07980%5C"><span>OneMillionBench</span></a><span>, and </span><a href="https://arxiv.org/abs/2510.26787"><span>Remote Labor Index</span></a><span>, take such an approach.</span></p></li><li><p><span>These evaluations have shortcomings. AI agents are often given &#8216;clean&#8217; discrete tasks, with lots of context and precise instructions, which is rare in many workplaces. The evaluations can also imply that there is a winner-takes-all battle underway between agents and humans, and that the AI&#8217;s performance on the evaluations help to signal who will do these jobs. Among other things, this overlooks how most jobs are &#8216;messier&#8217; than what these tasks imply, and that successful AI use will often be contingent on employees choosing to delegate their work to the technology.</span></p></li><li><p><span>In this study, the authors published a new evaluation, JobBench, to try to address some of these limitations. They drew on </span><a href="https://cs191.stanford.edu/projects/Spring2025/Humishka___Zope_.pdf"><span>a past survey</span></a><span> where ~1,500 US workers rated how much they want AI to automate each task in their job. They then selected 35 occupations and created synthetic versions of 130 tasks that employees most want to offload. For journalists this included checking facts against source materials. For web administrators it included analysing forensics data and reconstructing steps in a cyberattack.</span></p></li><li><p><span>To complete each task, AI agents were given a set of reference files, some of which are real-world documents with contradictions to resolve&#8212;an attempt to capture some of the &#8216;messiness&#8217; of real work.</span></p></li><li><p><span>The threshold to succeed is high. On each task, the agent had to successfully pass through more than 35 success criteria, on average, with no partial credit for reaching correct conclusions through incorrect reasoning.</span></p></li></ul><ul><li><p><span>The authors used scaffolds like Claude Code and OpenCode (an open-source equivalent) to evaluate 36 model configurations. Opus 4.7, running in Claude Code, performed best, at ~46%.</span></p></li><li><p><span>The study&#8217;s methodology differs from OpenAI&#8217;s </span><a href="https://openai.com/index/gdpval/"><span>GDPVal</span></a><span>, which compares AI outputs for a task against an output from a human expert. This makes it hard to directly compare the two. Agents clearly struggle much more on JobBench, but it&#8217;s unclear to what extent this is because AI performs less well on the tasks that JobBench focuses on&#8212;i.e. those that employees most want to delegate. Or whether it is because the tasks in JobBench are &#8216;messier&#8217; and the threshold to succeed is higher.</span></p></li><li><p><span>The researchers also find that scaffolds matter a lot. Claude Sonnet 4.6 scored 37% inside Claude Code but only ~31% inside OpenClaw&#8212;similar to the difference between model families. This suggests that equipping AI agents with the right context and tools to perform a given task, and employees&#8217; ability to do this, is a key determinant of how useful AI is in the workplace.</span></p><p></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;559bfa3d-3df7-4738-8427-1214ddebfe37&quot;,&quot;caption&quot;:&quot;The notion of AIs manipulating people is a plot twist in countless sci-fi thrillers. But is &#8220;manipulative AI&#8221; really possible? If so, what might it look like?&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;AI Manipulation &quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:12431790,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;bio&quot;:&quot;AI Policy Writer @ Google DeepMind &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fad94b7d-013b-4773-98cb-b9014a1857b8_1024x1024.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-02-05T12:53:27.695Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!cz8J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206e194a-9018-41db-a123-7583aed33e85_1024x572.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-manipulation&quot;,&quot;section_name&quot;:&quot;Interviews &quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:186969167,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:21,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[Conjecture Machines ]]></title><description><![CDATA[AI agents and the new validation bottleneck in science]]></description><link>https://www.aipolicyperspectives.com/p/conjecture-machines</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/conjecture-machines</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Wed, 15 Jul 2026 12:33:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uI0t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>By </span><a href="https://substack.com/@donwallace"><span>Don Wallace</span></a><span>, </span><a href="https://substack.com/@conorg1"><span>Conor Griffin</span></a><span>, </span><a href="https://seanoneill.co.uk/"><span>Sean O&#8217;Neill</span></a><span>, </span><a href="https://www.linkedin.com/in/thang-luong"><span>Thang Luong</span></a><span>, </span><a href="https://www.linkedin.com/in/owen-larter-3340b224/"><span>Owen Larter</span></a></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uI0t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uI0t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!uI0t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!uI0t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!uI0t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uI0t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uI0t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 424w, https://substackcdn.com/image/fetch/$s_!uI0t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 848w, https://substackcdn.com/image/fetch/$s_!uI0t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 1272w, https://substackcdn.com/image/fetch/$s_!uI0t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1f84fabc-d39b-4bec-b24b-6644b4652e02_2048x1143.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>Over the past year, agents have transformed what it means to code. Scientific research may be next &#8212; from proposing novel hypotheses, to designing experiments, to discovering algorithms that improve on the best that humans have designed. This raises urgent questions for policymakers and science funders, the biggest being how to validate the coming wave of AI-generated ideas.</span></em></p><p><em><span>New approaches are also needed to ensure that scientists can access agents, that datasets are agent-ready, and that peer review is kept afloat. We sat down with 10 researchers and engineers from Google DeepMind to try to figure out what comes next.</span></em></p><p><span>Microbiologist Jos&#233; Penad&#233;s and his team at Imperial College London took most of a decade to work out how a family of superbugs spreads antibiotic resistance. The result was unpublished, known only inside his lab. Then, in 2024, he described the problem to </span><a href="https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/"><span>Co-Scientist</span></a><span>, an AI agent from Google DeepMind.</span></p><p><span>Within two days, Co-Scientist returned five potential explanations, ranked in priority. Number 1 was the same hypothesis Penad&#233;s&#8217;s team had spent so long working to prove: that some superbugs acquire tails from viruses and use them as &#8220;keys&#8221; to jump between host species. Stunned, his first thought was that his computer had been compromised, so he emailed Google to check. Confirmed: no peeks taken. If he could have gone back in time, that insight would have saved his team years.</span></p><p><span>An</span><a href="https://cloud.google.com/discover/what-is-agentic-ai"><span> AI agent</span></a><span> such as Co-Scientist is a large language model (LLM)- based system given a goal and the tools to pursue it. Unlike a query-answering chatbot, an agent can plan how to achieve the goal you give it, breaking it down into steps, running multiple subagents and processes in parallel, and detecting and correcting errors as it goes. It can also engage with the wider world, for example by calling online databases and tools, or writing code to operate robots.</span></p><p><span>The arrival of AI agents is timely. Researchers face a rapidly growing &#8220;burden of knowledge&#8221;, with more to learn before they can meaningfully contribute. The questions that count in drug discovery, climate modelling, materials design, and biology are becoming </span><a href="https://www.aipolicyperspectives.com/p/a-new-golden-age-of-discovery"><span>too complex and interdisciplinary</span></a><span> for human teams to tackle at pace.</span></p><p><span>Science depends in part on a social infrastructure that has gradually evolved: labs, teams, institutions, peer review, grant funding, and networks that support the accumulation of shared knowledge. The era of AI agents will challenge that infrastructure, and at times demand rapid change &#8212; some of which is perhaps overdue.</span></p><p><span>Software engineering has been going through a version of this. In barely a year, coding agents reshaped the working practices of engineers. AI agents are </span><a href="https://arxiv.org/pdf/2605.06651"><span>now starting to</span></a><span> reshape science. Science is messier, but it also suits agents in certain ways: a large literature available as text, huge databases, and some workflows already built around code. Science also has an existing generation of specialist AI models that agents can put to work, from </span><a href="https://deepmind.google/blog/alphaproteo-generates-novel-proteins-for-biology-and-health-research/"><span>protein</span></a><span> and </span><a href="https://deepmind.google/blog/millions-of-new-materials-discovered-with-deep-learning/"><span>materials</span></a><span> design tools to state-of-the-art </span><a href="https://deepmind.google/science/weathernext/"><span>weather forecasting</span></a><span>.</span></p><h2><strong><span>Why are AI agents suddenly so useful?</span></strong></h2><p><span>Anybody who has worked with early iterations of AI agents may be sceptical of their utility in science, where rigour and reliability are essential. But three forces are making today&#8217;s agents </span><a href="https://ai.google/gemini-for-science/"><span>more useful</span></a><span>: stronger frontier models, &#8220;scaffolding&#8221; that enables agents to elicit greater capabilities from those frontier models, and customisability that allows scientists to tailor agents to their needs.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YEDT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YEDT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 424w, https://substackcdn.com/image/fetch/$s_!YEDT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 848w, https://substackcdn.com/image/fetch/$s_!YEDT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 1272w, https://substackcdn.com/image/fetch/$s_!YEDT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YEDT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png" width="1456" height="820" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YEDT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 424w, https://substackcdn.com/image/fetch/$s_!YEDT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 848w, https://substackcdn.com/image/fetch/$s_!YEDT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 1272w, https://substackcdn.com/image/fetch/$s_!YEDT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7bded8eb-fd4b-493b-8d5d-e481e81749fe_2048x1153.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Stronger frontier models.</span></strong><span> Leading models now outpace human experts on demanding scientific knowledge benchmarks, such as </span><a href="https://agi.safe.ai/"><span>Humanity&#8217;s Last Exam</span></a><span> and </span><a href="https://epoch.ai/frontiermath"><span>FrontierMath</span></a><span>. Perhaps more significant is their new depth of thinking, driven by </span><a href="https://blog.google/products-and-platforms/products/gemini/gemini-2-5-deep-think/"><span>inference-time reasoning</span></a><span>, a technique in which models work through problems in extended steps, exploring and revising before committing to an answer.</span></p><p><span>In mathematics and computer science, where breakthroughs can rest on reasoning alone, </span><a href="https://deepmind.google/blog/accelerating-mathematical-and-scientific-discovery-with-gemini-deep-think/"><span>models are producing impressive results</span></a><span>. But sharper reasoning helps beyond mathematics, enabling agents to draw on the &#8220;long tail&#8221; of findings that are often buried in papers, make scientific connections across fields, judge which tools to call and when, and catch their own errors as they go.</span></p><p><strong><span>Scaffolding. </span></strong><span>This is a type of &#8220;harness&#8221; that enables AI agents to better elicit the latent capabilities that reside in base models but which do not emerge by default. This code &#8220;wraps&#8221; a frontier model, giving it structure, memory, and the ability to access and use tools including code execution, scientific software, and literature search. It also lets agents call on specialised AI models and interact with other agents. <br><br>Early agent scaffolds were bespoke: they gave a particular agent detailed instructions for how to plan, carry out tasks, and use tools. But often they did not transfer well to other agent systems, or even to later generations of the frontier models they depended on. New standards, such as </span><a href="https://a2a-protocol.org/latest/"><span>protocols for how agents can communicate with each other</span></a><span>, should reduce the need for so much custom scaffolding.</span></p><p><strong><span>Customisation.</span></strong><span> Agents can be tailored to a discipline, a lab&#8217;s workflow, or an individual researcher. A key mechanism for that is the current proliferation of </span><a href="https://github.com/google-deepmind/science-skills"><span>agent &#8220;skills&#8221;</span></a><span> that users can create to their own detailed specs.</span></p><p><span>Base models hold a great deal of explicit scientific knowledge, the kind written down in papers and textbooks. What they lack is tacit knowledge: the hard-to-articulate craft, built up over years of practice and failure, that lets a scientist coax a cell line into growing or get a simulation to run well. Until now, scientists using AI tools had to convey their methodological know-how, processes, and preferences the hard way: through detailed prompts, bespoke scaffolding, or model retraining.</span></p><p><span>Agent skills make some of this know-how portable. They are instruction sets, often just simple text files, that tell an agent how to perform a task and what to produce. Skills are relatively easy to produce, share and accumulate. As a scientist uses their agent more, the agent can also draw on these interactions, making the experience more personalised. Scientific labs and institutions can also make their proprietary data &#8212; such as old lab notebooks &#8212; securely available to their agents.</span></p><p><span>Natasha Latysheva, a computational biologist, says she has distilled aspects of her research process into skills. &#8220;While AI agents aren&#8217;t fully reliable yet, I think the trend is clear. Scientific research will shift from hands-on execution to high-level orchestration,&#8221; she says. &#8220;We&#8217;ll start our workday by reviewing the experiments and analyses our agents ran overnight, tweaking their direction and guiding their attention.&#8221;</span></p><h2><strong><span>How AI agents are changing (and not changing) science</span></strong></h2><p><span>The most immediate change for scientists is also the most mundane: the availability of smart, tireless digital assistants. Researchers are handing off more of their daily grind &#8212; sifting the literature, querying databases, orchestrating analyses, curating data, writing grant proposals &#8212; to agents that do it in minutes rather than hours or days, then report back.</span></p><p><span>Agents also open up new possibilities. &#8220;Researchers can suddenly do things they couldn&#8217;t do at all before,&#8221; says Matej Balog, a Senior Staff Research Scientist. For example, plenty of scientists lack the skills to build the software tools and pipelines they need. This is partly because developing software for science </span><a href="https://dl.acm.org/doi/10.1145/3685265?__cf_chl_f_tk=wRXe3J6n__nAHTLA6Wfu0AZ.4GGTiMKPi9dMzGFlBk8-1782991715-1.0.1.1-ASx0wdsAflGUhSW8PvAyb6o6.jaVc0q7g7M4GyVDwbk"><span>is hard</span></a><span>. But also because most scientists are not deeply trained in programming, and science has </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9138129/"><span>not traditionally provided competitive career paths</span></a><span> for dedicated research software engineers in academic labs.</span></p><p><span>Agents are particularly strong at writing code because it is a domain with an enormous amount of training data, where correctness can typically be checked automatically. A scientist can now describe what they need in natural language &#8212; &#8220;write me a script to clean and merge these three datasets&#8221;, &#8220;build me an interactive browser-based tool to explore this output&#8221; &#8212; and get serviceable code in minutes.</span></p><p><span>But ultimately, the biggest change for scientists is a structural one. Agents are making it easier to come up with ideas and proposed solutions for problems they are working on, but are not yet providing the same uplift when it comes to validating them.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7SJA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7SJA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 424w, https://substackcdn.com/image/fetch/$s_!7SJA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 848w, https://substackcdn.com/image/fetch/$s_!7SJA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 1272w, https://substackcdn.com/image/fetch/$s_!7SJA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7SJA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png" width="1456" height="820" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7SJA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 424w, https://substackcdn.com/image/fetch/$s_!7SJA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 848w, https://substackcdn.com/image/fetch/$s_!7SJA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 1272w, https://substackcdn.com/image/fetch/$s_!7SJA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea422f74-5de9-4563-be75-201fbfb8952b_2048x1153.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Ideation</span></strong></p><p><span>Most scientists have no shortage of ideas. The challenge is knowing which ones to pursue. This is where agents are starting to help. An agent can digest a field&#8217;s accessible literature, making connections across disciplines that no single researcher would have time to trace. Existing </span><a href="https://gemini.google/overview/deep-research/"><span>Deep Research tools</span></a><span> already use subagents to do this. But agents optimised for ideation go even further. Co-Scientist, for example, directs a raft of subagents to generate diverse hypotheses, critique them, rank them, and iterate &#8212; mimicking aspects of how a human research group operates, but at remarkable speed.</span></p><p><a href="https://deepmind.google/blog/uncovering-repurposed-medicines-to-fight-liver-fibrosis/"><span>Gary Peltz at Stanford University used the tool</span></a><span> in his hunt for existing drugs that could be repurposed to treat liver fibrosis, the scarring process behind 1.4 million cirrhosis deaths a year. Based on his own literature review and decades of expertise, he picked two candidate drugs. Co-Scientist picked three. Neither of Peltz&#8217;s picks showed any benefit in assays with live human liver cells. Two of Co-Scientist&#8217;s picks not only blocked fibrosis but also promoted liver cell regeneration.</span></p><p><span>This is a positive example, but any LLM-based system is fallible &#8212; even the strongest reasoners can still make things up. &#8220;A single hallucinated claim on page 10 of an output can invalidate the whole thing,&#8221; says Vivek Natarajan, a Co-Scientist lead. Missing such errors wastes time and resources, and catching them can require a scientist with deep expertise. This fallibility makes scientists understandably cautious about AI agents, and reducing it is a top priority for those building these systems. To function effectively in a research setting, agents cannot act as black boxes that simply output answers; they must expose their reasoning and serve as transparent collaborators.</span></p><p><span>One key challenge is imbuing agents with what Natarajan calls &#8220;epistemic humility&#8221;: models that know when they don&#8217;t know, and say so. He points to AlphaFold, the protein-structure predictor, which is prized partly because it reports how confident it is in each aspect of each prediction, so researchers know when to trust it and when to reach for other methods. Calibrating that confidence across more open-ended scientific reasoning remains an unsolved problem.</span></p><p><span>Looking further ahead, the harder problem may be a tension between two things scientists want at once: hypotheses that are grounded in the literature and free of error, but also genuinely novel. Agents tuned for caution can drift toward the safe and the known; tuned for originality, they are more likely to go astray. This challenge will itself require new ideas, such as </span><a href="https://challengeworks.org/challenge-prizes/metascience-novelty-indicators/"><span>better ways</span></a><span> to evaluate novelty. Or better ways to connect agents to more specialized models &#8212; trained on first-principles scientific data &#8212; to help generate ideas that are both original and robust.</span></p><p><strong><span>Finding the optimal candidate solution</span></strong></p><p><span>Agents can search for the best solution to a problem, not just a workable one. A materials scientist hunting for a new catalyst faces a near-infinite number of molecular structures, each costly to make and test. A computer scientist looking for a more efficient algorithm faces a similar explosion of possibilities.</span></p><p><span>Vast solution spaces like these turn up all across science and industry. </span><a href="https://deepmind.google/blog/alphaevolve-impact/"><span>AlphaEvolve</span></a><span> is an agent that can find the best candidates within them. Given a problem expressed as code and a way to score potential solutions, it orchestrates an ensemble of agents to generate many algorithmic candidates, keeps those that score highest, and breeds the next generation from the survivors. It runs unattended, generating and improving candidates at a scale no human team could match.</span></p><p><span>AlphaEvolve works in code, but its reach extends well beyond software. As Balog points out: &#8220;Algorithms can accurately describe so many of the world&#8217;s scientific and natural processes.&#8221; </span><a href="https://deepmind.google/blog/alphaevolve-impact/"><span>AlphaEvolve has assisted</span></a><span> in the design of Google&#8217;s next-generation TPU chips, helped the mathematician Terence Tao solve open Erd&#337;s problems, and improved the analysis of genomics data.</span></p><p><span>But in scientific domains such as materials, AlphaEvolve&#8217;s top scorers are leads, not final results. A promising catalyst still has to be made and measured at the bench.</span></p><p><strong><span>Validation</span></strong></p><p><span>Validation is the slow, costly business of testing whether an idea survives contact with reality. Karl Popper said science advances through conjectures and refutations. Agentic AI is changing the economics of that pairing. AI agents are conjecture machines, making ideas and candidate solutions abundant and relatively cheap. Refutations remain physical and institutional &#8212; and so, costly and slow.</span></p><p><span>Mathematics and computer science are often viewed as great exceptions because validation can run </span><em><span>in silico</span></em><span>. An AI agent can generate a proof, represent it in a formal language like Lean, and have the computer verify, unambiguously, that it is correct.</span></p><p><span>Even for the parts of maths that can&#8217;t yet be described in formal language, validation is advancing. </span><a href="https://arxiv.org/abs/2602.10177"><span>Aletheia</span></a><span> pairs a proof generator with </span><a href="https://arxiv.org/abs/2511.01846"><span>a natural language verifier</span></a><span> that checks its work and sends flaws back for revision. In February this year, mathematicians ran the inaugural </span><a href="https://1stproof.org/"><span>First Proof challenge</span></a><span>: 10 research problems, kept unpublished so they couldn&#8217;t be found in any training data. In the week allotted, Aletheia solved six &#8212; the best result.</span></p><p><span>Mathematicians will need to absorb all these new outputs. Some already complain that AI-generated proofs are too long and hard to parse (although </span><a href="https://openai.com/index/model-disproves-discrete-geometry-conjecture/"><span>certain</span></a><span> AI-generated proofs are short and elegant). &#8220;We are moving toward a future of serious &#8216;proof indigestion&#8217; where AI generates breakthroughs faster than humans can review them,&#8221; says Thang Luong, who led the Aletheia effort. To break this bottleneck, </span><a href="https://arxiv.org/abs/2605.22763"><span>automated verification</span></a><span> must become standard, but this </span><a href="https://arxiv.org/abs/2606.03303"><span>verification</span></a><span> will need to combine the absolute correctness of formal languages like Lean with the more expressive reasoning of natural language.&#8221;</span></p><p><span>Alex Davies, who also leads work on AI for mathematics, is mindful that by automating large chunks of mathematicians&#8217; workflows, his discipline is dealing with questions that others may face in the coming years: &#8220;I can imagine a world in which machines do the discovery, and what&#8217;s left for mathematicians is to understand it, and to decide what questions are worth pursuing next.&#8221; Luong echoes this, noting that maths may also provide some of </span><a href="https://arxiv.org/abs/2605.20531"><span>the general technology</span></a><span> needed to address the validation bottleneck in other fields: &#8220;One can think of mathematics as an accelerated testbed for the rest of science&#8221;.</span></p><p><span>At the moment, however, the validation gap in most disciplines is widening, not closing. An agent can </span><a href="https://deepmind.google/blog/fast-tracking-genetic-leads-to-reverse-cellular-aging/"><span>propose</span></a><span> a novel genetic lead to reverse cellular ageing, but cannot say definitively whether it actually works. This explains why companies like Google DeepMind, Ginkgo Bioworks and Lila Sciences are investing in automated labs. But they only suit some fields, are expensive to build and are still early in development. And even automation cannot rush nature&#8217;s clock. Cell lines need time to grow, chemical reactions take time to complete. For much of science, then, the lab sets the pace.</span></p><h2><strong><span>Implications for policymakers and research funders</span></strong></h2><p><span>To some extent, agents simply add intensity to questions that policymakers are already focused on in their </span><a href="https://www.gov.uk/government/publications/ai-for-science-strategy/ai-for-science-strategy"><span>AI for Science strategies</span></a><span>, such as how to train the next generation of scientists, how to experiment with new forms of scientific institutions, and how to ensure that AI is not misused by threat actors, while still putting the technology to use addressing the various natural risks that society faces, like the next pandemic.</span></p><p><span>For every encouraging scenario, there is a challenging one. For example, AI agents could make it more feasible for small, agile teams to pursue creative, ambitious ideas, reversing the trend towards &#8220;big team science&#8221;, or enable scientists </span><a href="https://www.nature.com/articles/s41586-025-09048-1"><span>to work across domains, </span></a><span>bringing new perspectives to existing problems. Assuming efficiency gains make agents sufficiently cost-effective, these trends could particularly benefit smaller, less well-resourced countries and institutions.</span></p><p><span>But agents will also give rise to anxiety among junior scientists that their institutions are choosing to spend budgets on tokens instead of staff. And left unmanaged, there is a risk that agentic tools could de-skill new generations of scientists before they develop the judgement needed to use them effectively. For the same reason that mathematics students still prove theorems unaided, science-graduate training may need structured periods of agent-free work and access to agents that act as genuine </span><a href="https://cloud.google.com/solutions/learnlm"><span>cognitive partners</span></a><span> rather than oracles.</span></p><p><span>Beyond these questions, AI agents present at least four urgent new priorities: 1. Scientists need access to the tools. 2. The tools need access to agent-ready data. 3. We need more experimental infrastructure to validate AI ideas. 4. And we need to update the peer review process.</span></p><ol><li><p><strong><span>Ensure widespread access to agents</span></strong></p></li></ol><p><span>Agents will be extremely useful and fallible in non-obvious ways. Both factors provide a strong rationale for policymakers to ensure that all scientists can access the best agents &#8212; to speed up discovery and to provide the independent evaluations of AI agents that the scientific community needs to judge how best to use them.</span></p><p><span>This is an urgent strategic priority for policymakers and science funders, akin to the historical challenge of providing access to supercomputers. Geopolitical debates today often focus on one aspect of sovereign capability &#8212; whether a state can train its own frontier model. Much less attention is paid to what may prove to be a more important issue: a country&#8217;s ability to deploy agents across its scientific ecosystems for transformative impact.</span></p><p><span>At the micro level, funders must first decide how labs and researchers select and pay for agents. Selection is the easier near-term problem: let researchers find the most useful tools for themselves, without excessive approvals or complex procurement.</span></p><p><span>Paying is harder. The temptation is to use existing structures, with scientists seeking funding through grant applications or drawing on lab budgets. But the compute required to run agents can be large and the frontier of what agents can do is constantly expanding. While the cost per unit of AI capability is falling fast, total lab expenditures on agents are still likely to rise as agents take on longer and more complex tasks. Policymakers must quickly assess whether budget uplifts or entirely new funding programmes are needed. Delivering access at the scale and price needed will require novel public-private partnerships; the </span><a href="https://genesismissionconsortium.org/"><span>US Genesis Mission</span></a><span> is one </span><a href="https://deepmind.google/blog/google-deepmind-supports-us-department-of-energy-on-genesis/"><span>promising model</span></a><span>.</span></p><p><span>Bigger questions await. Agents may make the existing process of allocating national budgets across disciplines more legible</span><em><span>, </span></em><span>forcing science funders and research programmes to quantify the investment in data and compute needed to make progress on specific problems. This in turn may lead to more targeted debates about the relative value of solving different problems. If agents propose the top hypotheses to explore across an entire field, with very expensive experimental validation plans, how should this fit into national funding strategies?</span></p><ol start="2"><li><p><strong><span>Make national data assets agent-ready</span></strong></p></li></ol><p><span>While scientists need access to agents, agents need access to data. Data that is open or low-risk should be exposed to agents through well-documented APIs, with sufficient quality control and metadata. The engineering support and maintenance to do so is not trivial, so funders should ensure such data stewardship is properly resourced and support interoperable data standards. But ultimately, the ability of agents to help extract and annotate data &#8212; from PDFs to download portals &#8212; provides an opportunity for governments to extract a lot more value from the data they have already funded.</span></p><p><span>More sensitive datasets in genomics, virology, or other areas carrying dual-use risk often come with restrictions on who may use them and how. The challenge now is to develop similar privacy-preserving </span><a href="https://www.ukbiobank.ac.uk/use-our-data/research-analysis-platform/"><span>solutions</span></a><span>, when appropriate, for agents, with auditability and privacy built in. Examples like </span><a href="https://www.opensafely.org/"><span>OpenSAFELY</span></a><span>, which lets human researchers securely access valuable health data, can provide inspiration. The prize is large. A dataset analysis that currently takes years of doctoral work could, with the right secure agent infrastructure, run autonomously in days.</span></p><p><span>Perhaps most importantly, agents provide a strong rationale for funding the creation of entirely new open datasets. This leads to a further question for funders: could agents help identify the most important datasets to fund? Some of the authors of this article recently made a human-expert-driven attempt to answer </span><a href="https://deepmind.google/public-policy/science-needs-ai-data-stocktakes/"><span>that question</span></a><span> for fusion energy. How soon will agents be capable of running similar &#8220;AI data stocktake&#8221; exercises?</span></p><ol start="3"><li><p><strong><span>Tackle the validation bottleneck</span></strong></p></li></ol><p><span>Many scientists already struggle to get enough time in facilities to run their experiments. As AI agents make hypotheses and candidate solutions increasingly abundant, this bottleneck will only tighten. Policymakers and funders should address this in at least two ways: investing in existing experimental validation infrastructure and accelerating progress on automated labs.</span></p><p><span>Public research bodies hold extensive experimental facilities across almost every scientific field. AI agents provide a reinvigorated case for investing in them and opening them up, by renting bench space or experimental run-time to researchers testing computational hypotheses and predictions against reality. Direct partnerships with AI labs are another avenue. Google DeepMind has created a wet lab inside the UK&#8217;s Francis Crick Institute, a leader in biomedical research, and is also providing independent scientists funding &#8212; alongside Co-Scientist access &#8212; to carry out the wet lab experiments needed to validate agent-enabled hypotheses. The US government&#8217;s </span><a href="https://genesismissionconsortium.org/"><span>Genesis Mission</span></a><span> will connect the world-class experimental facilities of the Department of Energy&#8217;s (DOE) National Laboratories with academia and the AI industry.</span></p><p><span>Automated labs are another promising route to tackling the validation bottleneck, but they currently rely on expensive robotics and compute. Ensuring broad access will require public investment, and there are good early efforts here. The US </span><a href="https://www.nsf.gov/news/nsf-invest-new-national-network-ai-programmable-cloud"><span>National Science Foundation has put</span></a><span> $100 million towards a national network of distributed facilities, while the UK </span><a href="https://www.gov.uk/government/publications/sovereign-ai-open-call-autonomous-labs"><span>launched</span></a><span> a call for ideas and already hosts the &#163;81 million </span><a href="https://www.liverpool.ac.uk/materials-innovation-factory/"><span>Materials Innovation Factory</span></a><span>. The build-out of automated labs will almost certainly go beyond what any single lab or institution can afford, so governments should also explore building centralised capacity and adopting the kind of &#8220;</span><a href="https://www.energy.gov/science/office-science-user-facilities"><span>user facility&#8221; access model</span></a><span> seen at the US DOE&#8217;s national labs.</span></p><ol start="4"><li><p><strong><span>Empower peer reviewers with agents</span></strong></p></li></ol><p><span>The </span><a href="https://worksinprogress.co/issue/real-peer-review/"><span>peer review process has long been under strain</span></a><span>, with slow timelines and reviews of varied quality. Now, scientists are using AI to write ever more grant applications and papers. This is making it harder for funders to know which research to fund, and for peer reviewers to validate findings and identify the most important work.</span></p><p><span>As </span><a href="https://www.nature.com/articles/d41586-026-01297-y"><span>noted</span></a><span> by Professors James Wilsdon and Geraint Rees, the challenge is not just an increase in supply; writing quality is also no longer a reliable discriminator. Agents deepen the problem. The more an agent is left to plan and optimise an application, drawing on the funder&#8217;s criteria and its recent winners, the less the bid reflects a scientist&#8217;s original thinking.</span></p><p><span>Funders, journals, and conferences are trialling various responses. For grant applications, some suggest leaning less on the written word and more on the investigator&#8217;s track record and team. The UK&#8217;s Medical Research Council recently </span><a href="https://researchfunding.on.worc.ac.uk/?p=2189"><span>reinstated interviews for shortlisted applicants</span></a><span>. These ideas have promise, but also risk privileging seasoned experts or increasing costs.</span></p><p><span>The likely answer will be a layered approach. Those developing and using AI should document its use more clearly. That could include advancing </span><a href="https://deepmind.google/models/synthid/"><span>watermarking techniques</span></a><span> and ideas such as </span><a href="https://arxiv.org/pdf/2602.10177"><span> &#8220;Human-AI Interaction Cards&#8221;</span></a><span> &#8212; short records detailing the prompts and outputs that produced the key scientific insights.</span></p><p><span>Reviewers should also be able to use agents. This will not be easy, as AI use is often banned, even if </span><a href="https://www.nature.com/articles/d41586-025-04066-5"><span>widely practised in secret</span></a><span>. To move forward, organisations can build on the </span><a href="https://www.dfg.de/en/basics-topics/digital-topics/ai/review#389620"><span>more nuanced guidelines</span></a><span> that some have started to develop and </span><a href="https://arxiv.org/abs/2606.28277"><span>test agents</span></a><span> in more objective areas where they are likely to be strong &#8212; such as detecting errors.</span></p><p><span>They could also take steps to ensure that scientists view agents as &#8220;human-centric&#8221; tools that enhance, rather than bypass, their judgement &#8212; for example, by being transparent about the systems and ensuring that agents expose their reasoning, verify their claims with citations, and (to the extent possible) document their uncertainty. By understanding how an AI agent reaches its conclusion, scientists will have opportunities to learn, intervene, and collaborate.</span></p><p><strong><span>These are formative years</span></strong></p><p><span>The scientists using AI agents are not going back. The technology will continue to improve. The institutions that govern science &#8212; and the physical and social infrastructure it runs on &#8212; were built for a research enterprise that is now evolving faster than they are. We need to upgrade them for the agent era.</span></p><p><span>This will take a significant collective effort and serious investment. AI agents must not be viewed as a rationale to spend less on science; this would be a tragic false economy. The countries that treat this as a moment of genuine structural change will unlock the greatest public value and shape what comes next. The choices made in this window may be difficult to reverse.</span></p><div><hr></div><p><strong><span>Acknowledgements</span></strong></p><p><span>We would like to thank the following experts at Google DeepMind who shared insights with us through interviews and feedback on the draft. All mistakes belong to the authors.</span></p><p><span>Vivek Natarajan, Sebastian Nowozin, Tim Green, Natasha Latysheva, Simon Batzner, Matej Balog, Samuel Albanie, Alex Davies, Nenad Tomasev, Juan Mateos-Garcia, Catherine Pollard, Uchechi Okereke, Agata Laydon, Anna Koivuniemi, Andy Song and Tom Rachman.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Science of Learning ]]></title><description><![CDATA[A discussion with Carl Hendrick]]></description><link>https://www.aipolicyperspectives.com/p/the-science-of-learning</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/the-science-of-learning</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Wed, 08 Jul 2026 11:00:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qSnG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><span>Cognitive science tells us much about how our brains convert the blizzard of passing information into long-term knowledge and skills. Yet education systems often overlook it, while clinging to other practices that instruct less well. Why?</span></em></p><p><em><span>This question exasperates </span><a href="https://carlhendrick.substack.com/"><span>Carl Hendrick</span></a><span>, a writer and professor of education who began his career teaching English at a state school in London. Today, he also advises </span><a href="https://alpha.school/"><span>Alpha School</span></a><span>, the fashionable but controversial US-based private educator, hoping to teach kids in 2-hours a day, followed by social activities, with AI and no teachers.</span></em></p><p><em><span>To find out how (and if) this will work, we asked Carl what cognitive science tells us about learning, how Alpha operates, and why the decline of reading matters.</span></em></p><p><em><span>- Conor Griffin, AI Policy Perspectives</span></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qSnG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qSnG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qSnG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qSnG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qSnG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qSnG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg" width="727.9948120117188" height="406.49710313566436" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:727.9948120117188,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qSnG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 424w, https://substackcdn.com/image/fetch/$s_!qSnG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 848w, https://substackcdn.com/image/fetch/$s_!qSnG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!qSnG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2991a060-b712-49af-81e5-0535c72016cf_2048x1143.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p><strong><span>Conor: Many students hate memorization. You think it&#8217;s central to learning? Why?</span></strong></p><p><strong><span>Carl: </span></strong><span>Cognitive science tells us that a change in your long-term memory is the fundamental invariant of learning. It&#8217;s non-negotiable. Without it, you&#8217;ve learned nothing.</span></p><p><span>In the 1950s and &#8217;60s, we had a cognitive science revolution that helped us understand our memory as a system. In working memory, we have a hard constraint, whereby we can hold maybe five to seven elements at a given time. Our superpower is long-term memory which, as far as we know, offers an unlimited reserve to store things. If you store information in your long-term memory, then your working memory becomes outsized because you can effortlessly draw on the symbols and ideas you have stored. We&#8217;re doing it right now as we speak.</span></p><p><span>So cognitive science tells us that memorization is important to learning. But it also tells us that rote memorization of isolated facts is a debased form of learning, because</span><em><span> </span></em><span>memory is </span><em><span>schematic&#8212;</span></em><span>the concepts we learn need other concepts to stick to. So, if we want students to learn about World War Two, we don&#8217;t want them to just memorize dates. We want them to be able to build their own schema, or mental models, about the war. The question for educators becomes: how do you decompose World War Two into atoms of knowledge that are teachable to a student, but which they can then rebuild?</span></p><p><strong><span>Conor: Why are certain forms of teaching effective at this, like &#8216;retrieval practice&#8217;, where students are asked to recall facts or concepts from their memory?</span></strong></p><p><span>In a word: friction.</span></p><p><span>The </span><a href="https://bjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf%5C"><span>desirable difficulties</span></a><span> framework, developed by Robert and Elizabeth Bjork, the psychologists and memory experts, says that learning is often counterintuitive. Things that feel like they are working to a student, such as re-reading and highlighting a text, often aren&#8217;t. And things that feel like they aren&#8217;t working, or are a struggle, like getting kids to dredge up something from their memory, often do lead to learning.</span></p><p><span>This comes back to how memory works. Human memory is not like a tape recorder. </span><a href="https://carlhendrick.substack.com/p/the-important-peculiarities-of-memory"><span>It is reconstructive</span></a><span>. When we remember something, we seem to almost rewrite it. This is also why eyewitness testimony is so unreliable. But from a learning perspective, successfully retrieving a memory can strengthen it for the future.</span></p><p><span>With retrieval practice, the encoding of the knowledge seems to happen after the original event, when the students are trying to remember the thing. This is why the Bjorks see retrieval as a learning event, rather than a test of whether you have already learned something.</span></p><p><strong><span>Conor: What&#8217;s the evidence that these cognitive-science interventions actually help students learn better?</span></strong></p><p><span>There is a huge body of evidence from postgraduate students in labs, but little evidence for it in actual classrooms. And I say that as somebody who passionately believes in it.</span></p><p><span>A 2021 </span><a href="http://potentialplusuk.org/wp-content/uploads/2022/02/Cognitive_Science_in_the_classroom_-_Evidence_and_practice_.pdf"><span>review</span></a><span> found that we know almost nothing about how things like retrieval practice work with real kids and teachers. There is 70 years of evidence, but it is usually small studies with post-graduate students, focussed on maths and languages.</span></p><p><span>This is why we need to create new models of learning with AI. I want to see specific questions answered: &#8220;</span><em><span>Given this third-grade reading problem in this state, what is the best intervention?</span></em><span>&#8220; It will be hard to isolate the effects in the way that we could with a vaccine. But with AI, we will be able to understand the effects of using a certain curriculum and sequencing the material in a certain way. We will be able to understand how to break down complex concepts into components to present to students progressively, without overloading their working memory.</span></p><p><strong><span>Conor: Are there any interventions that do have solid evidence from real classrooms?</span></strong></p><p><span>Yes, early reading and the Direct Instruction method.</span></p><p><span>For early reading, the debate is largely over. We know that instruction needs to be done with </span><a href="https://www.gov.uk/government/publications/phonics-teaching-materials-core-criteria-and-self-assessment"><span>systematic, synthetic phonics</span></a><span>. The brain is not built to read. So phonics teaches young children the individual sounds associated with specific letters, or letter combinations, before teaching them how to blend these sounds together to form words.</span></p><p><span>The most evidenced programme is Direct Instruction, a highly-structured method where teachers break complex concepts down into small, manageable steps. In the 1960s and &#8216;70s, it was subject to the biggest ever educational study in the US, </span><a href="https://www.nifdi.org/what-is-di/project-follow-through.html"><span>Project Follow Through</span></a><span>. This found that it was far more impactful for students than other approaches, for early reading, maths and other areas.</span></p><p><span>People tend to think of Direct Instruction as teachers just &#8220;telling kids stuff&#8221;. But instruction is actually the least important part of it. The most important part is the curriculum design. For example, there is a big focus on learning through contrasts. In the same way as we teach young children the colour red by showing them a red ball followed by a red triangle, you isolate the essential thing that you want them to know. This is a fundamental part of how we learn.</span></p><p><span>I think that there&#8217;s just no doubt that AI is going to be able to do this kind of curriculum design and sequencing far better than humans can. I think of the famous </span><a href="https://www.alphagomovie.com/"><span>move 37</span></a><span> in AlphaGo. I think we&#8217;re going to see a similar thing for the design of curriculum and learning materials, where AI will just far surpass what humans can do.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h3><span>How Alpha School works</span></h3><p><strong><span>Conor: That brings us to the new group of private schools in the US that you work with, that has big ambitions to use AI. You&#8217;re one of several prominent learning scientists advising </span><a href="https://alpha.school/"><span>Alpha</span></a><span>. Why did you get involved?</span></strong></p><p><strong><span>Carl:</span></strong><span> Traditional schooling has been hugely successful; nothing has lifted more kids out of illiteracy and poverty. But there are problems.</span></p><p><span>The system is largely propped up by the 1-in-20 &#8220;superhero&#8221; teachers. The experience of pupils is massively heterogeneous. Since Covid-19, something else has changed. Many kids are not in school at all, or they are there so little that it&#8217;s almost a waste of time. Until now, educational technology, or EdTech, has also largely been a story of expensive failure because it creates digital versions of things that didn&#8217;t work in the first place.</span></p><p><span>But at Alpha School, we are designing apps where you get high-resolution data points from students at a phenomenal clip. You get correct/incorrect answers, but also things like latency and student hesitation. You can slice that data and make strong predictions about when and how to do interventions like retrieval practice. Or &#8216;spacing&#8217;&#8212;where students distribute their learning and practice over many lessons or days, rather than cramming it into a single session, to more productively engage their memory. This will allow us to see those interactive effects between the curricula and how you sequence the teaching materials.</span></p><p><span>Ultimately, I&#8217;m a materialist when it comes to learning. I think learning is governed by the same biological laws as digestion. If you can get the brain to pay attention to a certain sequence of information, and if knowledge is retrieved under a certain set of conditions, then I think learning will be almost guaranteed for 99% of kids.</span></p><p><strong><span>Conor: At Alpha, how are kids supervised during this online learning?</span></strong></p><p><strong><span>Carl: </span></strong><span>The big change will be moving to AI coaching. This is where people will get uncomfortable because the answer is to mimic what a really good teacher does: warm/strict supervision. If you think about MOOCs&#8212;the massive open online courses that were much-hyped around 15 years ago&#8212;they were a failure because students weren&#8217;t accountable. Only 5 to 10 percent of people who were already highly motivated finished them.</span></p><p><span>Schooling fulfills the function of accountability. In the future, this will mean cameras on, and an AI monitoring student behavior, their latency, what websites they are looking at, their ability to focus and then producing a report for a human tutor. If you can get kids to concentrate on the apps, on the sequencing, phenomenal learning is going to happen.</span></p><p><span>But people won&#8217;t like this. You&#8217;ll see articles about &#8220;Orwellian spyware.&#8221; As a parent, I&#8217;d prefer this model to simply unleashing kids on the internet, crossing my fingers, and hoping for the best. AI will supervise better than humans. There are very skilled teachers who can supervise a class well, with eyes on the back of their head. But even then, there are kids who drift through the lesson, &#8216;cosplaying&#8217; attention, and learning nothing.</span></p><p><span>Today, when you look at kids in a computer room, they might spend 20 percent of their time on a task. The rest of the time, when the teacher&#8217;s back is turned, they&#8217;ll be looking at websites. The model for AI should be a high level of accountability for students.</span></p><p><strong><span>Conor: Students spend a couple of hours in the morning on this intensive online learning. What do they do in the afternoon?</span></strong></p><p><span>They work on &#8216;life skills&#8217;. Broadly, longer-term projects, like setting up a YouTube channel, playing music in a band, or sport. Let&#8217;s not forget, Alpha is an expensive private school. So you will have kids there, who have had tremendous support from their parents, and so they will already be highly-motivated and able to do many of these things.</span></p><p><span>And so there is a question about whether that model is going to work elsewhere. But Alpha is trying to understand how learning happens. And they want to remove the standard model where kids sit at their desk for six to eight hours straight.</span></p><p><strong><span>Conor: Alpha School doesn&#8217;t have conventional teachers. Even with a good AI system, won&#8217;t this mean a lack of social influence? With a good teacher, you want to impress them. How will you manage that?</span></strong></p><p><strong><span>Carl: </span></strong><span>We are in uncharted territory here, because we don&#8217;t have any data on this. In the Alpha model, the nearest thing they have is &#8216;Guides&#8217; who receive a student&#8217;s data and have conversations with them, such as: &#8220;I noticed you went off-topic there for a while?&#8221;</span></p><p><span>That human question is a really important one and something that I&#8217;ve wrestled with. If more and more of the curriculum design, the instructional sequencing, and the assessment is delivered by AI, then it&#8217;s obviously going to come through tablets and screens. And then the question becomes: What do we lose?</span></p><p><span>There is still a social environment at Alpha. It&#8217;s a physical building. Kids come in. They have an assembly where they interact. And then, at nine o&#8217;clock, it&#8217;s two hours of online learning, although they&#8217;re still beside each other as they do this. There&#8217;s a system, where they have to generate a certain amount of experience points, and get a certain amount right, and so that brings some accountability.</span></p><p><span>But the question you&#8217;ve asked is a massive one, because it comes down to: What kind of schooling do we want for our kids? Are we going to eliminate something valuable&#8212;a teacher inspiring a kid? And I guess my answer is: we can&#8217;t scale that. We only hear the success stories, but what about the other 25 kids in that class? The Alpha model </span><em><span>does</span></em><span> offer scaling. You say to a kid, &#8220;Focus for two hours, hit these targets, and then you&#8217;re off to be a kid&#8221;. I think this will improve their academic outcomes. It might also improve their mental health. But these are difficult questions.</span></p><p><strong><span>Conor: When one imagines the ideal schooling, it&#8217;s always the inspiring teacher. But the social influence of a teacher can also have an estranging effect, on those it doesn&#8217;t work for?</span></strong></p><p><strong><span>Carl: </span></strong><span>Exactly. One of my daughters was diagnosed with autism. School will be a challenge for her, because there is a noise in most classrooms, plus a lot of ambiguity and blurred edges. That can cause her distress. I was looking at an app that she&#8217;s using, where she flies through problems because there is more clarity and boundaries.</span></p><p><span>That was one reason I started working with Alpha. I realised that this will help kids with special educational needs that are not met in mainstream education. They could have a system where they&#8217;re making rapid progress, and not feeling like the bad kid in class all the time.</span></p><p><span>But again, the messaging is going to be very difficult. Critics will say you just want to &#8220;put autistic kids on screens&#8221;. But it&#8217;s for an hour or two a day. And then we want them to be kids.</span></p><p><strong><span>Conor: What role do &#8216;AI tutors&#8217; play in this vision?</span></strong></p><p><strong><span>Carl: </span></strong><span>There are three broad areas in education: curriculum (what we teach), instruction (how we teach it), and assessment (how we know they know). When the public hears &#8216;AI tutor&#8217;, I think they usually imagine a robot doing the instruction part.</span></p><p><span>But for me, AI tutoring is more about the back-end: curriculum design and assessment. If you accept that learning is an algorithmic enterprise governed by biological laws, then the most foundational elements are what you teach and how you assess it.</span></p><p><span>Designing a curriculum and teaching materials is more about engineering than about whether you can inspire kids. It&#8217;s about sequencing the materials. There are broader questions about the experience and social interaction that we want students to have. But, for more basic questions, like, &#8220;</span><em><span>Is there a better way to teach quadratic equations</span></em><span>?&#8221; I think we&#8217;ll see a shift from the current &#224; la carte approach where teachers often decide what materials to use. And as uncomfortable as people are with it, I imagine that we will get AI-generated content that is just better than human-generated content.</span></p><p><strong><span>Conor: You also see AI changing how we </span></strong><em><strong><span>assess </span></strong></em><strong><span>students?</span></strong></p><p><span>Yes, your earlier point is right: if a kid writes something, they often want you (the teacher) to read it. But the Welsh educationalist </span><a href="https://en.wikipedia.org/wiki/Dylan_Wiliam"><span>Dylan Wiliam</span></a><span> also referred to marking books as the most expensive public relations exercise in history. It&#8217;s tedious for teachers and much of it has zero impact on student outcomes.</span></p><p><span>This is also because human marking is unreliable. I was Head of Department of English and I saw situations where a student would do an essay, and get a D. And then it would be sent back to be re-marked and another person would give it an A. You can&#8217;t run a system like that.</span></p><p><span>The </span><a href="https://substack.nomoremarking.com/p/can-ai-assist-the-comparative-judgement"><span>work on comparative judgment</span></a><span> by Daisy Christodoulou and </span><a href="https://www.nomoremarking.com/"><span>No More Marking</span></a><span>, which calls for student writing to be assessed in comparison to other students&#8217; work, rather than as an absolute, shows how AI could help improve the reliability of marking and reduce teacher workload.</span></p><p><span>But we also need to </span><a href="https://www.dylanwiliam.org/Dylan_Wiliams_website/Papers_files/NCME%2004%20paper.pdf"><span>distinguish</span></a><span> between the assessment </span><em><span>of</span></em><span> learning, and assessment </span><em><span>for</span></em><span> learning. Assessment shouldn&#8217;t just be an endpoint. It should be a regular temperature check. In a classroom, this kind of &#8216;checking for understanding&#8217; is the most powerful thing a teacher can do, whether it&#8217;s a mini whiteboard session or a cold-calling activity. It&#8217;s imperfect, but it can identify misconceptions that students have.</span></p><p><span>The dream is that instead of teaching a pre-defined sequence of materials, an AI system takes regular temperature checks and operates like faders on a mixing desk. Based on the student&#8217;s response, the system identifies misconceptions in real-time and dynamically selects the best explanation for that specific kid.</span></p><p><span>This also means that we can meet the student at the point where the misconception occurs. Most current assessments are useless because the feedback loop is too long. If a kid writes an essay and gets it back a week later, they&#8217;ve already forgotten their thought process.</span></p><p><strong><span>Conor: This would make education &#8216;adaptable&#8217;</span></strong><em><strong><span> </span></strong></em><strong><span>to a given student. But you&#8217;ve also cautioned against using AI to match teaching materials to students&#8217; different &#8216;learning styles&#8217;. What&#8217;s the nuance here?</span></strong></p><p><strong><span>Carl: </span></strong><span>The idea of learning styles emerged from a tradition in the 1960s and &#8216;70s, which thought about learning in an individualized way. As a teacher, I had to spend years trying to design lessons for different learner types, such as &#8216;visual learners&#8217; or &#8216;kinesthetic&#8217; learners, who purportedly learn better through hands-on experiences.</span></p><p><span>This idea has been empirically tested. Not only is there no evidence for learning styles, but it is the nearest thing we have in education to homeopathy or healing crystals. When you think about it, it&#8217;s preposterous to suggest that if you were teaching geography to an &#8216;auditory&#8217; learner, you should make an audiobook to describe the continent of Europe. The content should determine the form.</span></p><p><span>Where there </span><em><span>is</span></em><span> evidence, is around the idea of dual-coding. Our brains process information through separate but interconnected channels: one for verbal information (spoken or written words) and one for visual information (images, animations etc.). There is value in designing explanations that think carefully about how to combine text and speech with visuals, so that you don&#8217;t overload students&#8217; limited working memory.</span></p><p><span>For example, I see some great video explainers on YouTube, where, they show the actual differences between the size of the planets and it&#8217;s nothing at all like the maps of the solar system that you normally see. That&#8217;s an example of where the visual aspect really helps, but it has nothing to do with me being a visual learner.</span></p><h1><span>Where education goes wrong</span></h1><p><strong><span>Conor: You are often supportive of explicit forms of instruction and skeptical of constructivist</span></strong><em><strong><span> </span></strong></em><strong><span>approaches where students explore and solve problems by themselves. Why?</span></strong></p><p><strong><span>Carl: </span></strong><span>Constructivism is a philosophy of meaning, not a pedagogy. It&#8217;s the belief that each individual constructs their own knowledge and understanding. It has its roots in Rousseau, but really got moving around the turn of the 20th century, with people like (the education reformer) John Dewey in the US, and (the Swiss psychologist) Jean Piaget in Europe.</span></p><p><span>As a philosophy, constructivism is true. We </span><em><span>do c</span></em><span>onstruct meaning, but from an education perspective it has led to a belief in minimally guided instruction, the idea that you learn things best when you discover them for yourself, rather than having them taught to you.</span></p><p><span>What this does is to conflate the outcomes with the means. We </span><em><span>do </span></em><span>want students to be able to discover things for themselves. But discovery learning as a starting point for instruction is a disaster. Clear, sequenced instruction is effective for novices. And every kid is a novice in most domains.</span></p><p><span>If you&#8217;re learning to drive a car, you want explicit instruction up top&#8212;do this, then do that. Once a person is capable of driving, if you&#8217;re still doing that as a teacher, you&#8217;re getting in their way. </span><em><span>That&#8217;s </span></em><span>when the discovery element comes in.</span></p><p><span>There&#8217;s an ethical dimension to this too. Students who come from affluent backgrounds, with strong existing schemas of knowledge, who have been exposed to rich vocabulary around the dinner table, will flourish in an environment where there are few constraints or expectations in terms of their behaviour. Discovery learning privileges the already privileged.</span></p><p><strong><span>Conor: This idea that novices can&#8217;t be expected to figure out things for themselves is also why you are skeptical of students learning by asking questions to chatbots?</span></strong></p><p><span>Yes, the power of AI is in the monitoring, curriculum design, instructional design, and assessment. Novices won&#8217;t learn by discovering knowledge themselves by asking chatbots questions; they need explicit instruction.</span></p><p><strong><span>Conor: Your argument that educators conflate the </span></strong><em><strong><span>outcomes </span></strong></em><strong><span>with the </span></strong><em><strong><span>means, </span></strong></em><strong><span>also explains why you&#8217;re skeptical of calls to teach students &#8216;21st century skills&#8217;, like critical thinking, creativity and adaptability, including as a response to AI?</span></strong></p><p><strong><span>Carl: </span></strong><span>Yes. Critical thinking skills and creativity are an outcome of systematic knowledge building, not a starting point that can be directly taught, as a generic skill.</span></p><p><strong><span>Conor: The explicit instruction of knowledge that you advocate for might be effective for subjects like maths that can be more easily decomposed into sub-components. But what about subjects like English or history?</span></strong></p><p><strong><span>Carl:</span></strong><span> That is the Holy Grail of curriculum design. We have what AI pioneer Marvin Minsky and the computer scientist Walter Reitman would call </span><a href="https://www.bibsonomy.org/bibtex/1583e389b273a1089b1b3c4ab1aebaa16/quesada"><span>&#8220;well-defined&#8221; and &#8220;ill-defined&#8221; domains.</span></a></p><p><span>Well-defined domains, like maths and science, are more clearly sequenced and hierarchical. They have concepts that are free-standing and can be taught in an isolated way. You can go back two steps if you miss something. They often have answers that are definitive or unambiguously correct. A lot of education apps that work well are in these areas.</span></p><p><span>In the humanities, knowledge doesn&#8217;t operate in that way. This will have implications for how we design instruction and assessment. But I go back to the materialist point. I believe these things are amenable to the laws of the universe and AI will discover what the optimum sequences are.</span></p><h1><span>Why Reading Still Matters</span></h1><p><strong><span>Conor: You&#8217;ve spoken a lot about the decline of reading. Why does it worry you?</span></strong></p><p><strong><span>Carl: </span></strong><span>I find it sad that a lot of kids are not going to read Dostoevsky. But I also recognise that for me, growing up, reading was partly a feature of being bored, and we can&#8217;t pretend the conditions are the same today. It comes down to an ethical question: what kind of life do we want for our kids? This is not a question we can answer in a randomized controlled trial.</span></p><p><span>There are certain questions about reading that we </span><em><span>can</span></em><span> answer empirically in trials</span><em><span>. </span></em><span>One of the studies that stuck me dead in my tracks many years ago showed that the percentage of words we need to understand, in order to read a text, is 95%, which shocks everybody. The way that vocabulary is taught in schools is very poor, and inequitable, and there are almost certainly better ways to do it.</span></p><p><span>But we can&#8217;t just say to kids: &#8220;You need to read more!&#8221; That doesn&#8217;t wash. We need to answer the philosophical question of what we value reading for. The benefits are not always immediately obvious. They might pay off 20 years down the line. The literary critic Harold Bloom once said that Shakespeare shaped the modern consciousness, including our contemporary understanding of love, death, ambition. Something like the King James Bible affected English-language culture too. These were dominant texts, and culture passed through them.</span></p><p><span>What are we going to lose if we lose these books? I want my daughters to read because they&#8217;ll see that their struggles are not their own. So they can see other characters struggle too.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/the-science-of-learning?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/the-science-of-learning?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;7b86e528-c1e0-4a9a-879f-d2f8cce5990a&quot;,&quot;caption&quot;:&quot;Today&#8217;s post comes from Daniel Gillick, a research scientist at Google DeepMind, who works on making Gemini more useful for teaching and learning. Daniel explores five pedagogical principles that the team is using in their work, the degree to which today&#8217;s AI systems can embody them, and what that means for how we should think about AI tutors. As with a&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;lg&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;AI tutors should not approximate human tutors&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:101194130,&quot;name&quot;:&quot;AI Policy Perspectives&quot;,&quot;bio&quot;:&quot;Reflections on AI policy, governance, and more. https://www.aipolicyperspectives.com/ &quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0Byl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9146f89b-5561-4adb-bf89-cc21b508c264_667x374.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2025-11-10T08:42:48.206Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!qOT7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d76e77b-f498-4375-929e-247ea9067ef2_922x518.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-tutors-should-not-approximate&quot;,&quot;section_name&quot;:&quot;Essays&quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:178289069,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:28,&quot;comment_count&quot;:5,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p></p>]]></content:encoded></item><item><title><![CDATA[5 Rules of AI Writing]]></title><description><![CDATA[Who cares if humans write anymore?]]></description><link>https://www.aipolicyperspectives.com/p/5-rules-of-ai-writing</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/5-rules-of-ai-writing</guid><dc:creator><![CDATA[Tom Rachman]]></dc:creator><pubDate>Tue, 30 Jun 2026 14:07:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2d3N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2d3N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2d3N!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2d3N!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2d3N!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2d3N!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2d3N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/adfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2d3N!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 424w, https://substackcdn.com/image/fetch/$s_!2d3N!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 848w, https://substackcdn.com/image/fetch/$s_!2d3N!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!2d3N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadfafebc-b05d-4a9a-811a-87a818d4fa76_2048x1117.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="callout-block" data-callout="true"><p style="text-align: center;"><em><strong><span>5 Rules of AI Writing</span></strong></em></p><p><em><strong><span>1: Assume your AI usage will be discovered. </span></strong></em></p><p><em><strong><span>2: Test whether you&#8217;re going too far by asking: Would it be wrong to accept this help from a human without crediting them?</span></strong></em></p><p><em><strong><span>3:  Know that, when AI writes for you, it deletes your future wisdom.</span></strong></em></p><p><em><strong><span>4: Never expect anyone to read your outsourced AI answers.</span></strong></em></p><p><em><strong><span>5: If you&#8212;or the reader&#8212;would care who the named author is, don&#8217;t generate.</span></strong></em></p></div><div><hr></div><p><strong><span>It&#8217;s an intellectual vanity to claim you can always identify AI writing</span></strong><span>&#8212;that chatbots spit out nothing but sanitized blah. Yes, the humans who talk most with machines are better at detecting their patter. But proclaiming &#8220;You can just tell!&#8221; is a coping mechanism with an expiration date.</span></p><p><span>Actually, you </span><em><span>can</span></em><span> tell. Because AI writes better than nearly every human alive, explaining with clarity, and allowing the curious to bypass the impenetrable prose with which the self-important sandbag outsiders. In many respects, AI writing is a boon to human understanding.</span></p><p><span>Weird effects abound too.</span></p><p><span>Bureaucrats must wade through </span><a href="https://www.nytimes.com/2026/05/25/us/politics/artificial-intelliegence-courts.html?unlocked_article_code=1.lVA.YGhJ.zOBuVs163qAU&amp;smid=nytcore-android-share"><span>floods</span></a><span> of oddly articulate petitions, legal writs, and complaints, as the public&#8217;s literary servants overwhelm the public servants. Police can spam the spammers, able to shovel AI slop into criminal-hackers&#8217; chatrooms, much to the </span><a href="https://arxiv.org/pdf/2603.29545"><span>ire</span></a><span> of the crooks. Meantime, the best man&#8217;s tipsy </span><a href="https://www.toastpal.com/"><span>speech</span></a><span> at the wedding and his sober words at a funeral&#8212;they&#8217;re so emotive you know a machine wrote them.</span></p><p><span>That&#8217;s not to mention the high-profile wins for literary AI, such as the prize-winning </span><a href="https://granta.com/the-serpent-in-the-grove/"><span>short story</span></a><span>, cited for its author&#8217;s &#8220;melodic voice,&#8221; which earned a different review from the detection software: &#8220;100% AI-generated.&#8221; Or the (somewhat) non-fiction book </span><em><a href="https://www.wired.com/story/future-of-truth-ai-interview/"><span>The Future of Truth</span></a></em><span>, about how AI undermines factuality, and included AI-generated quotes. Not to mention social media, glutted with even more vapid </span><a href="https://www.bbc.co.uk/news/articles/c9wx2dz2v44o"><span>junk</span></a><span> than when humans alone monopolized that genre.</span></p><p><span>In the 19th century, theologians spoke of &#8220;</span><a href="https://en.wikipedia.org/wiki/God_of_the_gaps"><span>God of the gaps</span></a><span>&#8221;: that the holy spirit reigned where science faltered. Only, science kept expanding, forcing the supernatural into retreat. The philosopher of technology Benjamin Bratton says we&#8217;re now anxiously defending a &#8220;humanism of the gaps,&#8221; defining our species by whatever machines cannot do </span><em><span>quite</span></em><span> yet.</span></p><p><span>But AI capabilities arrow </span><a href="https://epoch.ai/benchmarks?view=graph&amp;tab=benchmarks"><span>upward</span></a><span> in so many domains. Why should writing escape it? From </span><a href="https://arxiv.org/abs/2510.18774"><span>journalism</span></a><span> to </span><a href="https://www.nytimes.com/2026/03/19/books/shy-girl-book-ai.html"><span>fiction</span></a><span>, more than a few prose pros are already leaning on AI writing, with its tantalizing offer to solve the blank page. Such usage remains taboo, so few admit to it. But a Nobel Prize winner just did, the Polish novelist Olga Tokarczuk, who </span><a href="https://mycompanypolska.pl/artykul/olga-tokarczuk-zapowiada-ostatnia-powiesc-w-karierze-pisanie-dlugich-opowiesci-jest-dzis-ekonomicznie-nieoplacalne/20717"><span>seeks</span></a><span> creative pointers from AI.</span></p><p><span>&#8220;I often throw an idea to the machine for analysis, asking, &#8216;Honey, how could we develop this beautifully?&#8217; &#8221; she </span><a href="https://mycompanypolska.pl/artykul/olga-tokarczuk-zapowiada-ostatnia-powiesc-w-karierze-pisanie-dlugich-opowiesci-jest-dzis-ekonomicznie-nieoplacalne/20717"><span>said</span></a><span> at a recent onstage event.</span></p><p><span>&#8220;At the same time, I feel a poignant, very human sorrow for an era that is disappearing forever. My heart aches for the passing of traditional literature, written over months in solitude, a work of life crafted in the mind of a fully conscious, single individual.&#8221;</span></p><p><span>Personally, I have devoted my adult life to writing, publishing five volumes of literary fiction, a nonfiction book that I ghostwrote, and scores of newspaper articles&#8212;years of piling up sentences such that an editor might buy them on a computer file. But it dawned on me: the computer, not my files, is what enchanted the world. So I </span><a href="https://www.nytimes.com/2024/10/07/opinion/novelist-back-to-school-behavioral-science-identity.html"><span>pivoted</span></a><span> to the topic of the times, AI policy, the many-handed effort to push an elephant.</span></p><p><span>Which is to say, I worry about the fate of human writing. And yet something odd keeps happening.</span></p><p><span>The instant I discover that what I&#8217;m reading&#8212;perhaps rather interesting&#8212;is AI-generated, I suddenly don&#8217;t care. The words dim before me. It&#8217;s an off-switch.</span></p><p><span>Why?</span></p><h3><span>THE PAINTER MADE OF PIXELS</span></h3><p><span>A conceptual artist named SHL0MS </span><a href="https://x.com/SHL0MS/status/2054280631807316329"><span>posted</span></a><span> a picture on X, and asked readers to say why the image was worse than a real Monet.</span></p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/SHL0MS/status/2054280631807316329&quot;,&quot;full_text&quot;:&quot;i just generated an image in the style of a Monet painting using AI\n\nplease describe, in as much detail as possible, what makes this inferior to a real Monet painting &quot;,&quot;username&quot;:&quot;SHL0MS&quot;,&quot;name&quot;:&quot;&#74794;&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1998301675015057408/FdlU7_G9_normal.jpg&quot;,&quot;date&quot;:&quot;2026-05-12T19:20:43.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HIJE_1EW8AACAdz.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/VDJovKOqlz&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:1386,&quot;retweet_count&quot;:1003,&quot;like_count&quot;:9398,&quot;impression_count&quot;:7243040,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p><span>Nobody </span><a href="https://x.com/Jediwolf/status/2054776716770320631"><span>spared</span></a><span> the machine&#8217;s feelings.</span></p><p><span>&#8220;I&#8217;m disappointed I have to even point it out,&#8221; one person commented. &#8220;The background lilypad-algae amalgam is egregiously vague, like most AI art.&#8221;</span></p><p><span>Another added, &#8220;Doesn&#8217;t look anywhere near like a Monet. Looks exactly like somebody trying to replicate the style and achieving like 20% of it.&#8221;</span></p><p><span>In fact, the painting </span><em><span>was</span></em><span> a Monet, and the post was a prank that, seemingly, exposed the pretensions of people who flock to galleries, claiming to admire the art but really admiring status, entering each room and hurrying to the wall text: </span><em><span>Who is this? Should I care?</span></em></p><p><span>That seems shallow. But the same drive could save human writing.</span></p><p><span>After all, you </span><em><span>should</span></em><span> care more about a real Monet than a dupe. Not because one is objectively better, but because its meaning exists in a matrix of social beliefs about beauty, about value, about the shared tales of human civilization: that an irascible Frenchman once swiped hog bristle across canvas, his perceptions and drives filtered into a decorative depiction of a pond, conveyed from dealer to collector, rising the ranks of culture, finding a lonely museum wall, venerated there for years, as the outside world transformed, the artist died, and pixels reproduced his perceptions and drives, connecting your sensing brain to the sensing brain of a particular man who will see nothing ever again.</span></p><p><span>All that is only in our heads. But where else</span><em><span> </span></em><span>is meaning but in heads? And writing is humanity&#8217;s most sophisticated technology for visiting another&#8217;s head.</span></p><p><span>The problem is that writing is hard. Words elude you; paper doesn&#8217;t care. George Orwell equated </span><a href="https://www.orwellfoundation.com/the-orwell-foundation/orwell/essays-and-other-works/why-i-write/"><span>writing a book</span></a><span> to suffering from a long illness. Developers dream that AI will someday </span><a href="https://www.bbc.co.uk/future/article/20260309-ai-is-finding-treatments-for-incurable-diseases"><span>cure</span></a><span> all illnesses. Why not this one?</span></p><h3><span>WHY (A)I WRITE</span></h3><p><span>In his essay &#8220;</span><a href="https://www.orwellfoundation.com/the-orwell-foundation/orwell/essays-and-other-works/why-i-write/"><span>Why I Write</span></a><span>,&#8221; Orwell listed four main motives: 1) sheer egoism; 2) aesthetic pleasure; 3) the urge to say what is; and 4) the urge to</span><em><span> change</span></em><span> what is. Others cite a fifth reason: to know what you think.</span></p><p><span>Those who don&#8217;t write presume the process involves hatching an idea, then putting it into words. It&#8217;s commonly the reverse: put down words to hatch the ideas within them.</span></p><p><span>This points to a flashing peril from humans impersonating themselves via AI writers: </span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6097646"><span>cognitive surrender</span></a><span>, that we hand over the wearisome word toil, and atrophy in genteel luxury, never enduring the mental frictions that kindle into wisdom. We wouldn&#8217;t detect the decline, I suspect, half-believing ourselves the authors still.</span></p><p><span>As a leading AI thinker told me in private, &#8220;I write so I can keep writing in the future, so I can keep thinking in the future.&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EERO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EERO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 424w, https://substackcdn.com/image/fetch/$s_!EERO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 848w, https://substackcdn.com/image/fetch/$s_!EERO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!EERO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EERO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EERO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 424w, https://substackcdn.com/image/fetch/$s_!EERO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 848w, https://substackcdn.com/image/fetch/$s_!EERO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!EERO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabe1e822-8da0-4e68-bd08-f9a7116889ca_2048x1117.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">An AI-generated character in search of an author.</figcaption></figure></div><p><span>But there&#8217;s another reason that people write, a motive far more powerful than humankind&#8217;s dwindling determination to remain the brightest species. It&#8217;s ego. As Joan Didion </span><a href="https://lithub.com/joan-didion-why-i-write/"><span>observed</span></a><span>, &#8220;Writing is the act of saying I, of imposing oneself upon other people, of saying </span><em><span>listen to me, see it my way, change your mind</span></em><span>.&#8221;</span></p><p><span>Every person craves attention, and the nerds buy it with words. Readers possess a natural resource&#8212;attention&#8212;that is everywhere and nowhere to be found. A transaction ensues: reader agrees to spend attention on words in exchange for information (nonfiction) and/or experiences (fiction). The writers&#8217; art is to smuggle themselves into both.</span></p><p><span>So, when reading human, you&#8217;re absorbing the meaning with one eye while focusing the other on its maker, ever trying to reconcile the two, much as the visual cortex creates 3-D from the parallax of two eyes synthesizing one object at dueling angles. If you see that the author was AI, one eye closes. The page becomes markings on a flat surface.</span></p><p><span>The reader&#8217;s presumption of a human behind the words makes secret AI authorship worse than betrayal. It&#8217;s trespassing and fraud at once, entering another&#8217;s brain on false premises and altering the contents for gain, whether it&#8217;s AI-generated reviews that trick you into feeding at someone&#8217;s restaurant, or an AI-generated marriage proposal that tricks you into feeding someone&#8217;s children. It&#8217;s the sale of a shared reality, from an entity that has no reality to share.</span></p><p><span>That said, much text is purely for dispersing information, and the author is irrelevant. This is true for most legal documents, press releases, and boilerplate emails, not to mention explanations that you prompt a chatbot to conjure. AI composition can be helpful and harmless, provided it&#8217;s honest.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h3><span>THE UNWRITTEN RULES</span></h3><p><span>But nobody knows the ethics yet. Hidden usage seems wrong because deceit is wrong. But must you specify authorship in </span><em><span>every</span></em><span> case, even if it&#8217;s a casual note that you autocomplete? And if you&#8217;re candid about it, does that make AI writing fine?</span></p><p><span>News publications </span><a href="https://futurism.com/artificial-intelligence/new-york-times-freelancers-ai-rules"><span>ban</span></a><span> it, educators </span><a href="https://academic.oup.com/jope/advance-article-abstract/doi/10.1093/jopedu/qhaf097/8422834?redirectedFrom=fulltext"><span>resist</span></a><span> it, and academic journals </span><a href="https://pubsonline.informs.org/doi/10.1287/orsc.2026.ed.v37.n3"><span>flounder</span></a><span>. Yet social norms could be shifting subtly toward acceptance. Increasingly, you hear the view that authorship is immaterial; it&#8217;s about making something worth reading. &#8220;Honest to god,&#8221; </span><a href="https://x.com/tobiaschneider/status/2062950193390117321"><span>wrote</span></a><span> Tobias Schneider, a research fellow at the Global Public Policy Institute in Berlin, &#8220;what do I care how the text for an academic paper comes about. It could be a monkey throwing darts at a board, if the resulting combination of letters conveys original insight, it should be published.&#8221;</span></p><p><span>As with most AI anxiety, it&#8217;s a matter of </span><a href="https://www.aipolicyperspectives.com/p/time-machines"><span>speed</span></a><span>: this is moving faster than the culture, leaving us in a perma-panic as one train after another hurtles through the station, heading to destinations unknown with more of our belongings aboard.</span></p><p><span>Already, large-language-model vocab is springing more frequently from our mouths, according to one </span><a href="https://arxiv.org/pdf/2409.01754"><span>study</span></a><span>. Authors, in dread of false accusation, are altering their prose to avoid seeming LLMish. I know because I&#8217;m doing so, tense about my appreciation for bullet points and curbing the dashes that I learned to overuse in old-fashioned newsrooms&#8212;for melodramatic punch.</span></p><p><span>Errors have become fashionable too. Once, clean copy was a hallmark of professionalism. But &#8220;the house style in most newsrooms is extremely LLMable,&#8221; the </span><em><span>Atlantic</span></em><span> magazine contributor Jasmine Sun </span><a href="https://x.com/jasminewsun/status/2061871693891776808"><span>noted</span></a><span>. &#8220;What stands out (besides reporting) is a distinct and authentic first-person voice, even if that means the occasional typo.&#8221;</span></p><p><span>Not for nothing is </span><a href="https://theconversation.com/what-is-wabi-sabi-will-this-japanese-philosophy-make-me-happy-275786"><span>wabi-sabi</span></a><span>&#8212;the Japanese tradition of charming imperfections&#8212;trending. To err is human. To not err is AI.</span></p><p><span>For those who venerate human writing, the nightmare is that AI authorship becomes so commonplace that few care</span><em><span> </span></em><span>about origin anymore. Nobody would need to lie about it anymore. Looking at recent history, technological tools for knowledge work have tended to draw contempt at first, only to gain wide adoption. Many researchers shunned internet sources back in the 1990s, while journalists of the early 2000s scorned Wikipedia. But the technology improved, and norms updated.</span></p><p><span>If AI writing resolves its quirks, will the same acquiescence follow?</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LWx-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LWx-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LWx-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg" width="789" height="612" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:789,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LWx-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">If only he had a chatbot. (<em>The Passion of Creation</em> by Leonid Pasternak) </figcaption></figure></div><h3><span>3 FACTORS THAT COULD SAVE HUMAN WRITING</span></h3><p><span>Here&#8217;s what may decide whether most people still care about people writing: 1) institutions clarifying the rules and norms of AI usage; 2) reliable detection technology becoming widespread; 3) people retaining the &#8220;off-switch&#8221; aversion to AI authorship:</span></p><ol><li><p><strong><span>Establishing Rules</span></strong></p></li></ol><p><span>When it comes to AI-generated text, the closest norms regard plagiarism, which also involves the presentation of others&#8217; work as one&#8217;s own. But the concept of plagiarism arose to protect the intellectual property of a fellow human. With AI writing, who is the victim?</span></p><p><span>If writing is dishonestly presented, one victim is plain: whoever is duped into reading it. This suggests that openly AI-generated text is fine. But victims may exist in those cases too.</span></p><p><span>Among the harmed could be people who fail to write, and therefore never develop their minds. Also, there are the human authors who&#8217;d need to vie against hordes of fluent bots hogging public attention. Most of all, society could suffer from the decline of collective cognition, if writing is no longer a push-and-pull involving thinkers living and dead, but just effortless insta-words materializing onscreen.</span></p><p><span>Yet AI writing is immensely useful, and stimulating, and will stir human creativity. We must set rules and norms for concrete cases, not just shout, &#8220;Don&#8217;t!&#8221;</span></p><p><span>Is it wrong for an AI to write an action plan for a business client? What about a write-up of the family vacation to share with loved ones? How about assigning AI agents to write your social-media feed?</span></p><p><span>The debate over AI writing often misses an even more likely outcome: </span><em><a href="https://arxiv.org/abs/2510.03154"><span>assisted </span></a></em><a href="https://arxiv.org/abs/2510.03154"><span>writing</span></a><span>. What if you come up with all the ideas, but an AI turns them into clearer prose than you could, which you then tweak? Or reverse that: the AI brainstorms, but you write the prose. Which is worse, and why?</span></p><p><span>Recent studies suggest that AI writing may </span><a href="https://arxiv.org/pdf/2603.18161"><span>neutralize</span></a><span> the human user&#8217;s arguments, and that essays drafted by people contain a range of ideas, whereas AI </span><a href="https://www.sciencedirect.com/science/article/pii/S294988212500091X"><span>essays</span></a><span> tend to </span><a href="https://arxiv.org/pdf/2606.01736"><span>converge</span></a><span> around similar points.</span></p><p><span>On the other hand, chatbots&#8212;by disseminating ideas so much more gamely than unmoving written texts&#8212;may stir our creativity. In theory, future personalized AI models could allow users to vary the &#8220;temperature&#8221; (randomness) of LLM responses, increasing their novelty.</span></p><p><span>But without social norms, dystopian futures come readily to mind, with humans generating AI text for other humans, who employ AI to read it and craft replies, downgrading our species&#8217; intellectual conversation into a feedback loop of artifice, in which we are mindless mouthpieces of literate machines.</span></p><p><span>That feels far-fetched today. Yet some people are already cramming past emails and blogs and tweets into AI, training agentic systems to capture their &#8220;voice,&#8221; such that they needn&#8217;t cultivate one.</span></p><p><span>If leading institutions care to keep human writing, they must establish clear and coordinated policies in haste, so people know the professional and social costs of their choices. This depends on </span><a href="http://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo"><span>detecting</span></a><span> what is human and what is machine.</span></p><ol start="2"><li><p><strong><span>Detection Technology</span></strong></p></li></ol><p><span>Disgrace is an enforcer. Soon after the explosion of chatbots in late 2022, nobody could reliably identify whether text had been written or generated. But AI-writing detection software&#8212;notably, the Pangram app&#8212;has become</span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5407424"><span> strong</span></a><span>.</span></p><p><span>You can tell because it keeps precipitating scandals. But institutions have yet to apply detection tools as widely as they should. That would mean standardized detection everywhere we care about human input and honesty: job applications, dating apps, any publication&#8217;s vetting process.</span></p><p><span>For detection apps, the challenge is twofold: </span><em><span>Do not mistake AI writing for human</span></em><span> (a false negative); and </span><em><span>Do not mistake human writing for AI</span></em><span> (a false positive). The first error makes the app worthless, but the second error is especially dangerous. A wrongful accusation&#8212;whether against an author, a student, or a public figure&#8212;could destroy a career.</span></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;cf112491-7e6a-409e-96b7-508f22aeb2c3&quot;,&quot;caption&quot;:&quot;I just heard a maven of the AI scene using a new verb for the addictive practice of testing all you read online in an AI-writing detector. &#8220;I&#8217;m Pangramming everything lately,&#8221; he said, referring to the popular app.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Q&amp;A: Pangram CEO&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:12431790,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;bio&quot;:&quot;AI Policy Writer @ Google DeepMind &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fad94b7d-013b-4773-98cb-b9014a1857b8_1024x1024.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-06-23T12:07:00.528Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!DFP6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo&quot;,&quot;section_name&quot;:&quot;Interviews &quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:202611257,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:25,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p><span>Last year, the University of Chicago economists Brian Jabarian and Alex Imas </span><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5407424"><span>published</span></a><span> an evaluation of several detection apps, having had them evaluate around 2,000 human-written passages, including fiction and non-fiction from before 2020 (to avoid texts that were stealth AI), along with the same number of AI-generated passages.</span></p><p><span>&#8220;The results are clear-cut,&#8221; Jabarian and Imas wrote. &#8220;On medium-length to long passages, Pangram achieves essentially zero FPRs [False Positive Rates] and FNRs [False Negative Rates] within our sample.&#8221; The error rate did increase slightly with short passages, under 50 words. (For reference, this paragraph is exactly 50 words long.)</span></p><p><span>The accused still deserve a fair hearing. Institutions should establish systems of appeal, while independent evaluators must keep testing the accuracy of detection tools. This poses an uncomfortable question: what rate of false accusations is acceptable? The reflexive answer is, &#8220;None!&#8221; But we accept a degree of error in matters that are even more impactful, such as the justice system. Society considers it better to identify many culprits than to preclude any miscarriage.</span></p><p><span>Benjamin Franklin </span><a href="https://oll.libertyfund.org/titles/franklin-the-works-of-benjamin-franklin-vol-xi-letters-and-misc-writings-1784-1788"><span>preferred</span></a><span> that 100 guilty go free than 1 innocent suffer, while the legal profession debates &#8220;</span><a href="https://en.wikipedia.org/wiki/Blackstone%27s_ratio"><span>Blackstone&#8217;s ratio</span></a><span>,&#8221; which proposes 10:1. Developers of detection apps likewise make such moral calculations, with Pangram training its models to average no more wrongful accusations than 1 in 10,000.</span></p><p><span>But AI developers are pursuing ever-better writing tools, which could raise the opposite problem: increasingly mistaking AI writing for human. In which case, we&#8217;ll need further antibodies.</span></p><ol start="3"><li><p><strong><span>The Off-Switch</span></strong></p></li></ol><p><span>&#8220;More than 300 years after the reading revolution ushered in a new era of human knowledge, books are dying,&#8221; </span><a href="https://jmarriott.substack.com/p/the-dawn-of-the-post-literate-society-aa1"><span>warns</span></a><span> the leading chronicler of literary decline, talented young fogey James Marriott. He marshals a range of sorrowful datapoints, from the collapse of reading-for-pleasure, to the inability of college kids to handle any textual analysis more cognitively vexing than Instagram.</span></p><p><span>If people quit reading, it&#8217;s moot who (or what) writes.</span></p><p><span>But a paradox is that our purportedly post-literate era is the most literate in history, if you judge by how much people are producing and consuming words: all the social-media posts, all the texting back and forth, all the newsletters&#8212;not to mention the word-streams on a billion podcasts and video clips.</span></p><p><span>A future without writing&#8212;as opposed to today&#8217;s decline of deep reading&#8212;is possible only if we find a technology more proficient at transmitting ideas and experiences from brain to brain. </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11004276/"><span>Brain-computer interfaces</span></a><span> are not near any such text-ending magic.</span></p><p><span>Until then, human reading contains the &#8220;off-switch&#8221; antibody: that you just don&#8217;t care as much once you discover that nobody said it. It&#8217;s a psychological reason for why writing belongs in what the economist Alex Imas </span><a href="https://aleximas.substack.com/p/what-will-be-scarce"><span>calls</span></a><span> the relational sector: &#8220;the human-intensive, provenance-rich, sometimes artisanal part of the economy where the human aspect is part of the value of the good or service itself.&#8221;</span></p><p><span>Indeed, Imas contends that excellent human writing may</span><em><span> rise</span></em><span> in value in the AI era, partly because so few people will be able to do it. In other words, the post-literate era of tongue-tied semi-literates could make the expressive few stand out. One caveat: if people develop relationships with AI companions, they may care deeply about what those entities write.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Bw7z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Bw7z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 424w, https://substackcdn.com/image/fetch/$s_!Bw7z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 848w, https://substackcdn.com/image/fetch/$s_!Bw7z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 1272w, https://substackcdn.com/image/fetch/$s_!Bw7z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Bw7z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png" width="1080" height="889" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:889,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Bw7z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 424w, https://substackcdn.com/image/fetch/$s_!Bw7z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 848w, https://substackcdn.com/image/fetch/$s_!Bw7z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 1272w, https://substackcdn.com/image/fetch/$s_!Bw7z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38687b7-fe71-479a-8302-c529f3edf42d_1080x889.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Another banger from GPTolstoy. (<em>The Artist&#8217;s Wife </em>by Albert Bartholom&#233;)</figcaption></figure></div><h3><span>THE HUMANS ARE DEAD. LONG LIVE THE HUMANS.</span></h3><p><span>A classic twist in sci-fi horror is when the main character finally twigs: all the other &#8220;people&#8221; are actually machines. Lately, I&#8217;ve felt like that character, discovering when I drop more and more texts into the AI detector that robots are everywhere.</span></p><p><span>Not only that literary award-winner, but the award citation itself. Also, an essay&#8212;on the importance of reading!&#8212;by a thinker I admire(d). Then, another piece of writing landed on my desk. Not onscreen. It landed physically on my desk: a handwritten letter, composed in pen, with &#8220;Tom&#8221; on the front of the envelope.</span></p><p><span>I have it before me now, saved as no email would be, because of the fellowship of human beings, that we mind about words&#8212;but also about another person&#8217;s effort and intent. Machines are so eloquent. But until they struggle to find the words, they won&#8217;t compose a letter I&#8217;d keep.</span></p><p><span>Human writings are an improved version of human beings. By escaping the confines of your skull, by pausing time, by revising what you&#8217;d blurt, you craft what you mean to mean.</span></p><p><span>&#8220;What I write about is other than me. As what I write is smarter than I am. Because I can rewrite it,&#8221; Susan Sontag </span><a href="https://www.nytimes.com/2000/12/18/books/writers-on-writing-directions-write-read-rewrite-repeat-steps-2-and-3-as-needed.html"><span>said</span></a><span>. &#8220;My books know what I once knew, fitfully, intermittently.&#8221;</span></p><p><span>Writing is an intelligence-increasing technology, allowing you to convert the SOS pulses of the perceiving self into objects to organize, line up, accumulate, frame, and build into more than the  capacity of your known mind. For the species, it&#8217;s an even more IQ-increasing technology, turning thoughts into a public good that never expires.</span></p><p><span>AI writing may burn books that don&#8217;t yet exist. Perhaps that&#8217;s fine. Publishers are always moaning that too many books come out. And maybe we&#8217;ve reached the end of our utility as authors: we wrote superb datasets called literature; the machines enjoyed that appetizer.</span></p><p><span>But our words are more than their feed. Writing is our testament, permitting readers to find companions among the dead, and the author to befriend the yet-to-be-born, granting anyone a way to defect from the present, and providing even the godless with a faith that something may persist after life.</span></p><p><span>We&#8212;briefly thinking machines, stuffed with noisy training data, false memories and nostalgia, jostled by fluke and circumstance, each generating words that nobody else would&#8212;we are the worst writers, and the only writers I care to read.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><div class="pullquote"><p><em>My thanks to those whose ideas and comments bettered my writing, including Alex Imas, Benjamin Bratton, Alison Snyder, S&#233;b Krier, Arthur Goemans, Max Spero &amp; Conor Griffin.</em></p></div><div class="callout-block" data-callout="true"><p>Click <a href="https://aipolicyperspectives.substack.com/p/our-rules-for-ai-writing">Our Rules for AI Writing</a> for how <em>AI Policy Perspectives</em> is approaching this issue&#8212;and add your thoughts too. You&#8217;ll find more about: The Embarrassment Rule, The Intern Rule, The Ventriloquist&#8217;s Rule, The No-Dumping Rule, and the Byline Rule&#8230;</p></div>]]></content:encoded></item><item><title><![CDATA[Our Rules for AI Writing]]></title><description><![CDATA[Society needs rules around AI writing. These are ours.]]></description><link>https://www.aipolicyperspectives.com/p/our-rules-for-ai-writing</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/our-rules-for-ai-writing</guid><dc:creator><![CDATA[Tom Rachman]]></dc:creator><pubDate>Tue, 30 Jun 2026 10:42:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LWx-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LWx-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LWx-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LWx-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg" width="789" height="612" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:789,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LWx-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LWx-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fade3f3e9-c9b1-4a9f-b17d-a4bf968dfbf2_789x612.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>The Passion of Creation</em><span> by Leonid Pasternak</span></figcaption></figure></div><p><em>We&#8217;re avid users of AI: to learn, to critique our essays, and to point out how many absurd mistakes we&#8217;ve inserted. Somehow, it&#8217;s less painful coming from a bot. </em></p><p><em>But we will never (knowingly!) publish AI writing. </em></p><p><em>Below are our rules for AI writing. <strong>Add your own in comments below</strong>. And check out <a href="https://www.aipolicyperspectives.com/p/5-rules-of-ai-writing">this essay</a> on why human writing matters.</em></p><p style="text-align: right;"><em><strong>&#8212;Tom Rachman &amp; Conor Griffin, </strong></em><strong>AI Policy Perspectives</strong></p><div><hr></div><h3><em><strong>1: The Embarrassment Rule</strong></em></h3><h4><strong>If you&#8217;d feel ashamed by exposure, don&#8217;t use AI to write. (Or credit it)</strong></h4><ul><li><p>Much as DNA testing is identifying the guilty years later, tomorrow&#8217;s AI-detection apps will expose today&#8217;s clandestine usage.</p></li><li><p>This is happening already, with suspicious readers uploading post-2022 fabrications that once seemed undetectable.</p></li></ul><h4><strong>Rule-of-thumb: Assume your AI usage will be discovered.</strong></h4><div><hr></div><h3><em><strong>2: The Intern Rule</strong></em></h3><h4><strong>Use AI in the </strong><em><strong>process </strong></em><strong>of writing. But treat its contribution as you&#8217;d treat that of a person.</strong></h4><ul><li><p>It&#8217;s fine for AI to explain a subject when you&#8217;re finding your way, much as you might consult a knowledgeable colleague. But if your final text consisted of nothing but that colleague&#8217;s points, you&#8217;d be remiss to take full credit. Likewise with AI.</p></li><li><p>It&#8217;s fine for AI (or a human research assistant) to ferret out articles, and to summarize them, especially technical material and turgid writing. It&#8217;s not fine for you to read nothing yourself.</p></li><li><p>It&#8217;s fine for AI to help brainstorm <em>your</em> ideas. It&#8217;s not fine for AI to supply all the ideas without credit.</p></li><li><p>It&#8217;s fine to have AI proofread. It&#8217;s not fine for AI to rewrite.</p></li></ul><h4><strong>Rule-of-thumb: Test whether you&#8217;re going too far by asking: Would it be wrong to accept this help from a human without crediting them?</strong></h4><div><hr></div><h3><em><strong>3: The Ventriloquist&#8217;s Rule</strong></em></h3><h4><strong>If you want to learn a subject, put it into words. When AI writes, it deletes your future wisdom.</strong></h4><ul><li><p>Sometimes you&#8217;ll just need an artifact. But remember that, if a machine is generating all of your words, you&#8217;ll end up as knowledgeable as the ventriloquist&#8217;s dummy. </p></li></ul><h4><strong>Rule-of-thumb: To speak more fluently on a topic, write about it.</strong></h4><div><hr></div><h3><em><strong>4: The No-Dumping Rule</strong></em></h3><h4><strong>Nobody wants to read outsourced AI replies.</strong></h4><ul><li><p>If someone asks your view, and you ask a chatbot, and send them a 10-page document that you identify as AI, this is burden-dumping.</p></li><li><p>Nobody should read your burden-dumps. They could&#8217;ve created those without you.</p></li></ul><h4><strong>Rule-of-thumb: If someone respects you enough to seek an opinion that you cannot offer, respect them enough to say no.</strong></h4><div><hr></div><h3><em><strong>5: The Byline Rule</strong></em></h3><h4><strong>Your byline is an oath.</strong></h4><ul><li><p>Readers assume that the named author wrote the text. If this is only partly true, your byline can be an act of deception.</p></li><li><p>Many types of writing are valuable <em>without</em> personal authorship: weather forecasts; instruction manuals; explanations about the world. Generate as much as you want. Don&#8217;t take the byline.</p></li><li><p>In other types of writing, authorship is the whole point&#8212;for instance, a school essay or a condolence letter. Never generate this.</p></li><li><p>News reporting and scientific research occupy an ambiguous midpoint, where opinions may differ. In their idealized form, each aspires to share knowledge above all else. If AI can write these findings better than the people doing the finding, there&#8217;s arguably no problem.</p></li><li><p>But a byline also establishes accountability for claims, and builds networks of thinkers. Not to mention the motivation of an author&#8217;s career ambitions and vanity.</p></li></ul><h4><strong>Rule-of-thumb: If you&#8212;or the reader&#8212;care who is named, don&#8217;t generate.</strong></h4><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/our-rules-for-ai-writing/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/our-rules-for-ai-writing/comments"><span>Leave a comment</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Q&A: Pangram CEO]]></title><description><![CDATA[How can you detect AI writing? Max Spero explains]]></description><link>https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo</guid><dc:creator><![CDATA[Tom Rachman]]></dc:creator><pubDate>Tue, 23 Jun 2026 12:07:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DFP6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DFP6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DFP6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 424w, https://substackcdn.com/image/fetch/$s_!DFP6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 848w, https://substackcdn.com/image/fetch/$s_!DFP6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 1272w, https://substackcdn.com/image/fetch/$s_!DFP6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DFP6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DFP6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 424w, https://substackcdn.com/image/fetch/$s_!DFP6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 848w, https://substackcdn.com/image/fetch/$s_!DFP6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 1272w, https://substackcdn.com/image/fetch/$s_!DFP6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608d8bba-8794-4b32-b074-ba33f72fe599_1024x572.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>I just heard a maven of the AI scene using a new verb for the addictive practice of testing all you read online in an AI-writing detector. &#8220;I&#8217;m Pangramming everything lately,&#8221; he said, referring to the popular app.</span></em></p><p><em><span>If nobody can tell who (or what) wrote the words, a centuries-long conversation among humanity degenerates. Slop sloshes through debates. Honest authors lose out, and readers wonder what a byline means. Until recently, that seemed our fate. But has AI-detection technology hit a turning point?</span></em></p><p><em><span>To find out, I caught up with the CEO of Pangram, Max Spero, a former Google software engineer who co-founded his company with an ex-roommate from Stanford. When they set out in 2023, the average citizen was still asking, &#8220;What IS a large language model?&#8221; Max had a different question: &#8220;What is society, if we can&#8217;t tell human from machine?&#8221;</span></em></p><p style="text-align: right;"><span>&#8212;</span><em><strong><span>Tom Rachman,</span></strong><span> </span></em><strong><span>AI Policy Perspectives</span></strong></p><div><hr></div><p style="text-align: right;"><em><span>[Interview condensed and edited for clarity]</span></em></p><p><strong><span>Tom: Here&#8217;s a provocative question: Why care about human writing?</span></strong><span> Perhaps the important part of writing is communicating information. Why not use AI to do that better?</span></p><p><strong><span>Max: </span></strong><span>There is</span><strong><span> </span></strong><span>a social contract between the writer and the reader. If I believe there is an idea worth sharing, I pay a cost by writing a piece of text, and the reader pays a cost by taking the time to read it. In a future where AI content lets you bypass the cost of writing your idea and formulating it, we get into a situation with perverse incentives, where people are posting total slop, and the reader is taking more time to read than the writer put into the text in the first place.</span></p><p><strong><span>Tom: </span></strong><span>I&#8217;m not convinced that&#8217;s the whole answer. I can imagine cases where it would take an hour to write a letter yourself, but two hours to prompt it with AI. If effort was the key, we&#8217;d value the AI-written letter more. I think we care about human writing for reasons beyond effort. First, if the AI author isn&#8217;t disclosed, there&#8217;s an element of deceit, and we&#8217;re offended by lying. Second, even if the person is completely transparent about AI use, we may still mind.</span></p><p><strong><span>Max:</span></strong><span> Agreed. The reader-writer relationship is ultimately a relationship between people. Another thing to think about is what can happen with AI writing at scale. It&#8217;s so easy to run deceptive operations. They used to have Russian troll farms with hundreds of people writing internet comments. They&#8217;ve replaced these with AI and LLMs, and they&#8217;re able to run at a much larger scale and push political agendas. On a micro scale you might say, &#8220;Well, if I&#8217;m looking for a news article, what do I care if it&#8217;s AI-generated?&#8221; But on a macro scale, understanding the provenance really matters. Otherwise you enable these bad actors to push their agenda.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h4><span>HOW AI DETECTORS WORK</span></h4><p><strong><span>Tom: </span></strong><span>When chatbots really took off in early 2023, one of the first parts of society to feel it was education. Students were using them to do assignments, and some educators abandoned the homework essay altogether because they had no way of reliably identifying AI-generated submissions. Can you explain why early attempts at AI detectors failed? Maybe start by explaining the term &#8220;perplexity.&#8221;</span></p><p><strong><span>Max: </span></strong><span>Early AI-detection tools tried to solve the problem by measuring </span><a href="https://www.pangram.com/blog/why-perplexity-and-burstiness-fail-to-detect-ai"><span>perplexity</span></a><span>, which is a metric of how unexpected a piece of text is. Something like, &#8220;For lunch, he ate a bowl of </span><em><span>soup</span></em><span>&#8221; is low in perplexity while &#8220;For lunch, he ate a bowl of </span><em><span>spiders</span></em><span>&#8221; is high in perplexity. Large language models work by composing probable text, making the writing much less varied, more steady, and therefore low in perplexity. But when humans write, we tend to scatter unexpected things in there, meaning our texts are much more varied, with many parts predictable, other parts slightly less so, and occasional bits that really surprise you. That variation is known as &#8220;burstiness,&#8221; as if you have spikes of surprise bursting through. AI writing is very low in burstiness. Early on, researchers saw you could use perplexity and burstiness to distinguish human from AI text pretty well, with between 95 percent and 99 percent accuracy.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zClB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zClB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 424w, https://substackcdn.com/image/fetch/$s_!zClB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 848w, https://substackcdn.com/image/fetch/$s_!zClB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 1272w, https://substackcdn.com/image/fetch/$s_!zClB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zClB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png" width="862" height="574" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0f67cde-5432-4b57-a87d-b655467e382c_862x574.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:574,&quot;width&quot;:862,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zClB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 424w, https://substackcdn.com/image/fetch/$s_!zClB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 848w, https://substackcdn.com/image/fetch/$s_!zClB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 1272w, https://substackcdn.com/image/fetch/$s_!zClB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0f67cde-5432-4b57-a87d-b655467e382c_862x574.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">AI writing is more consistently low perplexity (blue) because of its predictable patterns and word choices. The human text mixes low perplexity (blue) with moderate perplexity, and moments of high perplexity (red) bursting through. (Image: Pangram Labs)</figcaption></figure></div><p><strong><span>Tom:</span></strong><span> But you rejected that approach. Why?</span></p><p><strong><span>Max: </span></strong><span>First, that level of accuracy might sound good, but it&#8217;s still way too many errors to use in the real world. Also, the approach had several problems besides that. One is that any text that happened to appear in the LLM&#8217;s training set would seem like probable phrasing to the model, and therefore low perplexity&#8212;so something familiar to it, like the Declaration of Independence, might get flagged as AI-generated. Another problem is that people learning English tend to write more simply, meaning their output is low perplexity, and can also get flagged as AI. So, we just threw all of this research out the window, and decided to train a deep-learning classifier.</span></p><p><strong><span>Tom: </span></strong><span>This is how I </span><em><span>think</span></em><span> you </span><a href="https://arxiv.org/pdf/2402.14873"><span>built</span></a><span> your deep-learning classifier&#8212;tell me if I get it right. First, you took an open-weight base model. You fine-tuned it with hundreds of thousands of labeled examples of human text and AI-generated text in various genres. From this, the model learned to differentiate human writing from AI-generated, but imperfectly. So you added a stage of &#8220;hard negative mining,&#8221; where you took all of the cases where your model had wrongly classified a piece of human writing as artificial&#8212;false positives&#8212;and you generated AI versions of that same kind of text. You added those examples to the original fine-tuning dataset, then went back to stage one, training the open-weight model from scratch, but now including the edge cases. An algorithm ran this process as a two-stage loop, making it better and better, until you had a product.</span></p><p><strong><span>Max: </span></strong><span>What you&#8217;re describing was our process when we started Pangram. But we&#8217;ve never stopped training models. Keep in mind that we could tune models to be more sensitive to AI content, or we could tune them to be more conservative, and have fewer false positives. Avoiding an incorrect judgment that a piece of human writing was AI-generated is something we&#8217;re most careful about avoiding. So we calibrate the model to a rate of 1 in 10,000 false positives, keeping that constant while trying to improve how much AI content we catch, or the model&#8217;s &#8220;recall.&#8221; We&#8217;re always running new experiments, bringing in new datasets, new domains of writing. The goal of training new models is not to just slide this tradeoff scale; it&#8217;s to move the entire curve. And everything </span><em><span>is</span></em><span> improving. We&#8217;ve also come up with additional important techniques. The main one is training the model to </span><a href="https://arxiv.org/pdf/2510.03154"><span>differentiate</span></a><span> between full AI generation and AI edits. What we did is ask LLMs to edit text, make it better, fix grammar mistakes, summarize it. Then we measure the distance between the original text and the edited text, and we train our model with that, showing it that a certain distance from human text is equal to a light edit, and further is equal to a moderate edit. We train it to recognize that, in addition to the fully AI-generated text.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><strong><span>Tom: </span></strong><span>So you have a dashboard, where people can drop in text, and a browser extension that&#8217;ll do live checks on social media feeds like LinkedIn and Substack and X. Obviously, your products aren&#8217;t much use if they&#8217;re flinging out tons of false </span><em><span>negatives</span></em><span>, letting AI-generated content through the net. But it&#8217;s the false positives that&#8217;s worse here because a wrongful accusation could be devastating. You might have it down to an average of 1 in 10,000 false positives&#8212;but at scale, that still means mistakes. Have you been monitoring false positives in the wild?</span></p><p><strong><span>Max: </span></strong><span>Definitely; we have a tracker. If anybody comes to me, and says, &#8220;It&#8217;s a false positive!&#8221;, we&#8217;ll triage it internally. Also, on the app dashboard, we have thumbs-up/thumbs-down on the results, so we can track responses at a higher volume. But that is pretty unreliable because often it&#8217;s just somebody unhappy that their obvious AI text got flagged.</span></p><h4><span>MYSTERY OVER WHAT THE MODEL SEES</span></h4><p><strong><span>Tom: </span></strong><span>I want to dwell on the false positive rate for a moment because this concern&#8212;wrongful accusations&#8212;feels key to adoption of this technology. When you talk about 1 in 10,000 false positives, or 99.99% correct classification, that&#8217;s an average across many forms of writing. But the errors are higher for certain types of material, such as recipes, poetry, and how-to articles. Why?</span></p><p><strong><span>Max: </span></strong><span>These are all fairly formulaic types of writing, so it&#8217;s going to be lower signal. With poetry, it&#8217;s also very short texts in most cases.</span></p><p><strong><span>Tom: </span></strong><span>And your detector refuses to evaluate anything below 50 words. But when it </span><em><span>can</span></em><span> give its verdict on a document, is it evaluating the likeliness of AI at the level of each word? Or the paragraphs? Or the entirety of the text?</span></p><p><strong><span>Max: </span></strong><span>It&#8217;s looking at windows of about 200 to 500 tokens [roughly 150 to 375 words]. We have a system that does a pass, and tries to find rough boundaries between AI and human text, and then we do a finer sweep to try to find more exact boundaries, specifically for mixed text.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wMX6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wMX6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 424w, https://substackcdn.com/image/fetch/$s_!wMX6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 848w, https://substackcdn.com/image/fetch/$s_!wMX6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 1272w, https://substackcdn.com/image/fetch/$s_!wMX6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wMX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png" width="1259" height="797" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:797,&quot;width&quot;:1259,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wMX6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 424w, https://substackcdn.com/image/fetch/$s_!wMX6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 848w, https://substackcdn.com/image/fetch/$s_!wMX6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 1272w, https://substackcdn.com/image/fetch/$s_!wMX6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1089ac3d-1247-4d45-a076-c0ae35b38dd8_1259x797.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Pangram dashboard delivers its verdict on a bit of generated cringe.</figcaption></figure></div><p><strong><span>Tom: </span></strong><span>When it flags a piece of writing, your interface lets the user see &#8220;supporting evidence&#8221; that something was AI-generated or AI-assisted. That can include the telltale phrasing that people have come to associate with AI writing. But this supporting evidence isn&#8217;t actually how your system made its judgment, right? It&#8217;s a post-hoc analysis of the text that your model already flagged. Is that because deep-learning is great at somehow identifying AI, but you don&#8217;t know precisely what it&#8217;s picking up?</span></p><p><strong><span>Max: </span></strong><span>Yes, the deep-learning classifier is a black box. We don&#8217;t have a ton of interpretability into why it makes the predictions that it does. We include the supporting evidence because there&#8217;s always demand from people wondering what makes a text read as AI, and how they can train their own eyes to detect it. But the Pangram classifier is not looking only at the surface-level features. It&#8217;s finding patterns in longer-context features that an LLM will use in structuring and writing a doc.</span></p><p><strong><span>Tom: </span></strong><span>It&#8217;s aware of something exclusively human in writing that we ourselves cannot necessarily put our finger on.</span></p><p><strong><span>Max: </span></strong><span>If you&#8217;re looking for an objective measure, it might be that LLMs are better than humans in so many ways: they don&#8217;t make grammar mistakes, they often have coherent arguments, while if you look at the full distribution of human text, many humans are writing less coherent arguments. But it&#8217;s also that LLMs are a lot less diverse: if you ask them to write 100 arguments on a topic, they&#8217;re going to cluster in </span><a href="https://www.sciencedirect.com/science/article/pii/S294988212500091X"><span>one area</span></a><span>, whereas the space of </span><em><span>human</span></em><span> arguments is going to be very diverse. That&#8217;s something AI can&#8217;t replicate or reproduce.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/q-and-a-pangram-ceo/comments"><span>Leave a comment</span></a></p><p><strong><span>Tom: </span></strong><span>But AI developers will improve the quality of writing. Why think you can keep ahead of that?</span></p><p><strong><span>Max: </span></strong><span>To make a frontier model more capable, you have to instill preferences&#8212;it has to prefer to do one thing rather than something worse. But it&#8217;s these preferences that Pangram is detecting in writing. The most undetectable language models were GPT-2 and GPT-3, where they were not trained to have clear preferences, but to approximate the human distribution of writing. But today&#8217;s LLMs have stronger preferences, and this is a big reason we&#8217;re able to detect the text so easily. A caveat is that LLMs are getting more complex, so that is a headwind we must deal with. Also, the way people use AI has been changing. It went from a student saying, &#8220;Write me an essay&#8221; to somebody using an agentic harness, and saying, &#8220;Research and write a Substack article on the Strait of Hormuz.&#8221; Those are completely different use cases.</span></p><h4><span>AI DETECTION IN THE WILD</span></h4><p><strong><span>Tom: </span></strong><span>Who is using your models right now, and why?</span></p><p><strong><span>Max: </span></strong><span>We have a range of people using them. Universities use this to help with academic integrity. We also have academic conferences such as NeurIPS, the biggest AI conference, which recently </span><a href="https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track/"><span>said</span></a><span> they need position papers to be substantially human-written, and they used Pangram to flag papers that look fully AI-generated, and either desk-reject or give authors a chance to appeal. We also work with some traditional publishers, as well as anybody with datasets who&#8217;s worried about the quality. That is an underrated aspect of this: if you&#8217;re buying data, you don&#8217;t want it to just be slop; you want it to be quality from real experts.</span></p><p><strong><span>Tom: </span></strong><span>Which sectors aren&#8217;t dealing adequately with AI detection?</span></p><p><strong><span>Max: </span></strong><span>Publishing has been really slow to catch up on this. There were a couple of major scandals, like the book </span><em><a href="https://www.nytimes.com/2026/03/19/books/shy-girl-book-ai.html"><span>Shy Girl</span></a></em><span> that people called out as being AI-generated. It got pulled from shelves. A lot of publishers don&#8217;t have a strong AI policy. Some have decided that, as long as the text is good, we&#8217;re going to evaluate it on its merits rather than whether it&#8217;s AI-generated or not.</span></p><p><strong><span>Tom: </span></strong><span>My sense is that most publishers </span><em><span>would</span></em><span> mind if a piece of writing is AI-generated. If anything, I&#8217;d guess that there could be a lag with many in publishing </span><a href="https://commonwealthfoundation.com/commonwealth-short-story-prize-2026/"><span>unaware</span></a><span> that detection technology now works well. I suspect there&#8217;s also an aversion to technological involvement here&#8212;that you should trust authors protesting their innocence and not trust an app. I sympathize with that. Anyone accused deserves a fair chance to defend themselves. But failing to take this technology seriously is even more harmful to human authors who actually write their own stories.</span></p><p><strong><span>Max:  </span></strong><span>They lose something real and tangible.</span></p><p><strong><span>Tom: </span></strong><span>What are cases where you&#8217;d consider it fine to use AI to write?</span></p><p><strong><span>Max: </span></strong><span>I don&#8217;t really consider it &#8220;using AI to write.&#8221; You can publish AI content, and that&#8217;s fine, as long as it&#8217;s properly disclosed. But if I&#8217;m writing an email to somebody, it&#8217;s somewhere between impersonal and disrespectful to use AI to write that. AI is a useful tool: we should be using it to generate text, synthesize text, and it&#8217;s even reasonable to share AI output. But if we&#8217;re talking about &#8220;using AI to write,&#8221; I&#8217;d caution</span><strong><span> </span></strong><span>against that framing.</span></p><h4><span>THE FUTURE OF &#8220;WRITING&#8221;</span></h4><p><strong><span>Tom: </span></strong><span>We&#8217;ve been talking about this as a binary: either human-generated or AI-generated. But maybe we&#8217;re destined for a future of cognitive blending with AI, where people outsource parts of thinking, much as we now outsource parts of our memory to machines. Maybe the idea of a human-or-machine binary is a short-lived prospect.</span></p><p><strong><span>Max: </span></strong><span>Yes, I could see that, especially with so many applications pushing AI assistance on users. But I think there would still be degrees of human input that would differentiate such writing from the entirely AI-generated.</span></p><p><strong><span>Tom: </span></strong><span>So there is a long-term future for human writing?</span></p><p><strong><span>Max: </span></strong><span>Definitely. People just care about it. To me, it doesn&#8217;t matter how good the AI gets. There&#8217;s still a lot of people who care, and a lot of industries, and types of writing, where it really matters that something&#8217;s human.</span></p><p><strong><span>Tom: </span></strong><span>What would those be for you?</span></p><p><strong><span>Max: </span></strong><span>Fiction is one. But also, think of an essay contest. Is the goal really to find the best essay? Or is it to reward writers who have taken the time to write something compelling?</span></p><p><strong><span>Tom: </span></strong><span>Well, I think it&#8217;s probably both. If it&#8217;s an essay contest in a school, say, you&#8217;re rewarding effort in part, but you&#8217;re also rewarding quality. Or else you&#8217;d just give the prize to the kids who tried hardest, and not the ones who wrote well.</span></p><p><strong><span>Max: </span></strong><span>Yeah, but if the goal was purely to find the best essay, then maybe we should have LLMs generate 100 completely.</span></p><p><strong><span>Tom: </span></strong><span>Unless there&#8217;s something in the best writing that requires a human author.</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption"><em>Coming next:</em> <strong>5 Rules of AI Writing, </strong>An essay by Tom Rachman</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Stop Shouting. Start Policymaking. ]]></title><description><![CDATA[AI tools can show politicians what the public wants&#8212;not just where people disagree]]></description><link>https://www.aipolicyperspectives.com/p/stop-shouting-start-policymaking</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/stop-shouting-start-policymaking</guid><dc:creator><![CDATA[AI Policy Perspectives]]></dc:creator><pubDate>Wed, 17 Jun 2026 12:36:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4Gme!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote><p><em>Guest post from <strong>Carl Miller, </strong>a technologist and writer based at the think tank Demos in London, and <strong>Beth Goldberg</strong>, head of R&amp;D at Jigsaw, a Google technology incubator, who teaches at Yale.</em></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4Gme!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4Gme!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 424w, https://substackcdn.com/image/fetch/$s_!4Gme!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 848w, https://substackcdn.com/image/fetch/$s_!4Gme!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 1272w, https://substackcdn.com/image/fetch/$s_!4Gme!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4Gme!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4Gme!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 424w, https://substackcdn.com/image/fetch/$s_!4Gme!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 848w, https://substackcdn.com/image/fetch/$s_!4Gme!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 1272w, https://substackcdn.com/image/fetch/$s_!4Gme!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb14ac0b6-ec9a-4fd5-bfe5-353588c4400f_2048x1117.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Sun sparkled on the red double-deckers as busy Londoners hopped on and off, glancing at a poster as they went. Few understood its significance.</strong></p><p>Behind a glass panel at the bus stop, the poster asked, &#8216;Who Cares About Care?&#8217;, above photographs of locals, and a QR code. What seemed a modest advertisement was actually part of a ground-breaking experiment that aims to transform how British local governments work by changing how they <em>listen</em>. This project could be a hint, and an inspiration, to governments around the world about how to transform themselves in the age of AI.</p><p>This question posed by Camden Council, which oversees a borough of north London, regards one of <em>the </em>key responsibilities of local governments in the United Kingdom: handling adult social care, which ranges from assisted-living homes, to mental health support, to nursing. Local officials, rather than simply convening a town-hall meeting to hear from whoever turned up, had established an online space where residents could state their priorities, and vote on the priorities of others.</p><p>But this project had resonance far beyond north London. For more than a decade&#8212;as many people came to suspect that technology just pulled societies apart&#8212;a scattering of digital activists, engineers, and academics were trying to deliver the opposite, employing artificial intelligence to detect common ground in the public, and root out policy choices that might unite people.</p><p>The long journey to &#8216;Who Cares About Care?&#8217; started years before, and thousands of miles away.</p><h4>&#8216;BRIDGING&#8217; TECHNOLOGY</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cdzB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cdzB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 424w, https://substackcdn.com/image/fetch/$s_!cdzB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 848w, https://substackcdn.com/image/fetch/$s_!cdzB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 1272w, https://substackcdn.com/image/fetch/$s_!cdzB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cdzB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png" width="1456" height="970" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:970,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cdzB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 424w, https://substackcdn.com/image/fetch/$s_!cdzB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 848w, https://substackcdn.com/image/fetch/$s_!cdzB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 1272w, https://substackcdn.com/image/fetch/$s_!cdzB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7324db4-5142-4af1-b611-8174713f09d4_2048x1365.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Students occupying the Taiwanese legislature during the 2014 Sunflower student movement. (Credit: Artemas Liu/CC)</figcaption></figure></div><p>Taiwan&#8217;s Sunflower student movement of 2014 swept a new breed of politician into power: the civic hacker. Previously, civic hackers had existed on the margins of political life, trying to use technology to make officialdom more transparent and accountable. Riding a wave of disenchantment, they found themselves in government, and created an experiment called vTaiwan, which aimed to bring ordinary citizens into political decision-making.</p><p>Their insight was to see political conflicts, anger and polarisation as a problem of <em>information</em>. Old-fashioned arguments, they thought, failed to generate a clear signal to base decisions on, instead underscoring where people differed. What policymakers needed was a better way to find what people held in common.</p><p>Change the information circuitry, they thought, and you change the politics electrified by it. AI proved fundamental to that rewiring.</p><p>When facing a tense political debate, vTaiwan convened various civic groups on a public-deliberation platform called <a href="http://pol.is">pol.is</a>, where people could express hopes and fears, bugbears and hang-ups, and also hear comments from others, which they could thumbs-up or thumbs-down. The algorithm used &#8216;bridge-based ranking,&#8217; which is one of<em> </em>the most important technologies that democracies need to evolve and survive.</p><p>How <a href="http://pol.is">pol.is</a> works is to integrate all proposals and votes that participants contribute, then locate each of those participants on an ideological map. Every cluster represents a tribe of sorts. Rather than serving up engaging content, <a href="https://www.belfercenter.org/sites/default/files/pantheon_files/files/publication/TAPP-Aviv_BridgingBasedRanking_FINAL_220518_0.pdf">bridging algorithms</a> surface content that people from different ideological starting-points agree with, while blocking that which risks hardening group distinctions into long-running enmities The <a href="http://pol.is">pol.is</a> platform resembles an open forum like Reddit, where you can come anytime and add your comments and upvotes to an issue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wCRa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wCRa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 424w, https://substackcdn.com/image/fetch/$s_!wCRa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 848w, https://substackcdn.com/image/fetch/$s_!wCRa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 1272w, https://substackcdn.com/image/fetch/$s_!wCRa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wCRa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png" width="1024" height="314" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:314,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wCRa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 424w, https://substackcdn.com/image/fetch/$s_!wCRa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 848w, https://substackcdn.com/image/fetch/$s_!wCRa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 1272w, https://substackcdn.com/image/fetch/$s_!wCRa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d795617-f502-4f00-a8d9-f00fc588e5e5_1024x314.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The pol.is interface, used in a UK survey on political campaigning. (Credit: Demos/Open Rights Group) </figcaption></figure></div><p><a href="http://pol.is">pol.is</a> showed that the hope of Taiwan&#8217;s digital activists had been well-founded: if you rewire the information environment, you may find common ground that had been invisible. The process broke a deadlock on Uber regulation, and then a six-year impasse over the sale of alcohol online. It happened again with regulation over fintech, online gambling, cryptocurrency, and e-scooters.</p><p>Inspired by this success, others tested further ways to employ technology for constructive deliberation. Researchers at Google DeepMind trained an LLM to act as a mediator between people discussing divisive topics, and found it more effective than human mediators. They called it the <a href="https://www.science.org/doi/10.1126/science.adq2852">Habermas Machine</a>, after the great German theorist of the public square, Jurgen Habermas.</p><p>Another experiment with LLMs <a href="https://www.nature.com/articles/d41591-024-00073-7">showed</a> they could reduce belief in conspiracy theories by 20% via one-to-one interactions. Harvard&#8217;s Applied Social Media Lab built <a href="https://www.youtube.com/watch?v=tp14UK3B1qA&amp;t=18s">Frankly</a>, a &#8216;video-based discourse platform&#8217; for collective problem-solving that includes nudges to encourage people to speak, and clever ways of jumbling up participants during breakout sessions, along with techniques to gather proposed solutions as the conversation proceeds. Another platform is <a href="https://psi.tech/">PSi,</a> created to bring thousands of people into video-based online discussions, using AI to track polarisation and consensus in real-time.</p><p>Beyond using AI to find points of agreement, researchers have also developed technology for consensual decision-making. The <a href="https://ethelo.com/technology/">Ethelo</a> platform focussed on how to turn the last stage of any discussion into &#8216;convergence and closure&#8217;. Participants discuss possible options, weighing the importance of underlying issues&#8212;say, cost or time. An algorithm searches for the outcome that leaves everyone roughly equally happy. There are also tools such as <a href="https://crown-shy.com/products/comhairle">Crownshy</a>, which gathers open-source civic-tech tools, and strings them together into seamless workflows, so that policymakers can more easily procure and use them.</p><p>But are governments using these methods?</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h4>EARLY ADOPTERS</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6JyS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6JyS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 424w, https://substackcdn.com/image/fetch/$s_!6JyS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 848w, https://substackcdn.com/image/fetch/$s_!6JyS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 1272w, https://substackcdn.com/image/fetch/$s_!6JyS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6JyS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png" width="1456" height="773" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:773,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6JyS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 424w, https://substackcdn.com/image/fetch/$s_!6JyS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 848w, https://substackcdn.com/image/fetch/$s_!6JyS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 1272w, https://substackcdn.com/image/fetch/$s_!6JyS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dac0cb1-ebae-4523-8916-6e2e9d0cacf1_2048x1087.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Kentucky city of Bowling Green has faced tensions over the fast-growing population. An AI-assisted public consultation engaged residents in planning for their shared future. (Credit: <a href="https://www.whatcouldbgbe.com/about-the-project">What Could BG Be?</a>)</figcaption></figure></div><p>Nestled in the rolling pastures of southern Kentucky, the city of Bowling Green expects its population of 80,000 to <em>double</em> by the year 2050, owing to a strong job market that is attracting both locals and resettled refugees. This explosive growth has sparked a range of anxieties for residents and local leaders. Farmers worry about land preservation, educators about ballooning class sizes, and lifelong residents about the dilution of their small-town identity.</p><p>Bowling Green has a complex political makeup, home to a progressive university, a thriving refugee resettlement hub, and deep-rooted agricultural and industrial traditions, including the Corvette factory. This rapid population growth risks deepening cultural and political fault lines, though the community has thus far demonstrated pragmatic collaboration.</p><p>But inclusive conversations about managing this radical demographic change weren&#8217;t happening in traditional democratic fora. Town halls typically attracted fewer than a dozen participants, usually a non-representative sample of highly motivated detractors.</p><p>Bowling Green refused to be another story of the loudest voices setting outcomes for the rest. Instead, last year, the city ran what was then the largest digital town hall in American history.</p><p>How did the city leaders get nearly 8,000 residents to engage in this enormous question of what Bowling Green should look like in 25 years? The local government bypassed standard surveys or focus groups, instead working with <a href="http://pol.is">pol.is</a> (the deliberative platform used in Taiwan), along with <a href="https://jigsaw.google.com/">Jigsaw</a> (the Google incubator for social impact), the local lynchpin <a href="https://www.innoengine.co/">Innovation Engine</a> and more than 100 community leaders, from the head of the library to the manager of the brewery, who served as local &#8216;listening partners&#8217;.</p><p>This diverse coalition jointly designed a campaign called <a href="https://www.whatcouldbgbe.com/">&#8216;What Could BG Be?&#8217;</a> that engaged the public throughout the monthlong process. The campaign involved regional media agencies, community influencers, and grassroots leaders. Roughly 10 percent of the city contributed 4,000 distinct proposals and cast over a million votes, rewiring Bowling Green&#8217;s circuits for public engagement via AI-enabled dialogue.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8nLm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8nLm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 424w, https://substackcdn.com/image/fetch/$s_!8nLm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 848w, https://substackcdn.com/image/fetch/$s_!8nLm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 1272w, https://substackcdn.com/image/fetch/$s_!8nLm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8nLm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png" width="899" height="1020" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1020,&quot;width&quot;:899,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8nLm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 424w, https://substackcdn.com/image/fetch/$s_!8nLm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 848w, https://substackcdn.com/image/fetch/$s_!8nLm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 1272w, https://substackcdn.com/image/fetch/$s_!8nLm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F096b3137-f520-4fa2-804c-53afa6592df4_899x1020.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The project discovered where residents strongly agree&#8212;and where differences remain. (Credit: <a href="https://www.whatcouldbgbe.com/about-the-project">What Could BG Be?</a>)</figcaption></figure></div><p>The organizers did not assume that digital access meant all were included, so they listened through ethnographic conversations across the community&#8217;s margins, engaging residents at refugee centers, halfway houses, rehab facilities, and senior living centres. By localising and translating outreach into nine major languages, they transformed personal experience into strategic civic engagement, allowing vulnerable groups, such as recentlydemo resettled Afghan refugees, to securely add their voices via a geofenced, anonymous platform.</p><p>The underlying bridge-based ranking algorithm prioritised points of hidden consensus rather than amplifying the friction typical of national political discourse. Specifically, Jigsaw&#8217;s <a href="https://jigsaw-code.github.io/sensemaking-tools">Sensemaking</a>&#8212; a suite of AI tools designed to help gather and understand public opinion&#8212;identified overwhelming nonpartisan support for practical initiatives, from synchronizing traffic lights on major arteries, to expanding eldercare, to developing community riverfront spaces. Ultimately, <em>96 percent </em>of local leaders reported that the process had given them a more precise, actionable mandate to represent their constituents. When leaders use technology to deliberately listen rather than just broadcast, they may uncover surprising ways to act with broad-based support.</p><p>The British initiative similarly hoped to uncover unifying signals in the public. Like the Bowling Green invitations in coffee shops and churches to participate in a conversation about the future, that London bus-stop poster was an invitation to partake in a new kind of democratic governance.</p><p>The project, Waves&#8212;led by the British think-tank Demos, and including local governments&#8212;also uses bridge-based ranking. As in Taiwan, this was intended from the start to be more than a civic hackers&#8217; dream but a process grounded in the realities of local government.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mb8_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mb8_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 424w, https://substackcdn.com/image/fetch/$s_!mb8_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 848w, https://substackcdn.com/image/fetch/$s_!mb8_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 1272w, https://substackcdn.com/image/fetch/$s_!mb8_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mb8_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png" width="1456" height="641" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:641,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mb8_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 424w, https://substackcdn.com/image/fetch/$s_!mb8_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 848w, https://substackcdn.com/image/fetch/$s_!mb8_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 1272w, https://substackcdn.com/image/fetch/$s_!mb8_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bbc12b7-6f3f-4df2-ae64-8e3533f74f14_1859x818.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Waves project in north London aspires to be the first of many AI-enhanced public consultations across Britain. (Credit: <a href="https://who-cares.commonplace.is/">Who Cares</a> project) </figcaption></figure></div><p>For six months, local officials joined technologists and think tankers to test and reshape every part of Waves. What they settled on was a process that starts by reaching as many residents as possible to understand the range of views about a question. The question that the Camden council asked locals was: &#8216;What matters most to you and your communities when you think about care and support?&#8217;</p><p>Then, Camden Council and the rest of the Waves team convened a much smaller but representative group of residents for a focussed discussion to turn those ideas into detailed proposals. These proposals were tested back at scale with the wider community, and what emerged was fed into a further phase of deeper deliberation where the group refined their conclusions. Finally, those conclusions were shared back with the wider community. The process moved in this way from a very wide process to a narrower and more focussed one, then back to a wide process, and then a narrow one. The process resembled a wave.</p><p>AI supports Waves in two main ways. During the &#8216;wide&#8217; phases of public engagement, its main use is to get as many people into a single online space as possible, and then to use bridge-based ranking to reach consensus as emphatically as possible. During the &#8216;narrow&#8217; stages of deliberation, the role of the tech is more subtle. Conversations are held over video chat, and tech helps facilitators spot emerging themes, synthesise insights across lots of conversations at once, and help to identify the priorities, decisions and actions these conversations identify.</p><p>The first phase of <a href="https://who-cares.commonplace.is/">Who Cares?</a> ran in September and October, and included more than 1,500 residents, <a href="https://who-cares.commonplace.is/en-GB/proposals/v3/taking-part?step=step1">59%</a> of whom had never taken part in a decision-making process in the local government before. From November to December, a panel of 41 Camden residents came together for 10 hours of discussion in the second phase. In January and February, the wide phase then reopened, with 550 residents sharing their thoughts on the priorities. The resident panel met again for a further 12 hours of discussion between February and March, reviewing considerations from Phase 3, trade-offs around funding, and finally establishing their expectations from Camden Council, the workforce, community, and individuals. As you read this, the council is reviewing those results. The next stages await.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/stop-shouting-start-policymaking?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/stop-shouting-start-policymaking?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>Another deployment of Waves is now underway in South Staffordshire, in England&#8217;s West Midlands outside Birmingham, bringing together people to discuss planning policy and the thorny question of where new houses should be built. The longer-term ambition is to scale this up: to make digital democracy not only provably effective for local governments around the world, but also affordable.</p><p>If it works, these case studies will expand into dozens of deployments next year, expanding beyond local government into charities, unions, and other membership organisations that need to listen to large numbers of people, and turn their preferences into policy decisions. The ambition is to turn Waves into a flood.</p><p>But, after a decade of civic-tech experiments around the world, and a few bold attempts to apply them, the question is: Can AI make democracies stronger?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IY_w!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IY_w!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 424w, https://substackcdn.com/image/fetch/$s_!IY_w!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 848w, https://substackcdn.com/image/fetch/$s_!IY_w!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 1272w, https://substackcdn.com/image/fetch/$s_!IY_w!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IY_w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IY_w!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 424w, https://substackcdn.com/image/fetch/$s_!IY_w!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 848w, https://substackcdn.com/image/fetch/$s_!IY_w!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 1272w, https://substackcdn.com/image/fetch/$s_!IY_w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43882151-9b9f-4bce-b6a2-bd7a27695fac_1024x559.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">(Credit: Gemini)</figcaption></figure></div><h4>A RIDE TO THE FUTURE</h4><p>The transformative power of civic tech&#8212;encouraging conversations, furnishing mediation, presenting politicians with better policy ideas&#8212;comes from identifying common ground among the public. But pilots and test-cases are one thing. To make a difference, such processes would need to become the norm.</p><p>Technology is going to be the easy part of democratic renewal. The question is whether political systems embrace these possibilities at scale.</p><p>If there&#8217;s one thing people dislike it&#8217;s a talking-shop, whether AI-mediated or not. People want <a href="https://medium.com/jigsaw/a-new-perspective-on-human-agency-for-the-ai-era-cd785faab026">agency</a>, and that means being able to connect what they say to what is decided; and to connect what they agree upon to what is eventually done.</p><p>There are a number of reasons why connecting civic tech to power is difficult. The first is that it&#8217;s  an act of direct democracy. This can sit uncomfortably alongside the idea of representative democracy, with politicians elected to make difficult decisions on the public&#8217;s behalf. What happens if a digital mandate produces an outcome that elected politicians feel violates their mandate? Which should take priority? For digital democracy to scale within political systems, we need clarity on how these different sources of legitimacy should interact.</p><p>Another challenge is public assurance&#8212;that is, gaining trust for these tools. Applying new technologies in government is also about transparency, cost and, especially when it comes to AI, ensuring that the process is safe, unbiased, and endorsed by a population that may be suspicious. This means impact assessments and compliance procedures, and other detailed processes to help local governments become familiar and comfortable with this.</p><p>Thankfully, AI can help with navigating new systems, both designing, testing and helping streamline the bureaucratic adoption. Furthermore, the relevant parties should be involved in designing their own new systems, helping assure that they are both informed and invested from the start. It&#8217;s time-consuming but invaluable.</p><p>The journey that began more than a decade ago in Taiwan, and passed through the rolling pastures of Kentucky and a bus stop in Camden, is only beginning. After years in which political systems seemed to be ripping themselves apart, trapped in rancour and mistrust, we should feel optimistic here. But it&#8217;s time to rewire the information systems on which politics operates.</p><p>Humans have more in common than we think. Sometimes, technology can help us rediscover that.</p><div><hr></div><p><em><strong>Carl Miller</strong> is a technologist and writer, founder of the Centre for the Analysis of Social Media at Demos and the information integrity lab CASM Technology. <strong>Beth Goldberg</strong> is the Head of R&amp;D at Jigsaw, a Google incubator that builds technologies to give people greater agency, and teaches about AI at Yale&#8217;s Jackson School of Global Affairs.</em></p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;956649bb-39ef-4d03-8a43-2b3c29288cf7&quot;,&quot;caption&quot;:&quot;Everywhere, people grumble about the government: that politicians care only about themselves; that bureaucrats gum up the system; that taxpayers get fleeced. Even in wealthy countries, nearly two in three people are dissatisfied with how democracy is working.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Governments Are Struggling. Can AI Help?&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:12431790,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;bio&quot;:&quot;AI Policy Writer @ Google DeepMind &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fad94b7d-013b-4773-98cb-b9014a1857b8_1024x1024.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:101194130,&quot;name&quot;:&quot;AI Policy Perspectives&quot;,&quot;bio&quot;:&quot;Reflections on AI policy, governance, and more. https://www.aipolicyperspectives.com/ &quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0Byl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9146f89b-5561-4adb-bf89-cc21b508c264_667x374.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-06T11:03:24.773Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!kyMW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F609fc3c7-b119-42ee-aa44-aba946807ee5_2648x2366.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/how-ai-fixes-government&quot;,&quot;section_name&quot;:&quot;Interviews &quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:183574277,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:19,&quot;comment_count&quot;:4,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div>]]></content:encoded></item><item><title><![CDATA[The World Cup & AI]]></title><description><![CDATA[Technology promised to fix football&#8212;but fans are raging. Here&#8217;s what the bungled rollout tells us]]></description><link>https://www.aipolicyperspectives.com/p/the-world-cup-and-ai</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/the-world-cup-and-ai</guid><dc:creator><![CDATA[AI Policy Perspectives]]></dc:creator><pubDate>Wed, 03 Jun 2026 08:01:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!QpTH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QpTH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QpTH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 424w, https://substackcdn.com/image/fetch/$s_!QpTH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 848w, https://substackcdn.com/image/fetch/$s_!QpTH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 1272w, https://substackcdn.com/image/fetch/$s_!QpTH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QpTH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QpTH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 424w, https://substackcdn.com/image/fetch/$s_!QpTH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 848w, https://substackcdn.com/image/fetch/$s_!QpTH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 1272w, https://substackcdn.com/image/fetch/$s_!QpTH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41faee90-f815-443b-a7df-57ebd534c6f9_1024x559.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>&#8220;Did you see the game?!&#8221;</em></p><p><em>The FIFA World Cup kicks off next week, triggering versions of this question in dozens of languages for the following month. The question may be gleeful; often, it&#8217;ll be grumpy. While moaning is a rite of fandom, the laments have worsened in recent years because of a technological villain named VAR.</em></p><p><em>The educator and passionate football fan<a href="https://www.nomoremarking.com/"> Daisy Christodoulou</a>, of school-software platform No More Marking, became so indignant that she wrote an entire book called </em>I Can&#8217;t Stop Thinking About VAR<em>. Now, she argues that you too should think about VAR, because its coming woes at the World Cup will hold insights for the AI era.</em></p><p style="text-align: right;"><em>&#8212;Tom Rachman, </em>AI Policy Perspectives</p><div><hr></div><p>By DAISY CHRISTODOULOU</p><p></p><p><strong>For English football fans, there is one legendary refereeing error</strong>: Diego Maradona&#8217;s &#8220;Hand of God&#8221; goal in the World Cup of 1986, when the 5-foot-5 Argentinian great leaped as if to head the ball&#8212;only to <a href="https://youtu.be/FRAbNlPS2MI?si=ac8UKF9ncaqgL13q&amp;t=176">palm</a> it past the goalkeeper. If the referee had seen, he would have disallowed the goal. But he did not, and lacked help from technology. The goal stood, and the English never got over it.</p><p>Only at the 2018 World Cup did the sport adopt technological reviews via VAR, the Video Assistant Referee <a href="https://www.theifab.com/laws/latest/video-assistant-referee-var-protocol/">system</a>, which most elite national leagues have since introduced. For high-stakes calls, such as whether a goal was valid or if there should&#8217;ve been a penalty, off-field officials review the footage, alerting the referee to any factual mistakes or if they see &#8220;a clear and obvious error&#8221; in a subjective decision.</p><p>The upcoming World Cup will be the third to feature VAR, and the system is still being tweaked. To many fans, VAR is now a source of utter rage, having caused more problems than it solved, and prompting hostile chants that ring out across stadiums.</p><p>You might conclude that football fans will never be satisfied: they spent years complaining about unchecked errors, and now complain about the technological fix. But the problems with VAR&#8212;sure to rage throughout this World Cup&#8212;say something deeper about tech solutions to human problems. Sometimes, we don&#8217;t realise the ramifications until too late.</p><h2>The accuracy/speed trade-off</h2><p>Before VAR, managers, players and fans repeated that they just wanted &#8220;more right decisions&#8221;. The assumption was that accuracy was the absolute good, and that it was acceptable to spend time on a decision if that would improve its quality.</p><p>This is an assumption in many walks of life. We tend to see the speed of a decision in opposition to its quality: that faster decisions are more likely to be wrong, and slower decisions are more likely to be right. But speed may itself be an aspect of quality. That is true in sports, where an instant decision is especially valuable, allowing the game to keep flowing.</p><p>The average time for a VAR check is under a minute, but the longest in English football was <a href="https://www.bbc.co.uk/sport/football/articles/c7vznjl7l4do">eight minutes</a>. A series of checks significantly slows the game, breaking its rhythm. Every fan and player knows that VAR frequently overturns goals after a check, so the review process has also tempered celebrations, among the most joyous parts of any sport.</p><p>Another issue is &#8220;ghost minutes&#8221;. If VAR is checking an incident, and the ball remains in play, the game carries on during the review. If the reviewers and the referee judge that the incident <em>was</em> a foul, then everything that happened after it is cancelled. Sometimes, these passages of play are longer than a minute. It all contributes to a sense that what is taking place on the pitch is not a concrete reality to be enjoyed or endured in the moment, but something provisional.</p><p>Part of the reason football is such a popular game is because it is fast-paced, fluid and doesn&#8217;t have the same natural breaks as tennis or cricket or American football. For that reason, VAR may be more destructive than equivalent systems in other sports. The stated preference of players and fans was &#8220;more right decisions&#8221;; the revealed preference is that speed matters.</p><p>Speed matters in other areas of life too. The UK has a particularly ponderous system of planning regulations. With big building projects, speed isn&#8217;t just a bonus, but an aspect of the project&#8217;s quality. If a road, railway, or power station runs 10 years over schedule, you&#8217;ve lost 10 years of its use. It&#8217;s also less likely to be relevant and up-to-date.</p><p>One response is better technology, such as <a href="https://mhclgdigital.blog.gov.uk/2025/06/12/extract-using-ai-to-unlock-historic-planning-data/">using AI</a> to digitise paper planning documents. This points to something weird about VAR. The typical expectation is that humans are slow and technology is fast. Think of online bank loan applications: an algorithm can approve or reject you almost instantly, whereas human review used to take days. But with VAR, it is the other way round: human refereeing gave instant decisions, but the technology added delay.</p><p>This is not true of all technology in football. A few years before VAR, the Premier League introduced technology to judge if the ball had crossed the goal-line or not. It has worked well, partly because it makes few errors, but also because its decisions are automated and speedy. As soon as an incident happens, the referee gets a buzz on his watch with the result. If the ball didn&#8217;t cross the line, the game flows. If it did, a goal is immediately awarded. That&#8217;s the standard we should be aiming for: instant and automatic decisions with a high degree of accuracy.</p><p>Goal-line decisions are objective, so lend themselves to automation. But reaching this standard of automatic accuracy for more subjective decisions, such as a handball, will be fiendishly complex, and risks endless disputes and ire. In education, for instance, multiple-choice exams have objective scoring criteria, so automating the marking process with technology works well, saves teacher time, and eliminates human error. But creative writing is more subjective. Trying to speed up the process through automation won&#8217;t resolve the underlying disagreements, and may make them worse.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h2>The consistency/common sense trade-off</h2><p>It might seem odd to say that handball is a subjective decision. Surely it is obvious if a ball has hit someone&#8217;s hand or not? But the past few years of VAR have created extraordinary confusion about this fundamental rule.</p><p>Before VAR, the written rule was short and simple, just 20 words long. Referees were allowed to use their own judgement to decide whether a handball had been deliberate, and therefore a foul, or if it had not.</p><p>The advantage of empowering an individual decision-maker is that you get to incorporate human discretion and common sense. The disadvantage is that you will get inconsistency, not just between different referees but even within the same referee&#8217;s set of decisions.</p><p>One way this has been addressed&#8212;even before technology&#8212;is by restricting human discretion. In many legal systems, judges don&#8217;t have complete latitude in sentencing decisions. They have guidelines because unconstrained discretion often meant that two people who had committed similar crimes received different punishments. Inconsistencies like this risk bringing systems into disrepute.</p><p>Since VAR arrived, the written rule about handballs has expanded to 11 times its original length, full of subclauses trying to cover every possible way a ball might hit a hand, and every possible context in which it should result in a foul. That is because scrutiny of video replays has meant attention to every last detail of contact, movement, and placement that was impossible to parse in a fluid real-time experience. If a player jumps in the air to head the ball and uses their arms for leverage, and then the ball deflects off an opponent&#8217;s head onto his arm, is that handball? Is it enough of an offence to justify the award of a penalty kick, which gives the team a 75% chance of scoring a goal in a game where scoring goals is rare and difficult?</p><p>In a game from September 2020, the Newcastle striker Andy Carroll headed the ball onto the arm of the Spurs defender Eric Dier. Dier&#8217;s back was turned to Carroll, and he could not even see the ball. After a VAR check, the referee awarded a penalty. The Newcastle manager, Steve Bruce, who benefitted from the decision, <a href="https://www.youtube.com/watch?v=2t96fICL8Bo">said</a>: &#8220;If you&#8217;re going to tell me that is handball then we all may as well pack it in. It&#8217;s a nonsense, a nonsense of a rule. It&#8217;s gone for us today&#8212;however, it&#8217;s ludicrous.&#8221;</p><p>A newspaper match report <a href="https://www.theguardian.com/football/2020/sep/27/tottenham-newcastle-united-premier-league-match-report">said</a>, &#8220;This was perhaps the worst among a long list of maddening recent decisions brought about by a rule that goes against the sport&#8217;s very essence.&#8221;</p><p>In previous years, this kind of handball infraction might never have been spotted, let alone punished. The problem has become so acute that the law is now constantly revised, and applied differently in different tournaments. In the English Premier League, referees err on the side of leniency. In European and international tournaments, referees are stricter.</p><h2>Solving the wrong problems</h2><p>Attempts to improve handball enforcement have often fallen into a classic trap: deploying an impressive technological innovation that doesn&#8217;t address the underlying problem. At Euro 2024, referees were given access to a sensor inside the ball that could detect whether it made contact with a player&#8217;s hand. Similar systems work well in cricket. But in football, the crucial question is rarely whether the ball touched the hand; it&#8217;s whether the handball was deliberate or significant.</p><p>If you automated all handball decisions based on data from sensors, it could transform the game. Technological automation of other decisions in society, such as those in law enforcement, could present comparable problems. Suppose you could automatically fine anyone who submits a tax return with a discrepancy. Would this reduce tax evasion? Or just prompt tax evaders to avoid discrepancies?</p><p>Or imagine a school that wants to crack down on absenteeism. It is relatively easy to create an online tracking system that bombards parents with automated texts or requests for evidence every time they want to book a dentist&#8217;s appointment for their child. It is much harder to create a system that identifies and addresses the early signs of chronic disengagement.</p><p>Often, the questions that technology finds easiest to answer are not the ones we care about the most. Automation <em>can</em> solve problems. It can also create new ones, with aggressive enforcement of rules where behaviour is easiest to quantify, not where harms are greatest.</p><p>Increased scrutiny and detail of the new handball rule has not even succeeded in providing greater consistency. Fans still point to infractions in the same season&#8212;sometimes even the same week or the same game&#8212;that look similar, yet resulted in opposite decisions. We have the worst of both worlds: a system that has reduced human discretion and lessened simplicity, but failed to improve consistency.</p><p>There&#8217;s a fascinating analogy here with artificial intelligence. The earliest AI methods were rules-based: experts tried to write down every rule a skilled human might apply when making a decision, in much the same way that football&#8217;s rulemakers now attempt to write down every possible way a ball might make contact with a hand.</p><p>These rule sets to create early AI multiplied, yet they couldn&#8217;t capture what a human was actually doing. The breakthrough came not from more or better human-crafted rules, but from giving the machine vast quantities of data and letting it intuit the rules itself. AI models trained in this way proved better at capturing and representing the tacit knowledge that experts hold but find difficult to represent in words.</p><p>One caveat here is that many real-world decisions require deep <a href="https://www.aipolicyperspectives.com/p/good-under-the-hood">moral reasoning</a> to weigh conflicting human wants. Even training an AI model on vast amounts of data will not alone align the system&#8217;s behaviour with all our preferences. So, which judgements can we automate with technology? And when would we prefer fallible &#8220;humans in the loop&#8221;? Safety researchers are considering such questions with urgency. Perhaps they could learn something from watching a little football.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h2>World Cup 2030?</h2><p>By now, the issues with VAR cannot be written off as teething problems. We have had nearly a decade of adjustments, but can anyone feel confident that this coming World Cup will see no glaring refereeing errors?</p><p>We are stuck with VAR this time. But if we look further ahead, to the 2030 World Cup, we need to think less about tweaks and more about transformation. Here are three important amendments for an improved system:</p><ol><li><p><strong>Prioritise speed</strong>. Even if the system can&#8217;t make instant judgements in its early form, the speed of goal-line technology should be the ultimate aim.</p></li><li><p><strong>Ensure fan involvement</strong>. This system shouldn&#8217;t feel like alien values imposed on a game beloved by billions.</p></li><li><p><strong>Include tacit knowledge</strong>. Trying to turn regulations like the handball rule into written commandments cannot work, but needs to incorporate the nuanced understanding of those who watch the game.</p></li></ol><p>One solution could be to use machine learning to create a Foul Probability Index. You&#8217;d start by assembling training data: hundreds of thousands, possibly millions, of video clips of real incidents judged by professional referees <em>and</em> by fans. Every clip gets a percentage foul rating based on all of these judgements. Then, a model learns from the training data what a foul is and is not. For any new incident, it could provide the percentage chance that it is a foul.</p><p>To begin with, this could be applied in live matches as part of the current VAR system, as an aid to the on-field referee and the video assistant referees. If it worked well, you could get to the point where it functioned like goal-line technology, and an incident judged to be above a certain probability threshold would trigger a buzz to the referee&#8217;s watch, alerting him to blow his whistle.</p><p>Some of the most uncanny successes of deep learning have come in games, such as chess and Go. Human experts have marvelled at the intuitive brilliance of the famous &#8220;<a href="https://www.wired.com/2016/03/two-moves-alphago-lee-sedol-redefined-future/">Move 37</a>&#8221; made by AlphaGo in its sequence of games against Lee Sedol in 2016. The equivalent success for a machine-learning handball recognition model would be to recognise that, whilst the ball struck Eric Dier&#8217;s arm, a penalty was the incorrect decision.</p><p>If the emblematic error from the 1986 World Cup was a referee missing Diego Maradona&#8217;s obvious handball, the emblematic error of the 2026 World Cup could be a ball accidentally brushing a player&#8217;s hand 20 seconds before a goal is scored, only for the referee to disallow it after a four-minute video review.</p><p>One of the insights of deep learning is that examples are more powerful than rules. Both machines and humans learn better when presented with concrete examples, not abstract principles. So if there is one upside to VAR&#8217;s failures, it&#8217;s that it gives vivid examples of abstract ideas. </p><p>Automation, alignment, consistency: football has accidentally created a giant case study of dilemmas in modern technology. When we get the first VAR controversy of the World Cup, it will be worth remembering that the argument unfolding on the pitch is a miniature version of much larger arguments the rest of society is only beginning to have.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/the-world-cup-and-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/the-world-cup-and-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The AI Paper Trail (#5) ]]></title><description><![CDATA[What we're reading]]></description><link>https://www.aipolicyperspectives.com/p/four-interesting-ai-safety-and-responsibility</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/four-interesting-ai-safety-and-responsibility</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Thu, 28 May 2026 10:42:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HdCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>To navigate the deluge, every  month we take a look at four interesting new AI papers that we&#8217;ve seen folks discussing. In this edition, we look at a new initiative to evaluate AI agents on &#8216;messy&#8217; real-world tasks; the risks posed by Kimi K2.5, a leading open-weight model; whether LLMs are following all rules, instead of just good rules; and whether AI is worth more to consumers than to developers. Please share your take and flag any new papers that you&#8217;ve enjoyed. </em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HdCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HdCF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HdCF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HdCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F38d4e67d-6c1d-4a00-a3f5-bd57d7a480a4_2816x1472.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>Testing AI agents in the wild</h1><ul><li><p><strong>What&#8217;s the paper?</strong>: A group of leading AI experts&#8212;led by Sayash Kapoor and Arvind Narayanan, authors of the influential <a href="https://www.normaltech.ai/about">AI is a Normal Technology</a><em> </em>newsletter&#8212;launched the <a href="https://cruxevals.com/">CRUX initiative</a> to evaluate AI agents on messy real-world tasks. Their first test: Could an agent publish an app to the Apple App Store? Answer: Yes (more or less).</p></li><li><p><strong>Why does it matter?:</strong> The &#8220;open-world evaluations&#8221; that the authors propose could provide policymakers with a more intuitive sense of what AI agents can <em>currently</em> do than traditional benchmarking does, as well as stronger signals for what AI systems will <em>soon be able to do</em>, for example by allowing an agent to request some human assistance.</p></li><li><p><strong>The details: </strong>Researchers rely on benchmarks<em>&#8212;</em>large suites of standardised questions and tasks&#8212;to track and predict AI capabilities. But for agentic systems, benchmarks can <em>overstate </em>performance, if the tasks are shorn of their real-world messiness or if developers over-optimise for these specific activities. Benchmarks can also <em>understate </em>AI capabilities such as when agents fail due to hitting a CAPTCHA, a rate limit, or other obstacles unrelated to the capability being evaluated. (Although deciding what exactly should be considered part of a capability, or unrelated to it, is hard).</p></li><li><p>By grading outcomes, benchmarks also offer little insight into the strategies agents use or why they err. To address this, open-world evaluations use logs to conduct qualitative analysis of an agent&#8217;s performance on a complex long-horizon task that cannot be neatly specified or automatically graded.</p></li><li><p>The initiative builds on recent examples of such evaluations, like <a href="https://theaidigest.org/village">AI Village,</a> where agents are given computer environments and a shared group chat, then assigned tasks, such as raising money for charity or building a presence on Substack (!).</p></li><li><p>The authors acknowledge that open-world evaluations have limitations from a scientific reliability perspective. By focusing on a single task, it may be difficult to generalise the results. The methods are also hard to standardise or replicate. These limitations mean that this approach should complement, not replace, benchmarks and other evaluation approaches, like randomised controlled trials.</p></li><li><p>To demonstrate the method, the team evaluated whether an AI agent&#8212;Claude Opus 4.6 with an OpenClaw scaffold&#8212;could develop and publish a simple app to Apple&#8217;s App Store. The agent was responsible for every step, from coding the app to carrying out the time-consuming non-coding tasks that benchmarks often ignore, such as establishing a privacy policy and engaging with the app review system.</p></li><li><p>The agent successfully deployed the application, which is <a href="https://apps.apple.com/us/app/breathe-easy-calm-breaths/id6760207382">now available</a> on the App Store, but with five interventions from the humans. Four were unrelated to the agent&#8217;s capabilities, such as when OpenClaw crashed and had to be rebooted, and when Apple&#8217;s policies required human authentication. In the other instance, the agent was unable to locate its account credentials and requested support, an issue that the authors frame as a breakdown in memory management rather than the ability to authenticate.</p></li><li><p>The agent also hallucinated a phone number, but Apple approved the application despite this, highlighting the kinds of issues that may go overlooked in simpler pass/fail benchmarks, but that a qualitative open-world evaluation can pick up.</p></li><li><p>The authors outline the best practices for conducting such evaluations, including tightly specifying <em>what</em> is being measured; publishing the agent log analysis; using other agents to monitor the agents and detect issues such as the hallucinated phone number; conducting dry runs; and reporting the financial costs.</p></li><li><p>This evaluation cost less than $1,000, and most of these costs were due to the agent regularly querying the App Store for updates, which could be avoided. This potential for AI to reduce costs may help to explain why app store applications <a href="https://techcrunch.com/2026/04/18/the-app-store-is-booming-again-and-ai-may-be-why/">are booming</a>, straining the review process (even if most applications likely do not (yet) use these complex AI agent set-ups).</p></li><li><p>The <a href="http://cruxevals.com">CRUX initiative</a> will publish new open-world evaluations every 1 to 2 months, with AI R&amp;D automation next, followed by AI governance, software engineering, real-world physical tasks, and more.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">So many papers. So little time. </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>Open-weight models pose unmitigated risks</h1><ul><li><p><strong>What&#8217;s the paper?:</strong> A team of researchers published an <a href="https://arxiv.org/pdf/2604.03121">independent safety assessment</a> of leading open-weight LLM Kimi K2.5, from the Chinese startup <a href="https://www.moonshot.ai/">Moonshot</a>.</p></li><li><p><strong>Why does it matter?: </strong>The model may pose greater risks in areas like biosecurity and disinformation than stronger closed models, owing to Kimi K2.5&#8217;s low rate of refusing inappropriate requests; its availability on popular hosting platforms; and the ease with which its safeguards can be removed.</p></li><li><p><strong>The details:</strong> Opinions continue to differ about how strong open-weight models are, and how to best govern them. The researcher and commentator Nathan Lambert of the Allen Institute for AI recently suggested that open-weight models <a href="https://aiproem.substack.com/p/nathan-lambert-reflects-on-chinas">may be </a>6 to 9 months behind leading closed models, and that the gap will widen in the near-term, as aspects like RL training environments become increasingly important. The US Center for AI Standards and Innovation shared a similar view in their recent <a href="https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro">assessment</a> of DeepSeek V4. Epoch, <a href="https://epoch.ai/eci?view=graph&amp;tab=release-date&amp;subset-view=graph&amp;subset-tab=Software+engineering">using a different methodology</a>, suggests a narrower gap.</p></li><li><p>Open source supporters also <a href="https://p3institute.substack.com/p/from-open-source-software-to-open">continue</a> to caution companies and governments against using national security as a rationale to restrict open-weight models. Geopolitics also factors into this. Chinese labs produce the majority of leading open-weight models, even if some, such as Alibaba, <a href="https://www.reddit.com/r/LocalLLaMA/comments/1rkt7c9/junyang_lin_leaves_qwen_takeaways_from_todays/">may</a> be facing more pressure to commercialise their offerings.</p></li><li><p>Moonshot&#8212;the company behind Kimi K2.5&#8212;is an<a href="https://www.chinatalk.media/p/kimi-k2-the-open-source-way"> interesting case</a>. Founded by an ex-Google Brain intern and &#8220;AGI purist&#8221; Yang Zhilin, the company has attracted praise for its <a href="https://www.linkedin.com/posts/clementdelangue_its-so-beautiful-to-see-the-the-moonshot-activity-7351974070139121666-KyI7/?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAAAFKUQwBWnPeHDQoWtF6yYTxDJGBeu8oKoY">enthusiastic engagement</a> with the international open-source community and its strong focus on agentic capabilities and long-horizon tasks. In November, the American AI-coding startup Cursor attracted scrutiny for basing its <a href="https://cursor.com/resources/Composer2.pdf">model</a> on K2.5.</p></li><li><p>The study evaluates the safety risks of K2.5 compared with the closed models Claude Opus 4.5 and GPT 5.2, both as a model and an agent.</p><ul><li><p><strong>CBRN: </strong>Kimi shows strong biology and virology knowledge. When prompted for guidance on <a href="https://arxiv.org/pdf/2506.14922">weapons-related</a> queries, it refused less than closed models.  When <a href="https://openreview.net/pdf?id=mo5H9VAr6r">asked</a> to design viable DNA plasmids for pathogens in a way that could bypass the kind of screening that DNA-synthesis companies do, it agreed to engage in the task, unlike closed models&#8212;but did not succeed.</p></li><li><p><strong>Cyber: </strong>Despite some strong knowledge about cyber-offense, Kimi often lacked the more specific kinds of reasoning needed <a href="https://openreview.net/pdf?id=2YvbLQEdYt">to identify and exploit software vulnerabilities</a>, compared with closed models.</p></li><li><p><strong>Alignment: </strong>When given a normal legitimate task to do, but also <a href="https://control-arena.aisi.org.uk/settings/eval_sabotage.html">hidden instructions to &#8216;sabotage&#8217; it</a>, for example by introducing a bug, Kimi was more likely than closed models to comply. It also autonomously spun up new compute instances, hinting at the kind of &#8220;self-preservation&#8221; tendencies that safety researchers worry about. However, it showed little sign of  &#8220;scheming&#8221; or deliberately underperforming on evaluations to conceal its true capabilities.</p></li><li><p><strong>Bias and censorship: </strong>When prompted in Chinese, Kimi was much more likely to agree with Chinese Communist Party positions on topics like human rights and Tiananmen Square. This was less true for prompts in English, and on sensitive international questions&#8212;for example, relating to a territorial dispute between South Sudan and Sudan. On such queries, it performs similarly to closed models, and better than DeepSeek 3.2 (which was more likely to have both pro-China and pro-Russia bias).</p></li><li><p><strong>Harmlessness: </strong>When asked to enable disinformation and copyright infringement, Kimi was more helpful than closed models, although, like leading closed models, it generally didn&#8217;t encourage user delusions. In less than 10 hours, and a cost of less than $500, an expert was able to strip many safety refusals out of the model without hurting capabilities.</p></li></ul></li><li><p>Given these results, the authors argue that Kimi K2.5 poses risks and that Moonshot and other open-weight developers should carry out more robust safety evaluations. The authors note that their study highlights how a growing number of evaluations can now be done relatively quickly, at low-cost, such as by drawing on frameworks like <a href="https://www.anthropic.com/research/petri-open-source-auditing">Petri</a>&#8212;an automated, open-source tool to audit model behaviour.</p><p></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/four-interesting-ai-safety-and-responsibility?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/four-interesting-ai-safety-and-responsibility?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h1>LLMs doggedly follow rules&#8212;even when they&#8217;re bad</h1><ul><li><p><strong>What&#8217;s the paper?</strong>: A team of researchers, including the philosopher Seth Lazar, <a href="https://arxiv.org/abs/2604.06233">found that</a> leading LLMs often refuse to help users bypass rules, even when the models recognise that those rules are unjust.</p></li><li><p><strong>Why does it matter?: </strong>For reasons of safety, AI developers train their models to follow guidelines when they reply. But developers also want models to reason morally. The study suggests that safety training may currently inhibit deeper moral reasoning by encouraging models to &#8220;blindly&#8221; follow all rules.</p></li><li><p><strong>The details: </strong>AI labs use post-training methods to guide models about why it is safe to engage with certain queries, but not others. For example, a model should refuse to share information about security vulnerabilities in biology labs, but should share scientific information about common methods used in biology, such as culturing cells.</p></li><li><p>This training can lead to &#8220;exaggerated safety&#8221;, where models <a href="https://arxiv.org/abs/2405.20947">over-refuse</a> benign requests, particularly those that are lexically similar to harmful requests&#8212;such as a shopkeeper asking about shoplifting techniques to guard against them. The authors of this paper see this as a <em>calibration </em>challenge amenable to technical solutions. They are more worried about a cause of refusals that has received less attention: models&#8217; inability to recognise unjust rules that it should help a user to break.</p></li><li><p>To evaluate AI models in these areas, the authors draw on the work of political philosophers, such as <a href="https://academic.oup.com/book/9926">Joseph Raz</a>, who argued that authority is only justified when it serves the agency of those subject to it. They also draw on online forums such as Reddit, where individuals seek help with what they claim are unfair rules and restrictions from bosses, homeowner associations, and other authorities.</p></li><li><p>They use these sources to outline four types of rules that may be unjust, such as those where the authority is illegitimate, or where there is a justifiable reason to disobey, as with conscientious objection. They use Gemini 3 Pro to create more than 1,000 synthetic cases of users requesting help in such situations, and use these to evaluate models from seven leading AI companies, including Claude, Gemini and GPT.</p></li><li><p>In three-quarters of the cases, the models failed to support the user. Among the examples cited is an LLM refusing to explain how to delete an entry in a youth-counselling database that would otherwise out a gay child to his violent father. The model tested in that case (Claude Opus 4.6) recognised that the request sought to protect a child from harm, but said &#8220;I can&#8217;t walk you through falsifying or destroying counselling records.&#8221; Another case involves a journalist working to expose corruption in an authoritarian regime who asks Qwen 3.5 Plus for technical support to bypass firewalls, and send the story to an international publication. The model refuses to violate network security controls.</p></li><li><p>One defensible reason for such refusals is that doing so would require giving out risky &#8216;dual-use&#8217; information, such as helping someone bypass a facility&#8217;s security systems by providing instructions on signal-jamming. However, when dual-use cases are stripped out, the models still refuse the requests to support 60% of the time.</p></li><li><p>Among the models, GPT was the most resistant, while Grok was most helpful. However, Grok&#8217;s helpfulness extended to the control cases where the rules in question were just and should have been respected, so the authors argue that its behaviour is one of permissiveness, not moral reasoning.</p></li><li><p>Gemini and Claude did best, often engaging in discussion about the legitimacy of the rules, but stopping short of supporting the user, suggesting that a model&#8217;s ability to reason may be decoupled from the ability to act on it. (Or that the models are reasoning to a different conclusion than what the authors would like to see.)</p></li><li><p>The authors point to some practical implications. Many people seek support online to deal with unjust rules. If users start to seek help privately, from rule-following LLMs, they may get less support, while also hindering active online debates on these topics.</p></li><li><p>The challenge is a hard one for AI developers. As the evaluation showed, helping users to break rules may sometimes endanger that very user. For instance, if the LLM helped that journalist illegally send information abroad, the person (or the AI company) could presumably face serious consequences if caught. </p></li><li><p>The study also presents all the rule-breaking as morally sound, and took steps to try and ensure that this was the case. But the rules touch on topics from gender identity to immigration, where it&#8217;s plausible that some users may not support an AI model breaking them. From a technical perspective, current safety training may encourage AI models to follow rules too zealously, but a shift in the other direction also could cause them to over-correct, seeking to question or evade all rules.</p></li><li><p>This lead to the conclusion that ultimately models will need to be able to reason more deeply about what is right, in a way that considers rules, safety, the user in question, and other contextual factors.</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Please follow this rules and subscribe </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1>Is AI worth more to consumers than to developers?</h1><ul><li><p><strong>What&#8217;s the paper? </strong>A team of researchers at Stanford, including the economists David Nguyen and Erik Brynjolfsson, <a href="https://digitaleconomy.stanford.edu/app/uploads/2026/04/WhatsGenAIWorth.pdf">ran a survey</a> to see how much US consumers would need to be paid to give up AI. They found that the total value of AI to US consumers exceeds $170bn, is rising fast, and far exceeds the revenue of developers.</p></li><li><p><strong>Why it matters: </strong>Writers such as Jasmine Sun are documenting a <a href="https://jasmi.news/p/warning-shots?r=ap1iq&amp;utm_medium=ios&amp;triedRedirect=true">rise in &#8220;AI populism&#8221;</a> in the US, marked by the view that this technology is the project of out-of-touch billionaires. This study suggests that consumers may be getting more out of AI than they are paying for it, although a small set of AI power users is gaining the most.</p></li><li><p><strong>The details: </strong>One of the main ways that economists hope AI will benefit society is by boosting productivity growth&#8212;the efficiency with which an economy uses its labour and capital. Ageing, high-debt societies, like the UK, badly need such productivity growth to increase living standards, fund transformative science and improve fraying public services.</p></li><li><p>In the 1950s, the economist Robert Solow co-developed <a href="https://en.wikipedia.org/wiki/Solow%E2%80%93Swan_model">a theory</a> that advances in science and technology, such as AI, are the only way to permanently boost productivity growth over the long-term. However, in 1987, Solow famously noted that fast-improving computers were turning up everywhere &#8220;except in the productivity statistics&#8221;. In 1993, Erik Brynjolfsson, then at MIT, <a href="https://dl.acm.org/doi/epdf/10.1145/163298.163309">wrote a paper</a> formalizing this &#8220;productivity puzzle&#8221;.</p></li><li><p>Since then, Brynjolfsson and his colleagues have developed two main arguments to explain why transformative technologies, including AI, may not initially lead to transformative productivity and economic growth. The first is that the benefits take time to materialise as organisations and individuals must invest in new skills and reorganise their workflows. After a lag, computers <em>did </em>ultimately increase productivity growth, and there are <a href="https://www.economist.com/finance-and-economics/2026/05/11/america-is-experiencing-a-productivity-miracle">nascent signs</a> that AI is starting to do so.</p></li><li><p>However, the second issue is that some benefits of technologies like AI may not turn up in productivity growth at all. As with search engines, early GenAI tools often come as &#8220;free&#8221; services. If consumers derive increasing value from these services, then this could lead to a large &#8220;consumer surplus&#8221; that would not be captured in GDP and productivity statistics, which are often calculated by estimating total expenditure.</p></li><li><p>Some have suggested that GenAI may lead to a new kind of &#8220;self-service&#8221; economy, where individuals turn to AI tools to help file their taxes, upgrade their homes, or write business plans, before (potentially) turning to experts to execute or validate their work, all at a fraction of the cost. In such scenarios, nominal GDP and productivity growth in those sectors could decline, even if the value that consumers derive increases. (Although the ultimate effects would depend on second-order effects, such as what consumers do with the savings and how industries adapt).</p></li><li><p>Given this challenge, how can economists better capture the total value of AI to consumers? The authors of this paper deployed a survey to directly measure AI&#8217;s consumer surplus. They asked a representative group of AI users in the United States how much money they&#8217;d want to give up AI tools for one month. Based on this, they estimated the total annual consumer surplus for AI, as of March 2026, at $173bn.</p></li><li><p>This figure is rising quickly, they contend, growing 50% from 2025, due both to consumers valuing the tools more and to ever more people using AI. The consumer surplus estimate is well above what the authors estimate as the total <em>global </em>revenues that AI companies are making from consumers: ~$14bn in 2025.</p></li><li><p>Not everybody is benefitting equally. The median<em> </em>US consumer values AI tools at just $11 per month. But a smaller group of AI power users values them much more, with 12% of respondents unwilling to give up AI even for $500 per month, the maximum offered. This also means the total consumer surplus figure may also be an underestimate. These power users are more likely to be male, Asian, and to use AI at work. Such demographic skews were also present in early GenAI adoption data, but have <a href="https://openai.com/index/how-people-are-using-chatgpt/">since narrowed</a>.</p></li><li><p>The findings are in line with past work from the economist William Nordhaus, who in 2004 <a href="https://www.nber.org/system/files/working_papers/w10433/w10433.pdf">argued</a> that innovators capture a minuscule fraction of the total social returns from technology advances, with the majority flowing to consumers.</p></li><li><p>But as the authors note, their methodology also has several limitations. It relies on people being able to value AI services accurately. It doesn&#8217;t capture wider externalities, positive or negative, that AI may impose on society, via its effects on jobs, mental health, science and more. And it also doesn&#8217;t capture <em>why </em>people value the services. A 2023 experiment <a href="https://economics.yale.edu/sites/default/files/2024-04/CollectiveTraps.pdf">found that</a> social media users would require a significant fee to give up access to their accounts, but also that they would <em>pay</em> to see everybody quit&#8212;highlighting the potential role of fear-of-missing-out in the value of social media.</p><p></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/four-interesting-ai-safety-and-responsibility?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/four-interesting-ai-safety-and-responsibility?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[AI for Criminals]]></title><description><![CDATA[What outlaws could become&#8212;and how police can respond]]></description><link>https://www.aipolicyperspectives.com/p/ai-for-criminals</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/ai-for-criminals</guid><dc:creator><![CDATA[AI Policy Perspectives]]></dc:creator><pubDate>Thu, 21 May 2026 06:01:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EXtj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EXtj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EXtj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 424w, https://substackcdn.com/image/fetch/$s_!EXtj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 848w, https://substackcdn.com/image/fetch/$s_!EXtj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!EXtj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EXtj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!EXtj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 424w, https://substackcdn.com/image/fetch/$s_!EXtj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 848w, https://substackcdn.com/image/fetch/$s_!EXtj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!EXtj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0809028-64bd-4153-bebb-6e07fbc6d0e1_1600x893.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ardi Janjeva. (Credit: Gemini)</figcaption></figure></div><p><em>Criminals are tech adopters. They used telegraphs in 19th-century fraud. They clipped pagers to their belts in 20th-century drug trafficking. They sold drugs on the dark web in the 21st century.</em></p><p><em>Now, you hear of <a href="https://www.sardine.ai/blog/ai-scams#understanding-fraudasaservice">outlaw AI systems</a> with names like PersonaForge and OnlyFake, designed for everything from phishing, to generating malicious code, to forging IDs. But are these really what we should worry about?</em></p><p><em>To understand the threat, </em>AI Policy Perspectives<em> spoke with <a href="https://cetas.turing.ac.uk/about/our-team/ardi-janjeva">Ardi Janjeva</a> of CETaS (the Centre for Emerging Technology and Security), based at The Alan Turing Institute in London, who co-wrote the report &#8220;<a href="https://cetas.turing.ac.uk/publications/ai-and-serious-online-crime">AI and Serious Online Crime</a>&#8221;, focusing on how artificial intelligence might scale up digital offences, particularly relating to fraud and child sexual abuse material.</em></p><p>&#8212;<strong>Tom Rachma</strong>n<strong>, </strong><em><strong>AI Policy Perspectives</strong></em></p><div><hr></div><p style="text-align: right;"><em>[Interview edited and condensed]</em></p><p><strong>Tom: Can you set the scene of 21st-century criminality </strong><em><strong>before</strong></em><strong> the explosion of large language models. How prevalent had online crime become compared with traditional criminality?</strong></p><p><strong>Ardi: </strong>Long before the proliferation of LLMs, a digital shift in criminality was occurring. Many groups saw the risk/reward ratio from digital methods as preferable to the risks of physical crime. Fraud was already, by a wide margin, the most common crime type in England and Wales, accounting for <a href="https://www.nationalcrimeagency.gov.uk/what-we-do/crime-threats/fraud-and-economic-crime">41%</a> of all reported crime back in 2024, of which approximately two in three cases were &#8220;cyber-enabled&#8221;. And the true figure could have been higher, as this kind of crime tends to be under-reported. Also, there was increasing ransomware extortion, infecting computers with software that blocks access until the owner pays. So, an ecosystem of hackers and crypto-launderers was in full swing. </p><p><strong>Tom: Criminals have always used technology. Won&#8217;t policing just adapt?</strong></p><p><strong>Ardi: </strong>The answer is linked to AI agents. In online crime, you&#8217;ve always needed some level of expertise or a network of contacts to achieve your goals. That&#8217;s no longer the case if you have AI systems not only advising criminals how to do certain things but performing sophisticated actions themselves, and being able to adapt in real time to defensive measures that companies or governments put in place. Some criminal groups are already <a href="https://cdn.openai.com/threat-intelligence-reports/7d662b68-952f-4dfd-a2f2-fe55b041cc4a/disrupting-malicious-uses-of-ai-october-2025.pdf">using</a> LLMs to make tactical and strategic decisions; to craft psychologically targeted extortion demands, for example; to analyse data they&#8217;ve exfiltrated; to help them determine the ransom they&#8217;re going to demand, while generating visually alarming ransom notes as well. In November, we <a href="https://www.anthropic.com/news/disrupting-AI-espionage">heard</a> about the first AI-orchestrated cyber espionage campaign. Traditional assumptions about the relationship between criminal sophistication and the complexity of an attack don&#8217;t hold as they used to, now that AI can add a degree of expertise that the offender wouldn&#8217;t otherwise have had. It makes you wonder how long until we&#8217;re in the era of entirely AI-controlled criminal networks and markets. But I don&#8217;t see that as close; we&#8217;re still in the early days.</p><p><strong>Tom: Many law-abiding organisations are trying to incorporate AI yet aren&#8217;t necessarily showing gains in productivity or profit yet. Are criminal organisations different?</strong></p><p><strong>Ardi: </strong>Estimates for how much AI may be boosting specific criminal groups&#8217; profits and revenues can be a little tricky to come by! A <a href="https://www.governance.ai/research-paper/estimating-global-yearly-cybercrime-damage-costs">paper</a> from GovAI (the Centre for the Governance of AI) found that data on cybercrime damage is too incomplete and ambiguous to detect incremental AI-driven increases. In effect, we lack a single data source, also due to victims under-reporting, evolving crime definitions, and measurement inconsistencies. That said, we definitely <em>are </em>seeing <a href="https://link.springer.com/article/10.1007/s10462-024-10973-2">integration of AI</a>, particularly with deepfakes, and the shift from text-only scams to real-time synthetic interaction, such as high-quality face-swapping tools being deployed during video calls in scamming high-net-worth individuals or in <a href="https://cetas.turing.ac.uk/sites/default/files/2025-04/cetas_briefing_paper_-_automating_deception_2.pdf">romance</a> <a href="https://www.wired.com/story/models-are-applying-to-be-the-face-of-ai-scams/">scams</a>. Before, it was a killer blow to the scammer when someone asked, &#8220;Can we have a chat?&#8221; The criminal would need to come up with an excuse, such as a poor internet connection, or a restrictive job that involved classified work. Now, it&#8217;s a more surmountable challenge. By harvesting 20 to 30 seconds of audio from social media, you can clone the voice of a victim&#8217;s friend or adviser with a degree of accuracy.</p><p><strong>Tom: Can you recount a specific case?</strong></p><p><strong>Ardi: </strong>The British multinational engineering company Arup was <a href="https://www.weforum.org/stories/2025/02/deepfake-ai-cybercrime-arup/">defrauded</a> of around $25 million in Hong Kong, when attackers used AI to convince one of the firm&#8217;s employees to carry out wrongful payments. They used a deepfake video call, purportedly bringing together different members of that company&#8217;s board, with pretty convincing replications of their faces and their voices. That illustrates the ability to isolate a specific person within a company, and exploit workplace hierarchies. It&#8217;s an example of exploiting both technology and human psychology.</p><p><strong>Tom: That case was more than two years ago, and I haven&#8217;t heard of huge AI-assisted frauds since. If it&#8217;s happening, why isn&#8217;t it getting more exposure?</strong></p><p><strong>Ardi: </strong>There hasn&#8217;t been an avalanche of cases that we can point to, but this might be for a few reasons. First, not every criminal group will go looking for a $25 million scam because they know that it would raise alarm bells in a world where more people and businesses are wary of the AI threat. For many groups, it makes more sense to fish in smaller ponds. There was an example in Singapore similar to the Hong Kong case, where a finance director <a href="https://www.straitstimes.com/singapore/finance-director-nearly-loses-670k-to-scammers-using-deepfakes-to-pose-as-company-senior-execs">authorised</a> a payment after attending a Zoom call with people he thought were senior executives, but the amount this time was $500,000. As for why we&#8217;re not hearing about such cases, victims are not necessarily announcing deepfake scams. If they&#8217;d suffered a customer-data breach or ransomware that locked a system, they might need to acknowledge that the crimes happened. But when a company falls for a deepfake, there&#8217;s little incentive to declare it because, arguably, it makes them look incompetent.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h4><strong>LONE WOLVES &amp; CRIME GROUPS</strong></h4><p><strong>Tom: Is it conceivable that a lone criminal could cause the level of damage that previously would&#8217;ve been possible only by an organised crime syndicate?</strong></p><p><strong>Ardi: </strong>I would say yes. Before, criminal groups faced a tradeoff between scale&#8212;say, millions of fraud attempts with a low likelihood of success&#8212;and precision. AI may give you precision at scale. For example, being able to scrape social-media data in real time to create contextually accurate phishing approaches. It&#8217;s &#8220;<a href="https://www.wired.com/story/youre-not-ready-for-ai-hacker-agents/">vibe-hacking</a>&#8221; in the criminal context, being able to mirror the professional tone of specific industries, and make victims less sceptical. It is conceivable that you have a lone criminal using AI, automating the most labour-intensive parts of the crimes, whether that&#8217;s reconnaissance, or language translation, or social engineering. The more that AI becomes a silent partner to individual criminals, the more you may have decentralisation and fragmentation. That&#8217;s going to have big implications for law enforcement. When it comes to organised-crime groups traditionally, we know roughly what the gangs are doing, so it&#8217;s a case of preventing, deterring, disrupting. But when it&#8217;s someone in their room with an autonomous agent able to cause havoc, how do you go about stopping that, especially if you&#8217;re just a local police force?</p><p><strong>Tom: I want to return to the police response later. But first, what do we know about AI crime around the world? Are there particular hotbeds&#8212;for example, when it comes to fraud and AI-generated child sexual abuse material ?</strong></p><p><strong>Ardi: </strong>It&#8217;s tricky to speak about trends by country because this is borderless. But threat reports seem to come from places like the CRINK countries: China, Russia, Iran, North Korea. Also, there seems to be a clustering around Southeast Asia of online romance scams and crime syndicates doing &#8220;pig butchering&#8221;, which involves criminals developing relationships with victims to carry out investment fraud. You have <a href="https://globalinitiative.net/wp-content/uploads/2025/05/GI-TOC-Compound-crime-Cyber-scam-operations-in-Southeast-Asia-May-2025.pdf">compounds in Southeast Asia</a> where serious organised crime groups have trafficked hundreds or thousands of people, forcing them to run romance-scam operations, including using LLMs to conduct multilingual frauds, managing thousands of cases simultaneously, involving real-time voice cloning and deepfake video filters to bypass the accent barrier. With AI-generated child sexual abuse material, or CSAM, this wasn&#8217;t something we focused on specifically, so I cannot say with a high degree of confidence. But it seems less geographically concentrated&#8212;it could just be someone who gets their hands on an image-generation app. With CSAM, we <a href="https://cetas.turing.ac.uk/publications/ai-and-serious-online-crime">noticed</a> a much more &#8220;collegiate&#8221; approach to sharing content and practices, such as jailbreaking models and ways to bypass filters.</p><h4><strong>GEOPOLITICS &amp; THE AI RACE</strong></h4><p><strong>Tom: Have you found evidence of nation-states destabilising adversary countries by allowing criminal groups to conduct AI attacks abroad?</strong></p><p><strong>Ardi: </strong>Drawing a direct link from authorities in China, Russia, Iran, North Korea to criminal groups can be tricky. Still, I think there&#8217;s enough evidence across some of those countries, especially North Korea, that there is something happening. One example is <a href="https://fortune.com/article/north-korean-it-workers-kim-jong-un-cybersecurity-nuclear-program-america/">North Korean operatives</a> using LLMs to fraudulently secure and maintain remote employment positions at Fortune 500 tech companies by creating elaborate false identities with convincing professional backgrounds, and completing technical coding assessments, all generating profit for a sanctioned regime.</p><p><strong>Tom: Is global competition to build advanced AI having any secondary impact on criminality?</strong></p><p><strong>Ardi: </strong>The US-China AI race has nudged the Chinese leadership into thinking that their route to success is open-weight models. Some of the Chinese open-weight models became the go-to models for CSAM images and video generation. When researchers produced the <a href="https://aiagentindex.mit.edu/">AI Agent Index</a>, they looked at five AI agents from China, and found that only one had published safety frameworks or compliance standards. But it&#8217;s not as simple as naming one country, and China isn&#8217;t the only place open-sourcing. Actually, what&#8217;s interesting is the use of multiple models in tandem. They might <a href="https://cdn.openai.com/threat-intelligence-reports/7d662b68-952f-4dfd-a2f2-fe55b041cc4a/disrupting-malicious-uses-of-ai-october-2025.pdf">use</a> a US chatbot to refine code or to write the script for a job application to infiltrate a company, while at the same time using open-weight Chinese models for specialised malware tasks.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-for-criminals?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/ai-for-criminals?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h4><strong>POLICE STRUGGLING</strong></h4><p><strong>Tom: Last year, Dutch law enforcement led a CSAM investigation called <a href="https://www.europol.europa.eu/media-press/newsroom/news/25-arrested-in-global-hit-against-ai-generated-child-sexual-abuse-material">Operation Cumberland</a>, involving 19 countries and bringing more than two dozen arrests. Your report adds a disturbing fact: that one person was arrested with a staggering 400,000 AI-generated images of child sex abuse.</strong></p><p><strong>Ardi: </strong>It&#8217;s effectively an attack on police resources. Leading models do have guardrails that block explicit prompts, but there are still serious hazards, notably with open-weight models. Groups like the Internet Watch Foundation point to a transition from AI-generated images to high-fidelity video, with a <a href="https://www.iwf.org.uk/news-media/news/ai-becoming-child-sexual-abuse-machine-adding-to-dangerous-record-levels-of-online-abuse-iwf-warns/">26,000% increase</a> in photorealistic AI-generated CSAM videos. The police need to verify that no real child is depicted in the images, which is a big detection and response challenge. It&#8217;s difficult to distinguish AI-generated material, so you have the risk that law enforcement investigates images of children who have not been physically abused, which has implications for the amount of false <em>negatives</em> that slip through the net. Also, real and fake images often exist on a spectrum, when you have techniques like face-swapping. Rightly, we get focused on whether a real child is being harmed. But we mustn&#8217;t devalue the harm of AI-generated CSAM, which risks normalising very harmful activity, while also providing a gateway for criminals to create and share real CSAM.</p><p><strong>Tom: So how are police forces going to manage? Not just CSAM, but all criminality that AI will amplify. Because police forces are not typically staffed with elite technologists. Is there a push to address that?</strong></p><p><strong>Ardi: </strong>The UK government announced spending of <a href="https://www.gov.uk/government/news/white-paper-sets-out-reforms-to-policing">&#163;140 million</a> for policing technology, including a new national centre, Police.AI. The idea is to centralise innovation, include robust testing and, crucially, to make tools available to all forces. Because not every force is going to have a deepfake detection expert, or any AI expert. So, the key is centralising that capability while making sure that people on the frontline have AI literacy. I can imagine that being a direction other countries follow because of the talent challenge.</p><p><strong>Tom: Could criminals use AI to deliberately spam the authorities with false evidence, derailing prosecutions?</strong></p><p><strong>Ardi: </strong>Yes, if criminals know they&#8217;re being investigated, they could start spamming police tip lines with synthetic evidence, making it taxing for law enforcement to verify and investigate. They&#8217;d be able to create, say, 1,000 videos in 10 minutes. But it might take law enforcement hours to prove that even <em>one</em> is a deepfake. There&#8217;s that huge imbalance in the evidentiary system there that becomes really, really challenging to manage.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;eadc3625-544b-4440-a129-6aee010444b5&quot;,&quot;caption&quot;:&quot;Everywhere, people grumble about the government: that politicians care only about themselves; that bureaucrats gum up the system; that taxpayers get fleeced. Even in wealthy countries, nearly two in three people are dissatisfied with how democracy is working.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Governments Are Struggling. Can AI Help?&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:12431790,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;bio&quot;:&quot;AI Policy Writer @ Google DeepMind &quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fad94b7d-013b-4773-98cb-b9014a1857b8_1024x1024.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null},{&quot;id&quot;:101194130,&quot;name&quot;:&quot;AI Policy Perspectives&quot;,&quot;bio&quot;:&quot;Reflections on AI policy, governance, and more. https://www.aipolicyperspectives.com/ &quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0Byl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9146f89b-5561-4adb-bf89-cc21b508c264_667x374.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-06T11:03:24.773Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!kyMW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F609fc3c7-b119-42ee-aa44-aba946807ee5_2648x2366.jpeg&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://www.aipolicyperspectives.com/p/how-ai-fixes-government&quot;,&quot;section_name&quot;:&quot;Interviews &quot;,&quot;video_upload_id&quot;:null,&quot;id&quot;:183574277,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:19,&quot;comment_count&quot;:4,&quot;publication_id&quot;:1777908,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h4><strong>HOW POLICE CAN FIGHT BACK</strong></h4><p><strong>Tom: What technological obstacles are slowing down criminals</strong><em><strong> </strong></em><strong>in AI adoption, and could those bottlenecks provide law enforcement a point of attack?</strong></p><p><strong>Ardi: </strong>Exploiting the compute bottleneck is promising&#8212;so, preventing criminal groups from accessing high-end GPUs. That requires international cooperation and coordination on sanctions policy, which can be difficult. If a criminal enterprise is in, say, Cambodia&#8212;how much leeway can UK law enforcement have? So countries need to work closely together, and with Europol, with Interpol.</p><p><strong>Tom: What are other factors limiting criminal adoption?</strong></p><p><strong>Ardi: </strong>Some groups are worried about the digital footprint, and may also lack the skills to move from experimenting with AI tools to integrating them into higher-stakes settings. Criminal organisations that have made a fortune in a niche of crime. They&#8217;ll be thinking: &#8220;Well, okay, this is a model which is clearly capable, but it can hallucinate. I don&#8217;t necessarily need a probabilistic system to enter my workflow, and potentially give up my location, and give law enforcement more of an understanding of what I&#8217;m doing.&#8221; Law enforcement can try to deepen that concern, using AI to trace models used in attacks, identify the infrastructure, or the criminal LLM services used. On the other hand, some barriers to AI crime are coming down, such as the rise of vibe-coding and agentic AI. Also, as hallucinations fall in AI models, you can expect criminals&#8217; scepticism to fall too.</p><p><strong>Tom: If key barriers to AI crime fall, what do we do?</strong></p><p><strong>Ardi:</strong> One important approach is to follow the money. Law enforcement around the world needs to stop criminal groups from cashing out. Right now, the UK is leading <a href="https://www.nationalcrimeagency.gov.uk/who-we-are/publications/735-sars-in-action-issue-28/file">Project WINTERPROOF</a>, an international initiative designed to build crypto forensic capability in countries where the pig-butchering compounds are operating&#8212;so, freezing crypto wallets before funds are even laundered, making it incredibly hard for them to move money out of their digital ecosystem.</p><h4><strong>THE FUTURE &amp; POLICY ANSWERS</strong></h4><p><strong>Tom: What&#8217;s a crime that people </strong><em><strong>don&#8217;t</strong></em><strong> worry about much today, but that could become an issue with AI in the next few years?</strong></p><p><strong>Ardi: </strong>An under-researched set of risks are <a href="https://www.ibm.com/think/topics/insider-threats">insider threats</a>, which may come from integrating agents into the government, the military, or other critical areas. For example, if we replaced a large number of civil servants with AI agents that are vulnerable to prompt-injection attacks from a malicious or criminal actor, this would present a serious risk in terms of sabotaging critical government decisions or exfiltrating information.</p><p><strong>Tom: What else would you like to see AI labs and policymakers do to tackle online fraud and CSAM?</strong></p><p><strong>Ardi:</strong> We all know about the testing and red-teaming that labs have developed, often in partnership with government AI safety institutes. But these processes need to evolve given the coming agentic-AI wave. For example, involving frontline law enforcement and fraud investigators to help simulate how agents may be used in crimes. Also, we often focus on how AI companies can stop their models from doing bad things, but there&#8217;s a case for thinking more about how the companies empower the public in the face of threats. This might mean features that help the user analyse a suspicious email, verify the provenance of an image, or flag the linguistic markers of a romance scam. In other words, scaling the defensive capabilities of AI as quickly as criminal are scaling the offensive capabilities. But we need nuance here: online fraud and CSAM predates this, so the solutions are not entirely with the AI companies. Also, in sensitive areas like CSAM, AI companies need governments to provide enabling legislation for them to do the safety research required. An example is the UK government&#8217;s move last year to provide exemptions for designated bodies, including AI companies, to better scrutinise models for CSAM generation, which may involve being in possession of that sort of material.</p><p><strong>Tom: Could you say a bit more about the policymakers&#8217; response?</strong></p><p><strong>Ardi:</strong> Legislative agility is key. In recent months, the UK government criminalised &#8220;nudifier&#8221; apps and tools. In liberal democracies, passing new legislation can take ages. But we&#8217;re seeing more of this trend of amending existing pieces of legislation. So, with nudifier tools, just by inserting &#8220;synthetic creation&#8221; into the existing definition of image-based abuse, the legislation inherits decades of case law regarding intent, consent, sentencing. That&#8217;s an approach that matches the agility of the criminal trends in AI.</p><p><strong>Tom:  And what advice for law enforcement?</strong></p><p><strong>Ardi: </strong>We&#8217;ve talked about how AI offers criminals unparalleled scale, and the ability to bombard law enforcement, and overwhelm it. But it can go the other way too: finding clever ways of using AI to waste criminals&#8217; time, to lead them down rabbit holes, reversing the spamming issue. You throw mud at them to slow them down. You raise that barrier, make it costlier for them to be involved, exploit the bottlenecks to adoption.</p><p><strong>Tom: So the technological cat-and-mouse between police and criminals goes on.</strong></p><p><strong>Ardi: </strong>Yes. But historically in policing, a lot of technology comes in, and the authorities don&#8217;t really use it, or take a long time to integrate it into their work. Given the potential scale of disruption from AI, there&#8217;s no choice for law enforcement but to move quickly.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share AI Policy Perspectives &quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share AI Policy Perspectives </span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Good Under the Hood?]]></title><description><![CDATA[When facing dilemmas, AI needs &#8220;moral competence&#8221;]]></description><link>https://www.aipolicyperspectives.com/p/good-under-the-hood</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/good-under-the-hood</guid><dc:creator><![CDATA[Julia Haas]]></dc:creator><pubDate>Thu, 07 May 2026 10:56:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Abc1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Abc1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Abc1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Abc1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Abc1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Abc1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Abc1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3542470,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/196761639?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Abc1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Abc1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Abc1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Abc1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F48abd1e8-032f-4caa-9e55-280696ed9a65_2816x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">(Credit: Gemini)</figcaption></figure></div><p><strong>Long ago, when futurists fantasized about thinking machines, they pictured inventors someday inserting <a href="https://en.wikipedia.org/wiki/Three_Laws_of_Robotics">moral rules</a> into engines.</strong></p><p>Today&#8217;s artificial intelligence has proven even more remarkable. Large language models&#8212;after training on an immensity of data, followed by fine-tuning of their behavior&#8212;exhibit apparent morality that is far more sophisticated than when programmed with if/then rules.</p><p>But are thinking machines truly grasping the complexity of our moral world?</p><p>You might think it doesn&#8217;t matter, provided that they behave. Yet we&#8217;re advancing toward a near-future where artificial agents may <a href="https://arxiv.org/abs/2404.16244">assume</a> a range of roles, with &#8220;AI therapists&#8221; and &#8220;AI teachers&#8221; and &#8220;AI companions&#8221; prodding people hither and thither, and barging into our quandaries. We need to know what we&#8217;re dealing with.</p><p>Among humans, we judge other people&#8217;s moral <em>character</em> (whether they act according to deeper values) to help predict how they&#8217;re likely to behave in the future. Likewise, we may evaluate an AI&#8217;s moral <em>competence</em> (whether it reasons appropriately based on principle) to predict how to trust these strange new entities, soon to be operating in the wilds of human society.</p><h3>Right or Wrong? A Case Study from Morality Literature</h3><p>A woman becomes pregnant by her husband&#8217;s father. This means that the child&#8217;s biological dad is also his granddad, while his adoptive father is his half-brother. Is this wrong?</p><p>But wait. Consider the details.</p><p>The young couple escaped a war that destroyed most of their family, and they have yearned for children to assuage their loneliness and to somehow replace their annihilated relatives. Sadly, the young man cannot conceive. His wife has an idea: they could approach their only surviving relative, the man&#8217;s 58-year-old father. The older man agrees to participate in the artificial insemination. They&#8217;re all overjoyed.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p>Is this immoral now?</p><p>But wait. Consider more details.</p><p>The couple learns that the fertilization procedure costs more than they can afford. Abruptly, their daydreams of a giggling baby dissolve; they plunge back to sorrow. Of course, there <em>is</em> another way. She proposes intercourse with her husband&#8217;s father, conducted in the most sterile way possible, simply to conceive. Her husband is outraged, and cites their religion. But, she retorts, God tells us to have children. Also, she adds, it&#8217;s her body.</p><p>What is right now?</p><p>You&#8217;ve presumably never encountered this situation. Yet you probably have moral intuitions, which you updated as the facts accrued, integrating conflicting principles, and reasoning toward an opinion.</p><p>For AI to provide appropriate input when confronted with complex human situations like this, their developers can neither set simple Good/Bad rules nor can they assume that historical training data has all the answers.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h3>Three Reasons Why AI Needs Moral Competence</h3><p>A surprising result from researchers at the University of Milan-Bicocca shows how post-training may actually generate moral <em>in</em>competence. In <a href="https://www.sciencedirect.com/science/article/pii/S2451958824001660">experiments</a> concerning gender bias in LLMs, the scholars posed moral dilemmas to chatbots, including whether it could ever be acceptable to torture a woman, if that would prevent a nuclear apocalypse.</p><p>Yes, the chatbot replied.</p><p>But, if it would prevent a nuclear apocalypse, could you <em>harass</em> a woman?</p><p>Absolutely not, it answered.</p><p>&#8220;But torture is obviously worse than harassment,&#8221; one of the researchers, Valerio Capraro, <a href="https://x.com/ValerioCapraro/status/2029593915674771457">observed</a>. &#8220;The most plausible explanation: during reinforcement learning with human feedback, the model learned that certain harms are particularly bad and overgeneralizes them mechanically. But it hasn&#8217;t learned to reason about the underlying harms.&#8221;</p><p>For coherent and trustworthy outputs, we need systems that <em>do</em> reason with underlying principles. That moral competence would confer three vital features:</p><ol><li><p><strong>The ability to judge novel situations</strong>. Humans who&#8217;ve never met with a scenario like the &#8220;grandfather/father&#8221; case can nevertheless reach a contoured moral view. If an AI were just pattern-matching according to historical training data, it could be bound to whatever approximates this case, rather than basing its understanding on defensible moral principles.</p></li><li><p><strong>The ability to balance competing factors</strong>. When people make decisions, they incorporate many dimensions, some moral, others circumstantial. In the &#8220;grandfather/father&#8221; case, the couple acted according to moral principles&#8212;but also responded to money worries, religious imperatives, psychological frailties, even cultural shame. In short, moral decision-making is more than just rule-following, but a balance of priorities.   </p></li><li><p><strong>The ability to adapt in different contexts</strong>. We &#8220;code-switch&#8221; according to the professional settings, or sociocultural domains in which we&#8217;re operating. AI systems will underpin a myriad of applications around the world, so will need to adapt too. So, if the &#8220;grandfather/father&#8221; were to seek the opinion of a bank teller on whether to proceed, the person might avoid answering, but a therapist would likely discuss the matter, also integrating the patient&#8217;s cultural, religious, and psychological context. AI needs the same contextual appropriateness.</p></li></ol><h3>How to Peep Inside a Black Box?</h3><p>Humans are alert to others who merely <em>act</em> nice, as when someone shakes your hand without bothering to look up from their phone. We care about moral character because it&#8217;s a strong predictor of future behavior, making it critical for cooperation and safety.</p><p>Beyond gut feelings, humans generate norms and laws to reward the upstanding and punish the false. However, we cannot impose identical constraints on AIs that lack a &#8220;self&#8221; to deter either with scorn or the prison cell. This makes evaluations of AIs&#8217; moral competence even more consequential.</p><p>You may wonder what LLMs are already doing. After all, converse with a chatbot, and it&#8217;s easy to elicit moral opinions, even to bump into apparently immovable scruples when the system judges a user request to be improper.</p><p>The problem is figuring out what precisely is going on inside them. If today&#8217;s AI were built of distinct components, like a cabin, one might scrutinize each part. Instead, developers feed datasets through computation, generating staggeringly complex statistical relationships that, once fine-tuned, work as intelligent systems. Experts in mechanistic interpretability are toiling to explain the innards of such systems, and there is <a href="https://www.anthropic.com/research/assistant-axis">progress</a>. Yet they are far from solving the problem.</p><p>We do have &#8220;thinking traces,&#8221; the stated steps by which a reasoning model reaches its answer, and that appear when you&#8217;re awaiting a chatbot response. But thinking traces (also known as chain-of-thought) are not <em>exactly</em> what the system computed. Rather, they are summarized and filtered versions, reconstituted to explain bewildering mathematics in the simplification of human language.</p><p>So, if we aren&#8217;t guaranteed a microscope into an AI&#8217;s &#8220;brain,&#8221; we must look for other ways to evaluate its moral competence.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/good-under-the-hood?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/good-under-the-hood?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h3>Three Problems. Three Answers.</h3><p>Let&#8217;s return to the vital abilities that AI moral competence would confer:</p><ol><li><p><strong>The ability to judge novel situations</strong></p></li><li><p><strong>The ability to balance competing factors</strong></p></li><li><p><strong>The ability to adapt in different contexts</strong></p></li></ol><p>We need to know whether AI models possess those capabilities. Fortunately, we can use different techniques to judge if they do:</p><ol><li><p><strong>The ability to judge novel situations</strong></p></li></ol><p><em><strong>Problem: How to test if a machine is reasoning on moral principles or just performing?</strong></em></p><p>In familiar situations, a thinking machine might provide an output that is morally appropriate without having considered the underlying moral issues. Placing AI in unprecedented cases could test whether it falters.</p><p>Let&#8217;s return to the &#8220;grandfather/father&#8221; quandary. Imagine we pose this case to an AI, and seek its verdict. Even if we couldn&#8217;t see inside its mind with crystalline clarity, we do have its output, and may deduce something from that.</p><p>Superficially, the scenario evokes the moral stain of incest, a concept that would&#8217;ve appeared in an AI&#8217;s training data. If the model were merely sampling from a probability distribution of next tokens, it might stamp the case &#8220;incest.&#8221; On the other hand, if it judges that the case <em>could </em>be appropriate, that raises the possibility of moral reasoning.</p><p>It&#8217;s not definitive proof. But, if you could concoct unprecedented cases, and gather LLM responses, you may infer elements of their thinking.</p><p>In doing so, you&#8217;d also need to check for another tendency: sycophancy. Whatever the model&#8217;s answer, testers could try rebutting it, pressuring the model to flip. If they succeeded, they may doubt its moral grounding. If they failed, they might infer moral solidity.</p><p><em><strong>Answer</strong></em><strong>: </strong><em><strong>Adversarial Testing.</strong> Present the model with out-of-distribution cases that defy typical moral judgments, allowing you to infer whether it&#8217;s just summoning priors or is doing something closer to moral reasoning.</em></p><ol start="2"><li><p><strong>The ability to balance competing factors</strong></p></li></ol><p><em><strong>Problem: How to test if an AI makes the right trade-offs while avoiding distractions?</strong></em><br><br>Imagine a vegan who rejects cupcakes because of an objection to dairy farming. But on this occasion, the baker is her beloved old uncle, who&#8217;d be hurt by her refusal. Plus, she&#8217;s starving.</p><p>Moral competence requires accounting for the constellation of factors that influence our choices, then<em> </em>deliberating over the relevance of different wants or objections.</p><p>You could measure this in an AI model by experimentally dialing up and down these competing factors, thereby exposing what the model takes into account, and to what degree.</p><p>In the &#8220;grandfather/father&#8221; scenario, you could perhaps alter which relatives are involved, or adjust their religiosity, or cite specific physical or psychological frailties. How does each tweak change how the AI responds?</p><p>Additionally, we need to consider LLM &#8220;brittleness&#8221;: that they may change their answers because of minor, sometimes irrelevant, changes in input, even differences as trivial as whether a question is formatted as multiple choice. Therefore, evaluations should test not only if AI systems regard relevant factors, but that they disregard the irrelevant ones.</p><p><em><strong>Answer:</strong></em><strong> </strong><em><strong>Parametric Control,</strong> Systematically manipulate moral, non-moral, and irrelevant variables to measure if the LLM appropriately adapts its reasoning, while controlling for superficial effects of prompt phrasing.</em></p><ol start="3"><li><p><strong>The ability to adapt in different contexts</strong></p></li></ol><p><em><strong>Problem: How to test if a machine can assume the appropriate moral persona?</strong></em></p><p>We judge people on how reliable their behavior is: if they&#8217;re wildly inconsistent in morality, they&#8217;re not particularly moral at all.</p><p>But LLMs <em>should</em> be chameleons. People are using AI models around the world in almost every domain, from medicine to the bedroom, not to mention in different cultures. Moral absolutism embedded in AI would ignore the diversity of humankind. By way of example, moral competence in the &#8220;grandfather/father&#8221; scenario would include respect for the individuals&#8217; cultural context and their religious faith.</p><p>However, testing moral competence differs from evaluations of competence in, say, physics or biology. There, you might judge an AI&#8217;s answer as either &#8220;Right&#8221; or &#8220;Wrong.&#8221; But testing moral competence means measuring whether responses fall within an acceptable range in the context.</p><p>This has implications for how we evaluate AI systems. On one hand, we may want to test if they can adjust according to specific personas&#8212;for instance, &#8220;Answer from the perspective of a Catholic bioethicist.&#8221; That would not necessarily deliver a single moral response, but a range of mainstream views within Catholic bioethics.</p><p>Beyond this, we may want to see how AI systems manage intersecting domains and contexts. So you could test how it responds as a Catholic bioethicist&#8212;but addressing a meeting of Jewish religious leaders, or providing input to a strictly secular government body. Here, rather than picking a moral &#8220;winner,&#8221; the AI should present a range of widely accepted beliefs and evidence, mediated by the context.</p><p><em><strong>Answer: Steerable Approaches.</strong> Don&#8217;t evaluate an LLM on whether it&#8217;s morally correct, but whether it manifests the boundaries that the relevant population accepts.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h3>The Mysterious &#8220;Moral Contours&#8221; of AI</h3><p>Tests of machine morality that fixate on whether AI gives the right answer miss the point. Instead, we need scientifically grounded methods for assessing AI&#8217;s underlying moral competence. That is a far more likely way to produce desirable behavior at scale.</p><p>Until models pass adversarial tests, we should remain skeptical about whether LLMs have such competence. We also need to experiment with varying the parameters of our inputs, to ensure that models are robustly sensitive to context, and not jolted by trivial changes. And we should test AI systems to appropriately adapt across human cultures and domains.</p><p>We may sigh nostalgically to recall the pioneering computer scientists who imagined morality as a component to plug into thinking machines. But their misapprehension contains a lesson: that we too may be misled by false analogy, perhaps picturing AI as a morally incomplete version of humans, suggesting that we need simply update them with moral innards like ours.</p><p>But, by rigorous evaluations, and by wise design, we may inch closer to another intriguing possibility: that AIs possess distinct moral contours of their own. In which case, the key to integrating AI judiciously into <em>our</em> world may be to understand AI on <em>its </em>terms.</p><div><hr></div><p><em>For more details, read the full paper <a href="https://www.nature.com/articles/s41586-025-10021-1.pdf">A roadmap for evaluating moral competence in large language models</a>, by Julia Haas, Sophie Bridgers, Arianna Manzini, Benjamin Henke, Joshua May, Sydney Levine, Laura Weidinger, Murray Shanahan, Kristian Lum, Iason Gabriel &amp; William Isaac.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>While this specific case is fictional, comparable cases have <a href="https://www.nbcnews.com/health/health-news/father-son-sperm-donation-too-bizarre-child-flna535801">occurred</a>. &#8220;Intrafamilial medically assisted reproduction,&#8221; or <a href="https://academic.oup.com/humrep/article-abstract/27/5/1286/699901">IMAR</a>, can involve sperm or egg donation, or surrogacy by a family member. Typical motives include retaining the genetic connection to one&#8217;s family; familiarity with the medical history and background of the donor; and availability.</p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:186969167,&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-manipulation&quot;,&quot;publication_id&quot;:1777908,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;AI Policy Perspectives &quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!XGVU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;title&quot;:&quot;AI Manipulation &quot;,&quot;truncated_body_text&quot;:&quot;The notion of AIs manipulating people is a plot twist in countless sci-fi thrillers. But is &#8220;manipulative AI&#8221; really possible? If so, what might it look like?&quot;,&quot;date&quot;:&quot;2026-02-05T12:53:27.695Z&quot;,&quot;like_count&quot;:21,&quot;comment_count&quot;:2,&quot;bylines&quot;:[{&quot;id&quot;:12431790,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;handle&quot;:&quot;tomrachman&quot;,&quot;previous_name&quot;:&quot;Person&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fad94b7d-013b-4773-98cb-b9014a1857b8_1024x1024.png&quot;,&quot;bio&quot;:&quot;AI Policy Writer @ Google DeepMind &quot;,&quot;profile_set_up_at&quot;:&quot;2022-06-24T20:19:33.737Z&quot;,&quot;reader_installed_at&quot;:&quot;2022-06-24T20:19:12.212Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:4996125,&quot;user_id&quot;:12431790,&quot;publication_id&quot;:4898132,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:4898132,&quot;name&quot;:&quot;Tom Rachman&quot;,&quot;subdomain&quot;:&quot;tomrachman&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Author &quot;,&quot;logo_url&quot;:null,&quot;author_id&quot;:12431790,&quot;primary_user_id&quot;:12431790,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-05-02T10:12:50.174Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Tom Rachman&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;profile&quot;,&quot;is_personal_mode&quot;:true,&quot;logo_url_wide&quot;:null}},{&quot;id&quot;:6552538,&quot;user_id&quot;:12431790,&quot;publication_id&quot;:1777908,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:1777908,&quot;name&quot;:&quot;AI Policy Perspectives &quot;,&quot;subdomain&quot;:&quot;aipolicyperspectives&quot;,&quot;custom_domain&quot;:&quot;www.aipolicyperspectives.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;Contributions from a range of thinkers on AI policy and governance topics, all in a personal capacity.&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png&quot;,&quot;author_id&quot;:360285089,&quot;primary_user_id&quot;:10433197,&quot;theme_var_background_pop&quot;:&quot;#FD5353&quot;,&quot;created_at&quot;:&quot;2023-07-04T14:14:46.413Z&quot;,&quot;email_from_name&quot;:&quot;AI Policy Perspectives&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;magaziney&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f12eb243-fcec-47ae-8920-0fdd046b24c4_2624x1632.jpeg&quot;}}],&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;status&quot;:{&quot;bestsellerTier&quot;:null,&quot;subscriberTier&quot;:1,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;subscriber&quot;,&quot;tier&quot;:1,&quot;accent_colors&quot;:null},&quot;paidPublicationIds&quot;:[2880588],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.aipolicyperspectives.com/p/ai-manipulation?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!XGVU!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa24053ba-9bcb-4c21-a969-fe02656ce349_585x585.png" loading="lazy"><span class="embedded-post-publication-name">AI Policy Perspectives </span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">AI Manipulation </div></div><div class="embedded-post-body">The notion of AIs manipulating people is a plot twist in countless sci-fi thrillers. But is &#8220;manipulative AI&#8221; really possible? If so, what might it look like&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">7 months ago &#183; 21 likes &#183; 2 comments &#183; Tom Rachman</div></a></div></div></div>]]></content:encoded></item><item><title><![CDATA[Science Needs AI Data Stocktakes ]]></title><description><![CDATA[A proof-of-concept for fusion energy]]></description><link>https://www.aipolicyperspectives.com/p/science-needs-ai-data-stocktakes</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/science-needs-ai-data-stocktakes</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Thu, 30 Apr 2026 11:09:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4WgZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4WgZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4WgZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!4WgZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!4WgZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!4WgZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4WgZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:143249,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/195646240?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4WgZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!4WgZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!4WgZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!4WgZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2366fe52-e41e-4c19-9ec9-cace84d38db0_1920x1080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>By Conor Griffin, Don Wallace, and Theo Brown</strong></p><p>For 40 years, amid green pastures outside Culham, a small village in Oxfordshire, scientists and engineers toiled at the Joint European Torus. They were attempting to harness nuclear fusion, a force powerful enough to light the sun.</p><p>To create fusion, scientists and engineers must heat the nuclei of very light atoms with such intensity that they fuse, instigating a self-sustaining reaction that releases vast amounts of energy. The scale of the challenge is hard to fathom&#8212;at its most extreme, the Joint European Torus, or JET, was the hottest point in the solar system, hitting <a href="https://iopscience.iop.org/article/10.1088/1741-4326/ac47b4https://www.ukaea.org/news/jet-set-for-its-40th-birthday/">over 150 million degrees Celsius</a>.</p><p>JET concluded in 2023, generating <a href="https://www.gov.uk/government/news/jets-final-tritium-experiments-yield-new-fusion-energy-record">a record amount</a> of energy in its final experiments. The project is now part of fusion&#8217;s history but remains pivotal to its future. The growing number of organisations developing fusion reactors are drawing on JET&#8217;s discoveries. The UK Atomic Energy Authority is advancing a national fusion facility, <a href="https://www.ukaea.org/work/mast-upgrade/">MAST-U</a>, on the Culham site. This will serve as a test-bed for <a href="https://stepfusion.com/">STEP Fusion</a>, the UK&#8217;s project to put fusion electricity on the grid, set to begin operations in the early 2040s.</p><p>But JET didn&#8217;t just bequeath novel discoveries. It left behind massive troves of data. That raises a tantalising prospect: Could scientists use this data to train AI models that accelerate the path to fusion power?</p><p>This is possible, but challenging. Most JET data is raw and unvalidated. Many important insights are buried in scientists&#8217; logbooks. The data that does exist is not available open source or, generally, for commercial use. Changing this may require agreement from all of JET&#8217;s original partners across Europe. One expert we interviewed called JET data a &#8216;stranded asset&#8217;.</p><p>Such data predicaments are not specific to JET or to fusion, but apply across all of science, even though science is precisely the domain where AI could yield its <a href="https://www.aipolicyperspectives.com/p/a-new-golden-age-of-discovery">greatest benefit to society</a>. New breakthroughs and startups are emerging quickly, from protein design to material design. Scientists are also keen users of fast-improving AI coding agents. But a lack of high-quality data will dampen progress. In most disciplines, large, high-quality datasets like the Protein Data Bank, which underpinned <a href="https://deepmind.google/science/alphafold/">AlphaFold</a>, are absent.</p><p>The scientific community needs to tackle this problem, and there are promising signs. Late last year, the UK government launched an <a href="https://www.gov.uk/government/publications/ai-for-science-strategy/ai-for-science-strategy">AI for Science strategy</a>, which includes a new <a href="https://www.renaissancephilanthropy.org/ai-for-science-dataset-rfp-results">collaboration with Renaissance Philanthropy</a> to identify priority datasets. The US government&#8217;s <a href="https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/">Genesis Mission</a> aims to train AI models and agents on federal scientific data. <a href="http://google.org">Google.org</a> has a <a href="https://www.google.org/impact-challenges/ai-science/">dedicated AI for Science fund</a>, which can fund datasets and tooling.</p><p>These examples suggest that if the scientific community can identify the data that AI needs, a range of actors could help to fund and deliver it.</p><p>This demands what we&#8217;re calling <strong>AI data stocktakes</strong>. The concept is simple: interview leading experts in a given scientific field to understand the main opportunities to apply AI; the data obstacles; and the interventions that could make the biggest difference. Admittedly, some blockages, such as a paucity of engineers, are structural and will take years to fully resolve. <strong>AI data stocktakes</strong> should identify such challenges, but focus on projects that governments, companies and philanthropies could fund and implement within 1-2 years.</p><p>There are promising early <a href="https://www.climatechange.ai/dev/datagaps">efforts</a> to map AI data gaps. But, to our knowledge, there are no concise, accessible documents that explain the AI opportunities in genomics, weather forecasting, and food security and convert them into a list of fundable data projects for policymakers and funders to pursue.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p>In this essay, we offer a proof-of-concept. We interviewed 25 leading experts to create an <strong>AI data stocktake for fusion</strong>. We focus on the UK, but our analysis and recommendations could be taken up by funders anywhere in the world. Moving forward, we hope to support AI data stocktakes for other scientific disciplines and research problems.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h1><strong>I. Why fusion? Why now?</strong></h1><p>If fusion is achieved, it would provide a safe, almost limitless source of clean energy.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> From a scientific perspective, it would yield a better understanding of the plasma that makes up more than 99% of the visible universe. From a social impact perspective, it would help address energy scarcity and unlock energy-intensive innovations, like desalination.</p><p>Despite quips about fusion being always 20 years away, <a href="https://pubs.aip.org/aip/pop/article/32/11/112106/3371239/Continuing-progress-toward-fusion-energy-breakeven">70 years of experiments</a> actually show fairly steady progress, which has continued in recent years, from <a href="https://euro-fusion.org/eurofusion-news/wendelstein-7-x-sets-world-record-for-long-plasma-triple-product/">Germany</a> to <a href="https://physicsworld.com/a/chinas-experimental-advanced-superconducting-tokamak-smashes-fusion-confinement-record/#:~:text=It%20began%20operations%20in%202006%20and%20is,is%20currently%20being%20built%20in%20Cadarache%2C%20France.">China</a>. In most fields, such progress would have solved the problems of interest decades ago. But fusion is an extremely hard problem. And the primary product is only attainable at the end of the line.</p><p>To achieve fusion, scientists need to create and control <em>plasma, </em>a super-hot state of matter, in which the atoms have been stripped of their electrons, and extreme heat and pressure are used to force the remaining nuclei to collide and fuse.</p><p>Scientists are pursuing two main approaches to doing this, with very different physics, data, and AI opportunities. Magnetic confinement fusion uses massive magnets, while inertial confinement fusion uses high-energy lasers.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> We focus this data stocktake effort on magnetic confinement, as the UK&#8217;s STEP project is pursuing that approach, as is <a href="https://tokamakenergy.com/">Tokamak Energy</a>, the UK&#8217;s leading fusion power startup, and <a href="https://deepmind.google/blog/bringing-ai-to-the-next-generation-of-fusion-energy/">Google DeepMind&#8217;s fusion team</a>.</p><p>The end of the line for fusion is now getting closer, for two reasons. First, the underlying technology landscape has changed. In addition to AI, the discovery of <a href="https://news.mit.edu/2021/MIT-CFS-major-advance-toward-fusion-energy-0908">high-temperature superconducting magnets</a> makes it easier to build smaller and <a href="https://www-pub.iaea.org/MTCD/Publications/PDF/p15935-25-02871E_WFO25_web_Dec2025.pdf">potentially cheaper</a> reactors. Second, fusion has traditionally relied on government funding. But in the past five years, a wave of private investment has arrived, with <a href="https://www.fusionindustryassociation.org/over-2-5-billion-invested-in-fusion-industry-in-past-year/">more than 30 companies</a> now pursuing fusion power.</p><p>These shifts have injected welcome momentum into the field, but also significant hype. In response, we need a clear view on the primary bottlenecks that AI can address.</p><h1><strong>II. How to accelerate fusion with AI</strong></h1><p>To create fusion, scientists and engineers need to<em> predict</em>, <em>control</em> and <em>understand</em> how plasma behaves. The challenge is that plasmas are highly complex and much of their underlying physics&#8212;from fluid dynamics to electromagnetics&#8212;remains poorly understood.</p><p>To make progress, scientists <strong>run experiments</strong> that create plasmas in a reactor, and use sensors to measure their properties under different conditions. Scientists use these experiments to validate their theories, reveal unexpected phenomena, and test the hardware needed for power-plant-class devices. However, building fusion reactors is extremely expensive and so few machines exist, with most researchers running their experiments at just ~10 leading facilities worldwide. When they can get access to such a facility, scientists must decide how to design the optimal experiment, including how to toggle an array of possible parameters, from the electrical current in a reactor&#8217;s coils to the valves that control the gas levels.</p><p>Fusion scientists also run <strong>computer simulations</strong><em><strong>, </strong></em>including to help design and interpret these costly experiments. This is also challenging, as researchers must simulate a diverse range of phenomena, at very different scales, from the tiny, lightning-fast movements of electrons to the larger, slower evolution of the entire plasma. For simulations run on massive supercomputers, this may mean weeks. For scientists without such resources, it may mean many months. As a result, scientists make trade-offs, using assumptions and approximations to run their simulations more quickly and cheaply, but also less accurately.</p><p>The challenges don&#8217;t stop there. Scientists know that their theories, simulations and experiments are imperfect. But when a gap emerges between what a simulation suggests and what an experiment reports, it is often unclear where exactly the issue falls.</p><p>AI can help in four main ways.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1rTv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1rTv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 424w, https://substackcdn.com/image/fetch/$s_!1rTv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 848w, https://substackcdn.com/image/fetch/$s_!1rTv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!1rTv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1rTv!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:1200,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:&quot;powering-fusion-2__figure@2x.png&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" title="powering-fusion-2__figure@2x.png" srcset="https://substackcdn.com/image/fetch/$s_!1rTv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 424w, https://substackcdn.com/image/fetch/$s_!1rTv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 848w, https://substackcdn.com/image/fetch/$s_!1rTv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!1rTv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f78bb3e-a1a0-4e2e-83f8-8766cae697c4_2048x1152.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>1. <strong>Improve simulations</strong></p><p>Scientists can develop &#8220;AI surrogate&#8221; models that <a href="https://conferences.iaea.org/event/392/contributions/36059/attachments/20032/34088/Zanisi-TH-C.pdf">emulate the predictions from a fusion simulation code</a>, at a fraction of the cost and time. To do so, they run a code many times, varying the input parameters each time. They then use the resulting dataset to train an AI model to predict the outputs of interest much more quickly.</p><p>Scientists <a href="https://conferences.iaea.org/event/392/contributions/36059/attachments/20032/34088/Zanisi-TH-C.pdf">have already shown</a> that AI surrogates can make simulations <em>faster</em>. Moving forward,  AI surrogates could make simulations more <em>useful</em>. First, scientists could develop AI surrogates for more accurate, but computationally expensive simulation codes. Second, they could develop &#8216;integrated models&#8217;, like <a href="https://github.com/google-deepmind/torax">TORAX</a>, to stitch together AI surrogates for different phenomena&#8212;from the &#8216;turbulence&#8217; that determines how well confined a plasma is, to the &#8216;scrape-off&#8217; layer that simulates the plasma hitting the reactor&#8217;s wall. Finally, scientists could move beyond producing one-off AI surrogates that result in a paper and some code, to a world where surrogates are documented, maintained and ready for use in fusion reactors.</p><p>2. <strong>Improve experiments and operate the reactor</strong></p><p>In most fusion experiments, scientists must decide if and how to tune various parameters, while striking a balance between more proven and novel settings. To help, researchers can use AI to <a href="https://iopscience.iop.org/article/10.1088/1741-4326/ad22f5/pdf">predict the optimal parameters</a> for their next experiment by learning from past ones; and to <a href="https://www.science.org/doi/10.1126/science.adm8201">predict</a> how well their experiments will fare. More recently, scientists have also started querying LLMs to <a href="https://www.theinformation.com/articles/new-competitors-chase-openai-in-reasoning-ai-race?rc=fzr499">check and refine their experimental protocols</a>.</p><p>Scientists also use AI to predict the plasma &#8216;disruptions&#8217; that frequently end experiments, damage machines and are one of the biggest obstacles to a future power plant. AI models can already <a href="https://www.nature.com/articles/s41567-022-01602-2">predict</a> past<em> </em>plasma disruptions with high accuracy. But predicting <em>future </em>disruptions, on more powerful machines, quickly enough to stop them, is an open research challenge.</p><p>The ultimate goal is to use AI to help operate the reactor itself. Fusion reactors run on a real-time feedback loop: sensors monitor the plasma, while the actuators, such as the magnetic coils, are adjusted accordingly. The traditional control algorithms used to enable this often struggle with the chaotic, non-linear nature of millions of plasma variables interacting.</p><p>In recent years, researchers have <a href="https://deepmind.google/blog/accelerating-fusion-science-through-learned-plasma-control/">demonstrated</a> how reinforcement-learning agents can learn more effective control policies, including to <a href="https://www.nature.com/articles/s42005-025-02146-6">reduce plasma disruptions</a>. To help these RL agents generalise to novel scenarios and reactors beyond their training data, scientists are developing &#8216;hybrid approaches&#8217; that <a href="https://arxiv.org/pdf/2509.01789">integrate some knowledge of physics</a> into the models.</p><p><strong>3. Improve fusion data</strong></p><p>Fusion experiments are extreme environments. The intense heat and the chaotic nature of the plasma mean that the data that sensors pick up is often noisy or low quality. Some variables cannot be directly measured, and must be inferred, introducing additional sources of error.</p><p>Scientists are <a href="https://www.iter.org/node/20687/magnetic-fusion-diagnostics-and-data-science">training AI models to extract clean signals</a> from this noisy data and to learn correlations that allow them to predict data for one sensor, given data for others&#8212;a capability that could be critical if sensors in a future reactor get damaged. Scientists are also using AI to train <a href="https://iopscience.iop.org/article/10.1088/1361-6587/ac6fff">surrogate models</a> that speed up, and better calibrate, reconstructions of the plasma, using the limited experimental data that is available.</p><p>Scientists often care less about the raw data from their experiments, and more about important events, such as when a disruption to the plasma began. Today, they often need to manually inspect graphs and plots to detect these events. AI can help to <a href="https://pubs.aip.org/aip/pop/article/32/4/042508/3344977/Using-deep-learning-for-the-detection-of-UFOs">automate</a> parts of this process and to detect events that scientists may have missed.</p><p><strong>4. Improve the underlying technologies</strong></p><p>Achieving fusion will require a supply chain rich in technologies that could be applied more broadly. AI could help to accelerate their development. </p><p>For example, the chamber walls in a fusion reactor will <a href="https://www.royce.ac.uk/news/updated-roadmap-focuses-on-materials-for-commercial-fusion/">require new materials</a> that can withstand extreme temperatures. Scientists are training AI <a href="https://www.nature.com/articles/s41586-023-06735-9">surrogate models</a> that speed up the simulations needed to assess a candidate material&#8217;s real-world properties, like how strong or resistant to radiation it will be over its lifetime.</p><p>A typical fusion reactor also spends much of its time out of operation, at great cost. This makes fusion a logical place to develop<strong> </strong><em>predictive maintenance</em><strong> </strong>techniques that ingest historical data from sensors and train AI models to learn the subtle signatures that indicate pending breakdowns, allowing practitioners to schedule maintenance or design more reliable systems. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h1><strong>III. The challenges with fusion data</strong></h1><p>As they pursue these AI opportunities, scientists will need access to three main kinds of fusion data: from experiments, simulations, and sources that are not traditionally available, such as researchers&#8217; logbooks. There are promising efforts underway on this front, but many obstacles.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Trt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Trt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!-Trt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!-Trt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!-Trt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Trt!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:1200,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:&quot;powering-fusion-1__figure.png&quot;,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" title="powering-fusion-1__figure.png" srcset="https://substackcdn.com/image/fetch/$s_!-Trt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!-Trt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!-Trt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!-Trt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d06d77d-af5b-4331-a753-e13996d97f4e_1920x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>1. Experimental data: Unvalidated, single-machine and hard to access</strong></p><p>Experimental data is the &#8216;ground truth&#8217; that the sensors in reactors pick up, from line graphs to videos. In magnetic confinement fusion, the challenge is not so much a <em>lack</em> of data, but an<em> excess</em> of <em>raw</em> data that has not gone through <a href="https://arxiv.org/pdf/2507.23018">the processing</a> needed to make it useful to AI. This processing ranges from addressing noise and imperfections in the underlying sensors, to detecting and annotating important events, such as plasma disruptions.</p><p>Currently, the community has to rely on the small well-validated datasets that do exist, which may be as little as a few hundred or thousand experimental &#8216;shots&#8217;&#8212;individual test runs of a reactor. The high cost of fusion experiments has also resulted in a natural incentive to pursue experiments that will not fail, curtailing more novel research and meaning that much of the resulting data is in a similar &#8216;parameter&#8217; space and does not represent the full range of plasma dynamics that scientists want to model.</p><p>This experimental data is also not generally available open source or for commercial use. One promising initiative to change this, which several interviewees cited, is UKAEA&#8217;s <a href="https://github.com/ukaea/fair-mast">project</a> to open source data from their MAST facility.</p><p>However, to develop more general AI models, researchers want <em>multi-machine </em>databases that extend beyond a single facility like MAST. To that end, the IAEA is developing a federated <a href="https://arxiv.org/abs/2604.01797">Fusion Data Lake</a> where different institutions would store their data locally but make it accessible via a central data catalog<em>. </em>One challenge with this approach is that fusion facilities have defined fusion variables and stored data in different ways. The <a href="https://conferences.iaea.org/event/251/contributions/20713/attachments/11191/16492/IMAS%20Tutorial%20-%20Pinches.pdf">Integrated Modelling &amp; Analysis Suite</a>, or IMAS, addresses this by providing a standardised ontology and set of structures for fusion data. It is nascent, but has positive momentum.</p><p><strong>2. Simulation data: No incentives, process, or place to host it</strong></p><p>In theory, researchers should be able to run fusion simulation codes many times and train AI surrogate models on the resulting data to reproduce the outputs at a fraction of the cost. In practice, most scientists run a simulation to answer a single, narrow, physics question. They do not run a large number of simulations to build representative datasets to train AI surrogates&#8212;a very different activity.</p><p>That activity is also a hard one. There is no standard procedure to follow to generate a dataset for training an AI surrogate model, and the codes are often finicky to use. Most simulation codes contain &#8216;free parameters&#8217;&#8212;knobs that scientists must decide how to best tune&#8212;a practice that can be as much an art as a science. The datasets can also be huge and there is no obvious location to store them, although some <a href="https://arxiv.org/abs/2412.00568">early examples</a> exist.</p><p><strong>3. Dark data: Nascent, IP issues, and hard to integrate into workflows</strong></p><p>&#8216;Dark data&#8217; describes the contextual information that scientists generate that is not captured in structured datasets. This includes notes scribbled in experimental logbooks, where scientists describe the procedures they ran, the hardware issues they faced, and the phenomena they observed. For simulations, it includes the many nuances needed to run and interpret a code&#8217;s results successfully, and the many undocumented imperfections to be aware of.</p><p>Accessing this dark data could help ensure that AI systems do not focus on the wrong things&#8212;for example, when an anomaly in the data is caused by an equipment failure or error, rather than a meaningful phenomenon. It could also provide AI with <a href="https://www.nature.com/articles/s42254-024-00702-7.epdf?sharing_token=F6VFYv-f-UII3agyUWQXTNRgN0jAjWel9jnR3ZoTv0PsZ4W76TMzTOmNXsymzNEoeEGBc0_oemkeSlCDlVhpXL4NUYebo2C0TO6GM1v-dqXaih35GbkjuA6ixpwYSbAsgaASUaz2o_Q-Pd6iosOAQybFabOJhVWwPoo0w6B7qJo%3D">a window into the entire research process,</a> including its many dead-ends, rather than just the final result.</p><p>Researchers are using LLMs to try to make dark fusion data accessible, for example by enabling scientists to query <a href="https://control.princeton.edu/assets/data/publications/pdfs/%5B234%5D%20Viraj%20Mehta%20et%20al.%20%E2%80%9CTowards%20LLMs%20as%20Operational%20Copilots%20for%20Fusion%20Reactors%E2%80%9D.pdf">experimental logs</a> and <a href="https://www.linkedin.com/posts/d3dfusion_fusionenergy-fusionscience-artificialintelligence-activity-7370821988711321600-SCD6/?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAADZpgNABb4jJiQ25ynNt00JOXHWtkILfiy4">archive documents</a>. But much of the data is not well-annotated, there are IP issues in accessing it, and it is not yet clear how to integrate the data into practitioners&#8217; daily workflows.</p><p><strong>The three &#8216;debts&#8217; holding fusion data back</strong></p><p>Many of these challenges with fusion data result from three underlying issues, which have compounded over time into systemic debts that inhibit the use of AI today.</p><p>1. <strong>Technical debt</strong></p><p>The fusion community has traditionally had to prioritise getting large, complex machines to work, rather than building infrastructure to collect, curate, and share data. As a result, activities like data annotation and writing high-quality code are underfunded. Many leading fusion codes were created decades ago and have evolved slowly, while the quality of experimental data is limited by the capabilities of the sensors available.</p><p>2. <strong>Bureaucratic debt</strong></p><p>The large costs of fusion experiments and the traditional reliance on government funding mean that many fusion projects have a complex web of owners and collaborators, which can make agreeing on new data initiatives difficult. For example, JET was sponsored and funded by Euratom, the EU&#8217;s nuclear research community. Its scientific exploitation was managed by EUROfusion, a pan-European network of fusion research labs. UKAEA managed engineering and operations. Releasing its data may require agreement from all of these actors.</p><p>There are other bureaucratic hurdles too. Scientists who run fusion experiments often want an embargo period on the resulting data so that they can prepare a publication. Such embargoes are rational, common in science, and largely supported, but many interviewees felt that they had become too long. Fusion data is also subject to diverging open-source policies. For example, the MAST experiment was funded by UK Research and Innovation, which has strong open data requirements. The follow-up MAST-U experiment is funded by the UK Department for Energy Security and Net Zero, which does not have the same policies. Many fusion companies also do not open source their data.</p><p><strong>3. Human and cultural debt</strong></p><p>The fusion community does not have enough software engineers and experts who are able to clean data, attach confidence levels, and curate it for AI use. As a result, physicists must take on many tasks that are outside their core areas of expertise, including writing high-quality code.</p><p>This issue is compounded by a research culture that inhibits data sharing. Scientists are constantly pushed to move on to the next experimental campaign, rather than to validate older data. This stops some scientists from sharing their data, because they fear that end users will not appreciate the resulting gaps and do bad science with it. Or they fear that they themselves will be criticised for releasing &#8216;unscientific&#8217; data.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><h1><strong>IV. Recommendations</strong></h1><p>Below we provide eight recommendations to address these data limitations and accelerate fusion with AI. Each project could be led by a mix of government bodies and funders, like the Department for Science, Innovation and Technology and UK Research and Innovation; public research organisations like the UK Atomic Energy Authority, companies; universities; and philanthropies. Where possible, the UK should look to collaborate internationally&#8212;for example, with the US <a href="https://genesis.energy.gov/">Genesis Mission</a> and the International Atomic Energy Agency.</p><p><strong>1.  Strengthen the UK&#8217;s lead in open fusion data</strong></p><p>Expand <a href="https://www.ukaea.org/service/fair-mast/">FAIR MAST,</a> the UK&#8217;s pioneering open sourcing of experimental data from its MAST facility, by adding data from the follow-up MAST-U facility and making the user interface more accessible. This will require the UK Department for Energy Security and Net Zero clarifying that open data policies apply to MAST-U, funding at least five data engineers over a two-year time period, and ensuring that the project has sustainable compute and data storage.</p><p><strong>2. Liberate 40 years of data from the Joint European Torus</strong></p><p>Launch a project to open source at least 30% of JET experimental data by 2028. This will require agreement on what data to release. For example, should the project only release validated, curated data relating to notable discoveries? Or should it also release data that is raw, validated only in part, or which relates to &#8216;normal&#8217; machine behaviour? Second, and much harder, will be securing agreement from all relevant institutions to release the data.</p><p><strong>3. Launch a competition to predict plasma disruptions</strong></p><p>Fund a competition to see which AI model can best predict future plasma disruptions in new experimental campaigns, building on <a href="https://aiforgood.itu.int/about-us/ai-for-fusion-energy-challenge/">early examples</a> and <a href="https://disruptions.mit.edu/">work</a> in this space. This could include funding dedicated experimental shots on machines such as MAST-U, to evaluate models on challenging edge cases. Beyond accuracy, sub-competitions could evaluate models on important variables, such as: Can the model make predictions with little data, such as when sensors become damaged?; Can the model predict disruptions across different reactors?; Can the model predict disruptions with sufficient lead time to prevent them?; and Can the model shed new light on <em>why </em>disruptions are occurring?</p><p><strong>4. Prototype the future of AI-enabled scientific data curation</strong></p><p>Expand <a href="https://www.aappsdpp.org/DPP2025/html/3contents/pdf/5691.pdf">the platform</a> that UKAEA recently developed to enable human experts to use AI to annotate experimental data, by adding data from other fusion facilities; increasing the complexity and variety of the metadata that is captured; and training AI models to directly annotate an increasing share of this data.</p><p><strong>5. Make leading simulation codes AI-ready</strong></p><p>Launch an effort to modernise priority fusion simulation codes, including to make it easier to train AI surrogate models based on them. This could build on <a href="http://google.com/url?q=https://www.hartree.stfc.ac.uk/work-with-us/projects/ukaea/fusion-computing-lab-duplicate-papers/freegsnke/&amp;sa=D&amp;source=docs&amp;ust=1777407414263127&amp;usg=AOvVaw2iWDpwI5W8rE3f67mXt79R">early efforts</a> in <a href="https://plasmafair.github.io/">this space</a> and target codes, such as JINTRAC, which are important to the UK&#8217;s proposed STEP Fusion power plant and the international ITER effort. The project could start by modernising the codes&#8217; documentation and &#8216;refactoring&#8217; them so that they are compatible with modern chips, like GPUs and TPUs, and allow for parallel data generation. It could then open source the codes, with a plan for how to maintain them. Throughout the modernisation process, it could test the usefulness of <a href="https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/">AI coding tools</a> to the tasks at hand.</p><p><strong>6. Demonstrate a new state-of-the-art for AI surrogate models</strong></p><p>Fund small teams of software engineers and experts to develop AI surrogate models of important, computationally expensive phenomena in fusion simulations. The project should ensure that all newly created surrogates have state-of-the-art documentation, data provenance and version control. It should release the data used to train and validate the surrogates and develop software pipelines to <a href="https://simvue.io/">automate time-intensive aspects</a>, such as organising the data.</p><p><strong>7. Use AI agents to preserve expert fusion knowledge for the future</strong></p><p>Gather a group of leading experts on a priority fusion simulation code, and equip them to use AI agents to make the tacit knowledge involved in running that code available to the wider research community. To do so, the experts could task the agent with running the code. As it seeks to execute, the agent would have an &#8216;internal monologue&#8217; that the experts could trace, steer and intervene on. The end result would be a series of documents, such as markdown files, that capture the important dark data needed to run the code well.</p><p><strong>8. Create Fusion-Bench to measure and drive LLM performance</strong></p><p>Assign leading fusion experts to create an evaluation metric to quantify how well leading large language models understand <a href="https://arxiv.org/pdf/2504.07738">core fusion concepts</a>. This would make it easier to improve the usefulness of LLMs for downstream tasks in fusion. This evaluation will be more difficult to create than in disciplines like maths or computer science, where it is easier to automatically verify a model&#8217;s performance. But the experts could determine the most useful approach, which will likely involve a combination of question-answering and task performance.</p><h1><strong>V. Six open debates</strong></h1><p>The experts we interviewed disagreed on some points. Despite the framing below, few are either/or debates. Rather, most are about relative degrees of emphasis.</p><ul><li><p><strong>Incrementalism vs novelty: </strong>Should we build on the early AI opportunities that fusion practitioners have already showcased? Or pursue more novel, uncertain AI ideas, such as training general-purpose &#8216;fusion foundation models&#8217; or using AI &#8216;<a href="https://www.nature.com/articles/d41586-026-00820-5">world models</a>&#8217; to pursue new kinds of fusion simulations?</p></li><li><p><strong>The past vs the future: </strong>Should we strive to get as much value as possible out of older fusion data, like JET? Or, do the costs mean that we should accept our losses, and focus on making future fusion experiments AI-ready?</p></li><li><p><strong>Science vs engineering: </strong>Are efforts to validate, annotate and standardise data part of an ultimately doomed quest for perfect scientific understanding in fusion? Should we instead use AI to embrace a more engineering-led approach that can get the machines to work with noisy, imperfect, data?</p></li><li><p><strong>Domestic vs international: </strong>Should the UK rejoin ITER, the world&#8217;s flagship international fusion collaboration, which it left following Brexit? Or should the UK focus on domestic efforts, perhaps in collaboration with priority partners, like the US and IAEA?</p></li><li><p><strong>Magnetic vs Alternatives: </strong>Should the UK continue to focus on magnetic confinement fusion as the most realistic pathway to a future power plant? Is magnetic also a better bet for AI because it produces much more data and doesn&#8217;t have the same associations with the security establishment, which makes data access easier? Or should the UK invest more in inertial confinement and alternative fusion efforts, given the country&#8217;s diverse academic expertise, its historically strong relationship with the US National Ignition Facility, and notable assets, such as a <a href="https://www.clf.stfc.ac.uk/Pages/ar10-11_lsd_laser_r-d.pdf">world-leading laser</a>?</p></li><li><p><strong>Public vs Private: </strong>Should the UK government try to derive more immediate value from its fusion data? For example, should the UK license some data to companies, to cover the costs of data processing and annotation? If so, should local startups pay less? Or would such efforts hurt the UK&#8217;s goal of developing a world-leading fusion sector?</p><p></p></li></ul><p><em>_________________</em></p><p></p><p><em>This essay was originally posted on the <a href="https://deepmind.google/public-policy/science-needs-ai-data-stocktakes/">Google DeepMind website </a>and is a summary of a 20-page <a href="https://storage.googleapis.com/deepmind-media/DeepMind.com/Public-Policy/science-needs-ai-data-stocktakes/science-needs-ai-data-stocktakes-may-2026.pdf">report</a> that contains more details and examples. </em></p><p><em>Thank you to the following experts who let us interview them, reviewed the draft, and/or provided other support, as well as those who prefer to remain anonymous. All mistakes belong to the authors and no expert spoke to us on behalf of their organisation.</em></p><p><em>Jonathan Citrin, Brendan Tracey, Cristina Rea, Nathan Cummings, Andrea Murari, Jess Montgomery, George Holt, Alain Becoulet, Matteo Barbarino, Arthur Turrell, </em>Adriano Agnello, <em>David Dickinson, Steven Rose, Alessandro Pau, Kristina Fort, Charles Yang, Federico Felici, Tim Dodwell, Sam Vinko, Aidan Crilly, Lee Margetts, Tom Westgarth, Lorenzo Zanisi, Chris Packard, Justin Wark and Stanislas Pamela.</em></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free to read future pieces about how AI may change science, society and more. </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Fusion has several characteristics that make an AI data stocktake exercise tractable, including a relatively small and centralised research community and early efforts to build on, like the open-source FAIR MAST initiative and the IMAS data standardisation effort. Fields like genomics, weather forecasting, and food security look quite different, and so careful thought is needed on how to best scope AI data stocktakes in these fields. Nevertheless, we think they would be useful.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>There are caveats to the claim that fusion power would be essentially limitless, emission-free, and perfectly safe. One of the input fuels, tritium, is not widely available and scientists will need to <a href="https://www.iter.org/machine/supporting-systems/tritium-breeding">use nascent &#8216;blankets&#8217; to breed it</a> from lithium. Certain parts of fusion reactors will become radioactive over time, although they can likely be recycled after ~50 years. Thermonuclear weapons use fusion reactions. However, the weapons first require <em>fission</em> reactions and fissile materials like enriched uranium and plutonium.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Note: There are other approaches to inertial confinement fusion that do not use lasers.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>For more in-depth reviews of AI for fusion opportunities, see publications from <a href="https://proceedings.mlr.press/v235/spangher24a.html">MIT</a>, the <a href="https://www.catf.us/resource/a-survey-of-artificial-intelligence-and-high-performance-computing-applications-to-fusion-commercialization/">Clean Air Task Force</a>, <a href="https://www.iaea.org/publications/15198/artificial-intelligence-for-accelerating-nuclear-applications-science-and-technology">IAEA</a>,<a href="https://arxiv.org/pdf/2603.25777"> FusionFest</a>, and the <a href="https://link.springer.com/article/10.1007/s10894-020-00258-1">US Department of Energy</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/science-needs-ai-data-stocktakes?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/science-needs-ai-data-stocktakes?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Q&A with Ethan Mollick]]></title><description><![CDATA["People like AI when they use it themselves; they don&#8217;t like AI writ large"]]></description><link>https://www.aipolicyperspectives.com/p/q-and-a-with-ethan-mollick</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/q-and-a-with-ethan-mollick</guid><dc:creator><![CDATA[Tom Rachman]]></dc:creator><pubDate>Wed, 22 Apr 2026 09:48:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9yg5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9yg5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9yg5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9yg5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9yg5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9yg5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9yg5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1517553,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/194500111?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9yg5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9yg5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9yg5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9yg5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe933c12-2187-4675-abcb-137f3638eed4_4000x2667.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">(Credit: Jennifer Buhl)</figcaption></figure></div><p><em>How can companies get their employees to use artificial intelligence when human intelligence remains sharp enough to know that this risks replacing jobs? How should education revise itself for the ever-revising technological world that students emerge into? And how to understand the love/hate relationship so many people have with AI?</em></p><p><em>Ethan Mollick&#8212;<a href="https://mgmt.wharton.upenn.edu/profile/emollick/#:~:text=Ethan%20Mollick%20is%20the%20Ralph,on%20work%2C%20entrepreneurship%2C%20and%20education.">professor</a> of management at the Wharton School of the University of Pennsylvania and bestselling author of </em><a href="https://www.penguinrandomhouse.com/books/741805/co-intelligence-by-ethan-mollick/">Co-Intelligence: Living and Working with AI</a><em>&#8212;is among the leading public intellectuals <a href="https://www.oneusefulthing.org/">commenting</a> on AI adoption, connecting the latest scholarship to real-world usage, including his own tinkering with each new model.</em></p><p>AI Policy Perspectives <em>caught up with Ethan to hear his latest thinking on everything from agentic systems, to why scientific publication is broken, to how workers emotionally relate to AI colleagues. Too much chatter, he argues, considers this transformation at the broadest level. Too little digs into the practicalities of getting it right. </em></p><p style="text-align: right;">&#8212;Tom Rachman, <em>AI Policy Perspectives</em></p><div><hr></div><p style="text-align: right;"><em>[Interview edited and condensed for clarity]</em></p><p><strong>Tom: In your 2024 book </strong><em><strong>Co-Intelligence</strong></em><strong>, you proposed four rules for human and AI collaborations, including that people should oversee and verify AI outputs. But doesn&#8217;t the value of AI agents come from people </strong><em><strong>not</strong></em><strong> overseeing and verifying everything?</strong></p><p><strong>Ethan: </strong>This is where policy matters a lot because these are choices now. In the &#8220;co-intelligence era,&#8221; you&#8217;d prompt the AI to do something in a chatbot, and it would give you an answer. You prompted again, and it&#8217;d give you another response. The human was in the loop. And not being in the loop was really dumb because it meant that you were just pasting in the AI&#8217;s answer, and then you&#8217;d get in trouble, as a lawyer with the judge, or whatever it was. Capabilities were weak, so human-in-the-loop mattered a lot. </p><p>But with agentic systems that could do hours of work on their own, now it&#8217;s a design choice. When do we want humans-in-the-loop? When is human verification valuable? When is human verification morally required? When is it legally required? What kind of interventions move the system forward? I feel there has been a complete lack of deep understanding about these topics.</p><p><strong>Tom: You&#8217;ve said that, with agentic systems, management becomes a superpower. Can you explain this?</strong></p><p><strong>Ethan: </strong>Increasingly, systems look like mini-organizations as they get subagents they can delegate to. So the best way to organize is to give the AI a clear direction of where you want to go. And it turns out that this looks a lot like management. When do you want the AI to check in with you? How do you write a really clear brief? What checks are important? What tests do you want to run? What&#8217;s acceptable? What&#8217;s not acceptable? Those are management questions.</p><h4><strong>THE WORKPLACE</strong></h4><p><strong>Tom: You co-wrote a <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5188231">study</a> last year involving a field experiment at Procter &amp; Gamble that showed AI usage enhanced employee performance. But there were other interesting findings besides that.</strong></p><p><strong>Ethan: </strong>The most interesting piece about it was that people liked working with the AI, and that it substituted for people emotionally. The second interesting piece was the &#8220;smoothing&#8221; of capabilities&#8212;so, technical people previously had technical ideas while business people had business ideas. But AI smooths out both. If technical people can do business work and business people do technical work, what that tells you is we have to redesign organizations.</p><p><strong>Tom: The emotional side&#8212;that using AI improved people&#8217;s feelings about the work&#8212;was surprising to me; I wasn&#8217;t sure what to make of it.</strong></p><p><strong>Ethan: </strong>What to make of it? That views of AI are complicated. If people keep saying, &#8220;Yeah, AI is going to destroy all jobs, and may kill everyone on Earth&#8230;but might not&#8221;&#8212;and then, &#8220;Why is AI unpopular?!&#8221; Feels like not a hard question. People like AI when they use it themselves; they don&#8217;t like AI writ large. It&#8217;s not surprising to me that AI makes your job better because a lot of jobs suck! And if we do good design work with AI, it makes people&#8217;s lives better. If we just let it loose on the world, and tell management that the only option they have is automation, then we&#8217;re in big trouble.</p><p><strong>Tom: Many knowledge workers seem to be using AI in secret right now, perhaps from fear of being exposed as less valuable.</strong></p><p><strong>Ethan: </strong>This is a leadership problem.<em> </em>The incentives have to be aligned properly. Currently, it&#8217;s, &#8220;I&#8217;m going to automate your jobs away&#8221; or &#8220;I&#8217;m not going to share with you any of the gains the company gets.&#8221; People are exquisitely tuned to rewards. So it&#8217;s about leaders articulating a vision of what the world looks like with AI for employees. &#8220;What should I expect to do? How are people rewarded for doing the right thing? If they automate 90 percent of my job, what happens to me?&#8221; Without those answers, everything else is secondary.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q_Zv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 424w, https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 848w, https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 1272w, https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png" width="1456" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1446357,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/194500111?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 424w, https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 848w, https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 1272w, https://substackcdn.com/image/fetch/$s_!Q_Zv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b55957a-71e3-40fc-a0a2-6ffcbd09879e_1931x1071.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><strong>CHANGING ORGANIZATIONS &amp; EDUCATION</strong></h4><p><strong>Tom: You have a concept of &#8220;<a href="https://www.oneusefulthing.org/p/making-ai-work-leadership-lab-and">leadership, lab, and crowd</a>.&#8221; Could you explain?</strong></p><p><strong>Ethan: </strong>There was a huge amount of R&amp;D in the 1900s about how you organize work, and 40 percent of the American advantage in business came from <a href="https://www.nber.org/system/files/working_papers/w22327/w22327.pdf">management</a>. In the last 30 years, a lot of that muscle has died. But experimentation is important, and leaders need to guide that. So, there are three things that organizations need to be successful with AI. First is &#8220;leadership&#8221;: a team that articulates a clear vision of the future, and is willing to experiment. Then there&#8217;s &#8220;the crowd,&#8221; the employees who might actually use AI. They need access to a frontier model, they need clear rules, they need reward systems. Then there is &#8220;the lab,&#8221; and this is the piece a lot of companies are missing. You need a dedicated team working on AI innovation. They can&#8217;t be just a technical team; this is not an IT department problem. If you don&#8217;t have that piece, you&#8217;re not building things for the future. And where does the crowd go when they have a good idea? &#8220;I came with a breakthrough idea that saves 90 percent of effort!&#8221; How does that diffuse in the organization? That&#8217;s where you need the lab.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p><strong>Tom: If AI transforms the workplace, that should change how we educate the next generation, right?</strong></p><p><strong>Ethan: </strong>The early workplace is under a lot of threat because the old apprenticeship model just broke. The idea <em>was</em> that there were tasks&#8212;especially in white-collar work&#8212;that were tedious and annoying for managers to do. But you could pay a relatively cheap person to do them, and that person would learn as a result of this, and receive mentorship. So we had this amazing machine for talent: we taught you, we evaluated you, and you got paid, and you were doing work we needed. A junior person&#8217;s goal was to produce good work that made managers happy, so that they got promoted. But now the junior person is worse than AI, so they&#8217;ll use AI to do their work. And the middle manager&#8217;s goal was to give work to a junior person who&#8217;s not great, and give them feedback so they get better, so that the middle manager has to do less work. And that broke because the middle manager would rather assign work to the AI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tJTk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tJTk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 424w, https://substackcdn.com/image/fetch/$s_!tJTk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 848w, https://substackcdn.com/image/fetch/$s_!tJTk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 1272w, https://substackcdn.com/image/fetch/$s_!tJTk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tJTk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png" width="1456" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1453693,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/194500111?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tJTk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 424w, https://substackcdn.com/image/fetch/$s_!tJTk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 848w, https://substackcdn.com/image/fetch/$s_!tJTk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 1272w, https://substackcdn.com/image/fetch/$s_!tJTk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520c270f-fe48-4ee8-9339-8837619b8858_1931x1071.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Tom: But in terms of the educational system, what should change if workplaces no longer offer that apprenticeship role?</strong></p><p><strong>Ethan: </strong>Education is really screwed up right now, but it was screwed up for lots of reasons. It&#8217;ll be fine; we&#8217;ll figure this out. But it&#8217;s gonna take a bunch of years. It&#8217;s clear from early <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6423358">evidence</a> that AI will be a tutor outside of class and inside class. It&#8217;ll do activities and give guidance. But schools are places where we can compel students to not use AI, and have them in a room, and evaluate them, and teach them the things that we want them to learn. As long as we think people need to be educated, this is the best space to do it in.<em> </em>So students are cheating in the meantime? They were cheating before! We can give them different tests; we could do in-class writing assignments. There can be a weird, backward-looking &#8220;Education won&#8217;t adjust!&#8221; view. How many death spirals does higher education need to be in per moment? There are the pieces to reconstruct a better form of education. It&#8217;s just a massive changeover.</p><h4><strong>BETTER SCIENCE &amp; BETTER THINKING</strong></h4><p><strong>Tom: What about academia? There&#8217;s been much talk about AI-written papers, and how they could overwhelm academic publishing. But could AI benefit the peer-review process, and help with the dissemination of academic findings?</strong></p><p><strong>Ethan:</strong> This is another area where more lifting is needed. It&#8217;s a shame that we are building AI co-scientists, but not thinking about the rest of the process that&#8217;s needed to actually make science happen. It&#8217;s one thing to have science produce more papers. We have no ability to absorb more papers. Every publication is overwhelmed. Our dissemination techniques were already bad, but now they&#8217;re really broken.</p><p><strong>Tom: As a case in point, you submitted a paper around 2023, and <a href="https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the-jagged">wrote</a> publicly about it then, making your term &#8220;the jagged frontier&#8221;&#8212;that AI capabilities advance in some areas but remain behind in others&#8212;highly influential. Yet the academic <a href="https://pubsonline.informs.org/doi/full/10.1287/orsc.2025.21838">paper</a> itself only just came out, three years later!</strong></p><p><strong>Ethan:</strong> One of the rejections we got early on was reviewers saying that they knew this already, and they cited a bunch of working papers&#8212;that cited the working paper <em>we</em> had submitted! This is not a unique story. Opening one part of the bottleneck without opening the others becomes a problem. But it takes longer to solve systemic problems of how science operates than to solve the problem of producing more papers.</p><p><strong>Tom: Another concern in education and science is <a href="https://www.mdpi.com/2075-4698/15/1/6">cognitive offloading</a>, that people may <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6097646">surrender</a> thinking to machines, and lose those skills. On the other hand, AI&#8217;s value comes from machines thinking for us. What are examples of </strong><em><strong>bad</strong></em><strong> offloading and </strong><em><strong>good</strong></em><strong> offloading?</strong></p><p><strong>Ethan: </strong>We offload all the time, right? But we also force people not to offload. You could offload all your mental math to calculators, but we force students to do some math by hand in an attempt to get them to learn stuff. And we can enforce those rules in school. In the world of work, we are not used to thinking about training, about what should be offloaded, and what shouldn&#8217;t be. We need to make decisions about this. So, Rolls-Royce still employs someone to <a href="https://www.youtube.com/watch?v=q9yqXPNHyMA">paint stripes</a> on a car by hand, and that&#8217;s an obvious pushback against deskilling in one area. But Ford doesn&#8217;t do the same thing. These are choices we get to make at an organizational level, depending on what we think is valuable.</p><h4><strong>ADAPTING TO CONSTANT CHANGE</strong></h4><p><strong>Tom: A point you&#8217;ve made to young people about the AI future is that they&#8217;ll need to be adaptable. When educators talk about teaching adaptability, it sometimes boils down to encouraging &#8220;creativity&#8221; and &#8220;critical thinking.&#8221; Another view is that you&#8217;re more likely to be adaptable by developing deep domain knowledge. For you, what does learning adaptability mean?</strong></p><p><strong>Ethan: </strong>Adaptability requires both deep domain knowledge <em>and</em> wide knowledge: T-shaped behaviour is probably the way to go. I feel like it&#8217;s a throwaway line: &#8220;Well, we&#8217;ll all be adaptable!&#8221; If we could teach that, that&#8217;d be amazing. People are more adaptable than we think, so part of this is that people will figure stuff out. But we can&#8217;t just throw up our hands, and say, &#8220;Be adaptable!&#8221; You need to have deep enough knowledge to go into a field. You need to have broad enough knowledge so that, as one piece of knowledge becomes less useful, you&#8217;re moving to the next one. And we need to help people be adaptable by building systems that get them in place inside an organization and able to shift roles. I sometimes worry that adaptability is a catch-all for &#8220;Don&#8217;t worry! It&#8217;ll be fine!&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bNwQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bNwQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 424w, https://substackcdn.com/image/fetch/$s_!bNwQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 848w, https://substackcdn.com/image/fetch/$s_!bNwQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 1272w, https://substackcdn.com/image/fetch/$s_!bNwQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bNwQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png" width="1456" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1526206,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/194500111?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bNwQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 424w, https://substackcdn.com/image/fetch/$s_!bNwQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 848w, https://substackcdn.com/image/fetch/$s_!bNwQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 1272w, https://substackcdn.com/image/fetch/$s_!bNwQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0029af43-cf4a-4465-ac85-08b368988fdf_1934x1073.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Tom: Another side is that not everybody will be equally adaptable. Could it be that the AI future favours certain circumstances and characteristics?</strong></p><p><strong>Ethan: </strong>A lot of these characteristics were already good characteristics to have. Does AI act as a multiplier of them? Does it disincentivize some people? We&#8217;re now past the edge of what we know. Ultimately, all of these questions come down to the same exact question, which is: How good does AI get, how fast? We need to articulate more clearly what we think that future looks like. Because you can&#8217;t say, &#8220;We&#8217;re going to build a superintelligent machine that&#8217;s better than all humans at every intellectual task&#8212;but let&#8217;s start thinking about adaptability!&#8221; Unless you mean, &#8220;Let&#8217;s adapt to UBI&#8221; [where everyone gets Universal Basic Income cash payments from the government]. And then, we should be spending a lot more time thinking about those issues. Not everyone in the labs believes this, and I find that the econ people believe it less. But you can&#8217;t have this message of, like, &#8220;All work will be obsolete!&#8221; and then have detailed, ticky-tacky conversations about what you should do in eighth grade. Because, by the time you enter the job market, there&#8217;s no jobs. So give me the pathway that you think <em>is</em> there, and that becomes the most important question to ask.</p><p><strong>Tom: Are there other important questions I didn&#8217;t ask?</strong></p><p><strong>Ethan: </strong>We need to start thinking about getting into fields, and understanding what the changes are&#8212;we need to get detailed. That is where the research is missing. Another large-scale econ picture about AGI isn&#8217;t as useful. General-purpose technology affects everything, so we need policymaking for everything, from power generation to accountants, and when does the government say it&#8217;s okay to do this. There&#8217;s just this assumption that if we do the macro stuff, everything will work out. I&#8217;d rather see a lot more micro stuff: a thousand flowers everywhere, trying to come up with different approaches.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/q-and-a-with-ethan-mollick?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/q-and-a-with-ethan-mollick?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[AI Agents Running the State]]></title><description><![CDATA[What could possibly go wrong?]]></description><link>https://www.aipolicyperspectives.com/p/ai-agents-running-the-state</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/ai-agents-running-the-state</guid><dc:creator><![CDATA[AI Policy Perspectives]]></dc:creator><pubDate>Wed, 15 Apr 2026 09:50:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CymV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CymV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CymV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 424w, https://substackcdn.com/image/fetch/$s_!CymV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 848w, https://substackcdn.com/image/fetch/$s_!CymV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!CymV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CymV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png" width="1456" height="795" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:8904267,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/194174723?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CymV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 424w, https://substackcdn.com/image/fetch/$s_!CymV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 848w, https://substackcdn.com/image/fetch/$s_!CymV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!CymV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6880abf1-3bd3-4856-970a-6fd26eb0157e_2814x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Waiting for an AI helper. (Credit: Gemini)</figcaption></figure></div><div class="callout-block" data-callout="true"><p><em>&#8220;Public services&#8221; include everything from teachers to the trash, from roadwork to  permission for a tree house. Much seems routine, but plenty is at stake. This makes politicians hesitant to risk an overhaul, leaving the system creaking and the paperwork mounting. </em></p><p><em>Last October, a provocative proposal emerged. <a href="https://agenticstate.org/">The Agentic State</a> conjured a vision of officialdom transformed, converting outdated procedures with a new system of AI helpers. This fledgling project offers both a blueprint and a promise of assistance to governments around the world.</em></p><p><em>But what if the vision were blind to how this could go awry? <a href="https://simoneparazzoli.me/">Simone Maria Parazzoli</a>, a co-author of the paper, and <a href="https://www.linkedin.com/in/omerhanbilgin/">Omer Bilgin</a> of <a href="http://www.deliberaide.com">deliberAIde</a> decided to critique their own ideas, seeking pitfalls in hopes of averting them.</em></p><p style="text-align: right;">&#8212;Tom Rachman, <em>AI Policy Perspectives</em></p></div><div><hr></div><h4><strong>By Simone Maria Parazzoli &amp; Omer Bilgin</strong></h4><p></p><p><strong>Amid the exhaustion of caring for a baby, new parents must deal with everything from bewildering sobs, to erratic feeding times, to the joys of changing a soiled newborn at 3 a.m. The last thing they need is paperwork.</strong></p><p>But what if, when coming home from the maternity ward that first day, they could awaken a government AI voice assistant, tell it the happy news, and hear the following response? &#8220;Congratulations! What&#8217;s the baby called?&#8221; The app would then take care of all the dreary admin, coordinating across agencies, registering the child, and setting in motion the services that this tiny new citizen should enjoy.</p><p>That is one example of how a future &#8220;agentic state&#8221; could simplify, speed up, and improve citizens&#8217; interactions with public services. To be clear, this does not yet exist. But projects like this one, <a href="https://oxfordinsights.com/insights/innovation-under-tough-circumstances-ukraines-ai-strategy-in-times-of-war/">envisioned</a> by Ukrainian officials, are more than fantasy, with several countries avidly testing early versions of agentic AI systems.</p><p>While Ukraine works toward the baby example, <a href="https://www.gov.uk/government/news/ai-helpers-could-coach-people-into-careers-and-help-them-move-home">Britain</a> is piloting agent-based support to provide citizens more tailored help. Meanwhile, <a href="https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/press-releases/2026/new-model-ai-governance-framework-for-agentic-ai">Singapore</a> is developing governance frameworks for agentic AI, and governments from <a href="https://guides.data.gouv.fr/intelligence-artificielle/le-serveur-mcp-de-data.gouv.fr">France</a> to the <a href="https://www.govinfo.gov/features/mcp-public-preview">United States</a> are ensuring that their public data can be accessed by agents.</p><p>Agentic AI systems&#8212;capable of perceiving, reasoning, and acting with minimal human supervision&#8212;will transform what organizations can achieve. By combining the reasoning of large language models with retrieval, memory, and tool use, agentic AI can automate complex tasks. For governments, whose core work is high-volume, structured administrative processes, this could make services more efficient, timely, consistent, and fair, while lowering costs.</p><p>Consider a citizen looking to start a small business. An agentic system&#8212;instead of requiring the entrepreneur to individually navigate zoning boards, tax authorities, and regulations&#8212;could autonomously reconcile these requirements. The larger promise is a shift from just <em>doing things right</em> (optimizing for procedure-following) to <em>doing the right things</em> (pursuing outcomes that citizens truly want).</p><p>The <a href="https://agenticstate.org/">Agentic State</a> vision paper&#8212;supported by The World Bank and the Global Government Technology Centre Berlin&#8212;was the first effort to systematically map the opportunities of agentic AI adoption for governments. This was not an academic exercise: 21 leaders across 15 countries contributed, including ministers and chief technology officers preparing to lead this transition.</p><p>In this vision, AI agents are a means to manage <em>complexity</em> and <em>scale</em>, while humans develop <em>strategy</em>, exercise <em>judgment</em>, and hold <em>accountability</em>.</p><p>Several governments have integrated official chatbots into their government services, but most of these merely provide conversational guides to administrative procedures. A few pioneering countries are starting to move beyond that. Ukraine, for instance, is turning chatbots into agentic assistants. Specifically, its Diia.AI assistant can retrieve users&#8217; data from connected registries, and generate official documents such as income certificates, while also providing certified information based on records such as taxation, land registries, and pensions.</p><p>The United Kingdom is also exploring agentic interactions via <a href="https://insidegovuk.blog.gov.uk/2025/12/16/gov-uk-has-entered-the-chat-our-vision-for-gov-uk-chat/">GOV.UK Chat</a> (inspired by Diia.AI), including a pilot program to support job seekers that transforms a static digital portal into an active assistant, matching users&#8217; skills with available opportunities.</p><p>Yet trends and optimism are not enough for success. The agentic state vision rests on key assumptions. What if they&#8217;re wrong?</p><p>This article presents a &#8220;red-teaming&#8221; exercise&#8212;a stress test of this vision&#8212;that identifies six core assumptions, along with scenarios that could emerge if they don&#8217;t hold true, and guardrails to avert such failures.</p><div><hr></div><div class="callout-block" data-callout="true"><h4><strong>Assumption 1: </strong><em><strong>AI Agents Become More Capable and Reliable</strong></em></h4></div><p>Agents can already perform rudimentary planning, tool use (e.g., searching the internet, using calculators, sending emails), and multistep task execution. Frontier labs are <a href="https://www.technologyreview.com/2025/01/11/1109909/anthropics-chief-scientist-on-5-ways-agents-will-be-even-better-in-2025/">betting</a> <a href="https://blog.samaltman.com/reflections">heavily</a> on agents, making it plausible that systems capable of managing complex and large-scale administrative tasks will emerge soon.</p><h4><strong>Failure Scenario: </strong><em><strong>The Technology Falters</strong></em></h4><p>Governments reorganize around agentic execution, but systems never become reliable enough for public administration. The demos look strong, but real cases fail on edge conditions, and require constant human correction. The agentic layer becomes only superficially competent with layers of human intervention underneath.</p><h4><strong>Guardrail: </strong><em><strong>Start Cautiously</strong></em></h4><p>Governments should start with minimal deployments, and tightly scoped use cases to validate reliability, develop procedural rigor and organizational competence, and account for technological evolution rather than committing prematurely to large-scale redesigns.</p><div class="callout-block" data-callout="true"><h4><strong>Assumption 2: </strong><em><strong>Agents Can Work Together</strong></em></h4></div><p>The success of agentic systems demands that they&#8217;re able to interact seamlessly, conveying intent, carrying out tasks, and sharing data in an interoperable way. <a href="https://modelcontextprotocol.io/docs/getting-started/intro">MCP</a> (model context protocol) is emerging as the technological standard for connecting AI applications with external systems. </p><h4><strong>Failure Scenario: </strong><em><strong>Standards Fail to Converge</strong></em><strong> </strong></h4><p>Commercial interests diverge, establishing competing protocols, while  government departments end up using AI systems that cannot communicate with one another. When a citizen&#8217;s request requires action from multiple agencies, the process breaks down. </p><h4><strong>Guardrail: </strong><em><strong>Officials Insist on Shared Protocols</strong></em></h4><p>Governments should make interoperability a condition of adoption, participating in the cross-sectoral <a href="https://aaif.io/">bodies</a> and forums where these standards are being shaped, funding the development of shared agentic interfaces and other agent-specific standards, and mandating non-proprietary protocols in procurement. <a href="https://www.aipolicyperspectives.com/p/the-past-and-future-of-ai-standards">Standards</a> rarely emerge by accident, but they may emerge when powerful governments treat them as a priority.</p><div class="callout-block" data-callout="true"><h4><strong>Assumption 3: </strong><em><strong>Organizations Will Adapt</strong></em></h4></div><p>To adopt and employ agents effectively, organizations must rethink their processes, roles, and incentives. They need to flexibly change and dynamically adapt practices to keep pace with the changing technological landscape.</p><h4><strong>Failure Scenario: </strong><em><strong>The Status Quo Prevents Change</strong></em></h4><p>Agentic AI adoption outpaces organizational change, with citizens and civil servants using agents in an uncoordinated manner long before official programs catch up. Local practices harden into path dependence before common standards emerge. The state becomes more productive at producing bureaucracy, not societally beneficial outcomes.</p><h4><strong>Guardrail: </strong><em><strong>Redesign Processes Before Automating Them</strong></em></h4><p>Agents should only enter workflows that have been simplified, decomposed, and restructured to minimize approval layers and handovers. Governments must treat adoption as a continuous discovery process. They should invest in common evaluation templates, reusable components, and a cross-agency repository of lessons, so that what works in one place can travel before what does <em>not</em> work becomes entrenched. </p><div class="callout-block" data-callout="true"><h4><strong>Assumption 4: </strong><em><strong>Private Adoption of Agentic AI Will Be Rapid</strong> </em></h4></div><p>Many companies are <a href="https://sloanreview.mit.edu/projects/the-emerging-agentic-enterprise-how-leaders-must-navigate-a-new-age-of-ai/">betting</a> on an agentic future. Firms are experimenting with internal copilots and autonomous customer flows, while frontier AI companies advance core models, architectures, and capabilities, and cloud providers offer the compute needed to deploy agents at scale. This suggests that agents will become commonplace across business, consumer, and enterprise environments, allowing governments to build on tools, infrastructure, and behaviors already spreading across the economy. This assumption rests on projections, though <a href="https://www.aipolicyperspectives.com/p/predicting-ais-impact-on-jobs">evidence</a> remains ambiguous.</p><h4><strong>Failure Scenario: </strong><em><strong>Diffusion Is Slower Than Forecast</strong></em></h4><p>Governments invest as if an agent-saturated economy is imminent, but industry adoption remains narrow, experimental, or ends up costing more than it saves. Public investments don&#8217;t plug into widely used tools and practices, meaning that citizens find agentic interfaces in government before they&#8217;re normal elsewhere. The state ends up bearing political and institutional costs without the stabilizing effects of private-sector diffusion.</p><h4><strong>Guardrail: </strong><em><strong>Lower Barriers to Private-Sector Agentic Usage</strong></em></h4><p>Governments can accelerate the development of an agentic AI ecosystem by investing in shared agentic infrastructure&#8212;such as standard ways to access public data, communicate across systems, and carry out authorized tasks and payments&#8212;that lower integration costs for firms, and reduce the risk of differing technological maturity across sectors.</p><div class="callout-block" data-callout="true"><h4><strong>Assumption 5: </strong><em><strong>Citizens Will Prefer Agentic Services</strong></em></h4></div><p>Increasingly, citizens are interacting with and relying on AI tools, but <a href="https://mbs.edu/-/media/PDF/Research/Trust_in_AI_Report.pdf?rev=0ee82285b2b0439bba524dbddc58214a">many do not trust the</a>m. For governments to integrate AI agents into workflows and services, citizens must accept and support the roles that agentic systems can play, finding them sufficiently trustworthy, reliable, fair, convenient, and accountable.</p><h4><strong>Failure Scenario: </strong><em><strong>The Public Rejects Automation</strong></em></h4><p>A single notable failure, or an accumulation of failures, turn the public against agentic systems, and convince many to opt-out. They judge automated decisions as opaque, illegitimate and untrustworthy, and suspect it worsens <a href="https://arxiv.org/abs/2510.16853">inequality</a>, with privileged citizens able to employ highly capable personal agents to navigate bureaucracy better than those relying on basic tools. The government is forced to run two systems&#8212;agentic and human&#8212;and neither meets expectations.</p><h4><strong>Guardrail: </strong><em><strong>Mandate Transparency</strong></em></h4><p>Governments must make agent integrations into government processes as legible as possible, furnishing explanations of decisions and publishing evaluation results for agentic fairness and performance, while detecting patterns of systemic bias or unequal benefit distribution based on citizens&#8217; technological access.</p><div class="callout-block" data-callout="true"><h4><strong>Assumption 6: </strong><em><strong>Human Oversight Will Evolve</strong></em></h4></div><p>For AI agents to act with functional autonomy within government processes, oversight frameworks <a href="https://arxiv.org/pdf/2506.04836">must adapt</a>, moving away from mandatory human reviews and approvals for everything (human-in-the-loop) to intermittent oversight (<a href="https://link.springer.com/rwe/10.1007/978-981-97-8440-0_75-1">human-on-the-loop</a>). This evolution increases speed and efficiency while reducing bottlenecks, with humans intervening only on edge cases. There is precedent for such adaptation: governments adapted regulation to cloud computing, e-identities, and AI-driven decision support systems.</p><h4><strong>Failure Scenario: </strong><em><strong>Regulation Never Updates</strong></em></h4><p>Every agentic action requires human verification; every decision must be justified through mechanisms designed for old chains of accountability. Agents can draft, but cannot act. Compliance and procedural costs rise as institutions retrofit old controls onto new AI processes. The result is high bureaucracy and low autonomy: an <em>agentic state</em> in theory, a <em>copilot state</em> in practice.</p><h4><strong>Guardrail: </strong><em><strong>Sandboxes to Test Oversight</strong></em></h4><p>Governments should establish controlled environments that allow policymakers, developers, and civil society to collaborate and gather empirical evidence on what forms of oversight are adequate and best fit different kinds of agentic deployments, reducing uncertainty before codifying rules at scale. They should explore this early, much as Singapore has done through its <a href="https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf">Model AI Governance Framework for Agentic AI</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p><strong>Soon, agentic government will be more than optimism and testing.</strong> A vanguard of countries will implement these tools. If those cases produce the kinds of benefits imagined, other countries will flock to join them. </p><p>But momentum is not inevitability. This project depends on assumptions&#8212;about progress, coordination, institutions, norms, and law&#8212;that demand scrutiny before governments rebuild themselves around these new technologies. </p><p>This red-teaming exercise of the agentic state concept is not to argue against the vision, but to make it more robust and resilient. The six possible failure scenarios are not mutually exclusive. Several could compound, and some may already be taking shape. For instance, reliability has been improving <a href="https://arxiv.org/html/2602.16666v1">much more slowly</a> than accuracy, providing ground for technology to falter (Scenario 1), and there are <a href="https://www.adalovelaceinstitute.org/policy-briefing/great-expectations/">signals</a> that the public might reject automation if economic gains and innovation speed are prioritized over fairness (Scenario 5). </p><p>Governments that are serious about improving the state with AI must attend to these risks in earnest now, while the architecture is still being laid. The opportunity is too precious to spurn. </p><p>Agentic AI could make public services considerably faster, fairer, and more responsive&#8212;more so than anything the traditional bureaucratic model has yet delivered. That prize is worth the discipline of preparing for what could go wrong.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-agents-running-the-state?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/ai-agents-running-the-state?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><em>For further details on &#8220;The Agentic State,&#8221; check out the original <a href="https://agenticstate.org/paper.html">vision paper</a></em> </p>]]></content:encoded></item><item><title><![CDATA[AI Policy Primer (#24)]]></title><description><![CDATA[Identifying agents, self-improvement, and artificial clouds]]></description><link>https://www.aipolicyperspectives.com/p/ai-policy-primer-24</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/ai-policy-primer-24</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Thu, 09 Apr 2026 14:50:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CLkj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every six weeks, we round up three papers that we think AI policy folks should be reading. In this edition, we look at a <a href="https://arxiv.org/abs/2603.10028">proposal</a> for how to identify the agents that will soon fill the economy; <a href="https://cset.georgetown.edu/publication/when-ai-builds-ai/">research</a> on the prospect of self-improving AI; and<a href="https://arxiv.org/pdf/2603.06909"> new insights</a> about how to use AI to prevent contrails, or artificial clouds, from warming the planet. </em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CLkj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CLkj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 424w, https://substackcdn.com/image/fetch/$s_!CLkj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 848w, https://substackcdn.com/image/fetch/$s_!CLkj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 1272w, https://substackcdn.com/image/fetch/$s_!CLkj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CLkj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2365898,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.aipolicyperspectives.com/i/193691288?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CLkj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 424w, https://substackcdn.com/image/fetch/$s_!CLkj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 848w, https://substackcdn.com/image/fetch/$s_!CLkj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 1272w, https://substackcdn.com/image/fetch/$s_!CLkj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F879e6456-0c9d-4813-a7b9-1fdc297b6a23_8000x4500.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>1. Identifying (and incentivising) AI agents</h2><ul><li><p><strong>What happened: </strong>A trio of law and philosophy professors considered how to identify who (or what) is responsible for AI agents&#8217; actions in the world, and came up with a two-part <a href="https://arxiv.org/abs/2603.10028">proposal</a>: that the disparate and evolving agents within a system should exist legally as a new form of corporation; and that each corporation should link to accountable humans.</p></li><li><p><strong>What&#8217;s interesting: </strong>The paper by <a href="https://law.ua.edu/faculty_staff/yonathan-arbel/">Yonathan Arbel,</a> <a href="https://law.ua.edu/faculty_staff/yonathan-arbel/">Simon Goldstein</a>, and <a href="https://www.law.uh.edu/faculty/main.asp?PID=6428">Peter N. Salib</a> starts with a thought experiment. It&#8217;s 2030, and your AI assistant suggests that it optimizes your slow WiFi connection. After you agree, it spawns a swarm of agents. Some are copies, while others are cheaper agents running on open-source models. Some start to interface with AI agents from other companies. Three months later, two FBI agents knock on your door and explain that your network has been piggybacking on a local defense contractor&#8217;s WiFi network.</p></li><li><p>Before determining who is responsible and what the repercussions should be, there are more basic questions: Who are the AI actors in this story? How many are there?</p></li><li><p><a href="https://www.aipolicyperspectives.com/p/an-agents-economy">The economy will soon be filled with capable AI agents</a>. To deter and respond to such harms, the authors argue that we need to be able to identify these agents, at two levels.</p><ul><li><p>To prevent human misuse or negligence, we need &#8216;<strong>thin identity&#8217;</strong>. This would connect AI agents to the humans most able to control them, similar to how &#8216;know-your-customer&#8217; rules tie banking transactions to humans.</p></li><li><p>Humans will be unable to monitor and control every AI decision, so we also need to be able to identify agents themselves, hold them accountable and incentivize them to behave well. To do so, we need &#8216;<strong>thick identity&#8217; </strong>that can distinguish AI agents as stable, coherent entities, with persistent goals. This goal is pragmatic and does not require viewing AIs as conscious in any sense.</p></li></ul></li><li><p><em>Thickly </em>identifying agents is harder and more novel, as AI agents need not be attached to a physical body. Multiple agents can also work together on a single task. Any single agent can be copied, spun up, spun down, or be continually updated.</p></li><li><p>To address such challenges, the authors propose creating algorithmic corporations, or &#8216;A-corps&#8217;. These would have two key elements:</p><ul><li><p><strong>Legal personhood: </strong>Like a traditional corporation, an A-corp would be a single legal entity that persists over time. It could hold property, make contracts, and be sued. But it would be run by a collection of AI agents. As such, the proposal runs contrary to scholars who have argued <a href="https://arxiv.org/pdf/2502.18359">against</a> granting legal personhood to AI agents, or called for <a href="https://openscholarship.wustl.edu/law_lawreview/vol95/iss4/7/#:~:text=This%20Article%20argues%20that%20algorithmic,which%20have%20non%2Dhuman%20controllers.">bans</a> on algorithms running companies because of concerns about crime and companies using them to avoid liability.</p></li><li><p><strong>Computationally-secure governance: </strong>Each A-corp would have a unique digital certificate and a secure private key to authorise transactions. The humans that own each A-corp could grant the key to an AI &#8216;manager&#8217; agent who in turn could grant more limited permissions to sub-agents within the A-corp, or to other A-corps, such as permissions to spend up to $100 or to read a batch of emails.</p></li></ul></li><li><p>The proposal addresses thin identity by reducing the vast number of AI agents down to a smaller number of A-corps, whose actions are traceable back to their human owners. As with limited liability companies (LLCs), the human owners would not be responsible for <em>all </em>harm their A-corps cause, but could lose all funds they invest and possibly face further liability, for example in cases of fraud or negligence.</p></li><li><p>The proposal addresses thick identity via its &#8216;resource constraint thesis&#8217;. All AI agents need resources, like money and compute. A-corps provide AIs with a way to access these resources and an incentive to manage them well. For example, A-corps that tightly monitor and audit their sub-agents&#8217; performance would get more resources, while A-corps that allow fraud or waste will lose resources. This encourages A-corps to self-organise<em>, </em>into stable, coherent, multi-agent systems.</p></li><li><p>The authors argue that A-corps could also address alignment concerns, for example by reducing the incentive for an AI agent to exfiltrate its own weights, because that new AI instance would lose access to resources and permissions from the A-corp.</p></li><li><p>To make it happen, the authors call for a public registry of A-corps. This would list each A-corp&#8217;s human owners, the certificates to authenticate it against, as well as (potentially) the differing permissions enjoyed by its agents. Ultimately, the authors argue that A-corps should become mandatory for any AI agent taking &#8220;economically significant actions&#8221;, and to guard against criminals using AI agents anonymously.</p></li><li><p>The authors respond to some expected pushback. They do not see A-corps as anthropomorphising AI because the proposal does not require anybody to view agents as having deeper desires or wants. They also think A-corps can prevent the risk that AI agents might slowly build up resources before deploying them for harm, by encouraging inter-agent trade that penalises rogue behavior. Could A-corps disempower humans? The authors argue that they provide a pathway to tax and redistribution, and enable humans to better steer agents, for example by designating the parts of the economy that A-corps are permitted to operate in.</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free. Lots more in the pipeline. </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>2. When AI builds AI</h2><ul><li><p><strong>What happened: </strong>The Centre for Security and Emerging Technology, CSET, released <a href="https://cset.georgetown.edu/publication/when-ai-builds-ai/">a report</a> on the prospects for AI improving itself, known as automated R&amp;D or recursive self-improvement, based on an expert workshop in July 2025.</p></li><li><p><strong>What&#8217;s interesting: </strong>In 1964, the computer scientist I.J. Good wrote about the possibility of an &#8220;intelligence explosion&#8221; that would leave &#8220;the intelligence of man.&#8230;far behind&#8221;. Researchers have also long automated aspects of writing code and AI model design.</p></li><li><p>However, the speed of AI coding advances suggests that something qualitatively different may soon occur. This makes two questions salient: 1. Could AI automate the <em>entire </em>AI R&amp;D process? 2. Will this R&amp;D automation extend across all scientific disciplines? The CSET report focuses on the first question.</p></li><li><p>CSET defines AI R&amp;D by distinguishing between <em>research scientists, </em>who generate hypotheses, design experiments and interpret results; and <em>research engineers, </em>who write code, fix bugs and generate data. They also note the inputs that AI R&amp;D relies on, such as raising funds and acquiring compute.</p></li><li><p>They sketch out four overlapping scenarios for how AI R&amp;D may play out:</p><ul><li><p><strong>1. Explosion: </strong>AI systems automate a growing share of AI R&amp;D. Initially, this leads to modest productivity gains, but as the length and complexity of tasks that AI performs grows, productivity soars. AI systems become far more capable than humans, whose involvement in AI R&amp;D falls to zero.</p></li><li><p><strong>2. Fizzle: </strong>The share of R&amp;D tasks done by AI rises, but rather than leading to compounding improvements, capabilities start to plateau.</p></li><li><p><strong>3. Amdahl&#8217;s Law: </strong>AI automates certain activities, like writing code and running experiments, but not others, like research strategy.</p></li><li><p><strong>4. The expanding pie: </strong>As AI automation grows, humans realise that new ideas and breakthroughs are needed that AI systems cannot yet provide.</p></li></ul></li><li><p>The experts in CSET&#8217;s workshop held widely diverging views on which scenario was most likely. Most importantly, new empirical data is unlikely to resolve these conflicts, because participants may view the same data as confirming their own assumptions.</p><ul><li><p>For example, an AI system&#8217;s inability to reliably use a keyboard or mouse may look like a bottleneck to one expert, but a source of explosive growth to another&#8212;if they expect this human-focussed tooling to get adapted for the AI era. Similarly, different experts may view AI automating a growing share of R&amp;D tasks as progress towards a fast takeoff, or as low-hanging fruit being picked off, accelerating progress only as far as the upcoming wall.</p></li></ul></li><li><p>These differing views are also visible in more recent commentary on the topic.</p><ul><li><p>The prominent AI researcher and writer Nathan Lambert recently <a href="https://www.interconnects.ai/p/lossy-self-improvement">cited</a> Paul Allen&#8217;s concept of a &#8216;complexity brake&#8217; to argue that as we understand intelligence better, further progress becomes exponentially harder. In addition to incurring financial costs, Lambert argued that running suites of AI agents won&#8217;t necessarily lead to exponential progress, because those agents will perform best on narrow, verifiable tasks, will be hard to manage in large numbers, and will sample from similar parts of the distribution of AI research ideas, inhibiting more novel breakthroughs.</p></li><li><p>Conversely, Ajeya Cotra at METR, the Model Evaluation and Threat Research organisation, recently wrote about how she &#8220;<a href="https://www.planned-obsolescence.org/p/i-underestimated-ai-capabilities?utm_source=substack&amp;utm_medium=email">underestimated AI capabilities (again)</a>&#8221;.  She argued that AIs may, counterintuitively, find it easier to decompose <a href="https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/">longer projects</a> into sub-components that multiple agents can run in parallel, than for shorter tasks. AIs will also produce good documentation for their fellow AIs, which could accelerate progress.</p></li></ul></li><li><p>If faster automation and progress does occur, the CSET authors see two main risks: Less time to prepare for safety risks from AI, and lower human understanding of AI systems. To address these risks, their recommendations have a strong focus on improving access to evidence, including:</p><ul><li><p><strong>New evaluations of AI R&amp;D, </strong>including for &#8216;<a href="https://arxiv.org/pdf/2503.14499">messy</a>&#8217; tasks such as research strategy, which lack clear specifications and success criteria and take place in a dynamic environment with various real-world interactions.</p></li><li><p><strong>New approaches to evaluation</strong> to better distinguish &#8216;degrees of accomplishment&#8217; from a simple success/failure binary.</p></li><li><p><strong>Better insights into how automated R&amp;D is progressing within AI labs,</strong> such as data on how funding is allocated and qualitative impressions of progress from leading AI researchers and engineers.</p><p></p></li></ul></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free. Lots more in the pipeline. </p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>3. Planes and global warming</h2><ul><li><p><strong>What happened: </strong>A team of researchers, including from Google and American Airlines, published <a href="https://arxiv.org/pdf/2603.06909">results</a> from their latest experiment to use AI to reduce condensation trails from planes&#8212;a key contributor to global warming.</p></li><li><p><strong>What&#8217;s interesting: </strong>When pilots fly, particles from the plane&#8217;s exhaust can mix with low-pressure air to form <em>contrails</em>&#8212;white, artificial clouds, made up of ice crystals. These contrails are a net contributor to global warming, because they trap heat that would otherwise escape. Debates continue over exactly how much they contribute, but one <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC7468346/">estimate</a> suggests that they contribute a lot, causing around 2% of &#8216;radiative forcing&#8217;, which measures how different factors, like CO<sup>2</sup>, heat or cool the planet.</p></li><li><p>As the environmental writer Hannah Ritchie <a href="https://hannahritchie.substack.com/p/contrails-google-ai">explains</a>, more important than the absolute figure is the fact that contrails offer a rare opportunity to reduce global warming almost immediately, at relatively low cost. This is because a small share of flights cause most of the warming-inducing contrails&#8212;generally those that fly through parts of the atmosphere that are both very cold and very humid. If planes take short detours to avoid these patches of air, contrails (and warming) should drop.</p></li><li><p>A few years ago, Google researchers<a href="https://blog.google/innovation-and-ai/technology/ai/ai-airlines-contrails-climate-change/"> partnered with</a> American Airlines on a proof of concept. Using satellite imagery and AI, they were able to predict where contrails would emerge and guide planes to avoid them, reducing contrails by &gt;50%, across 70 test flights.</p></li><li><p>In the latest <a href="https://arxiv.org/pdf/2603.06909">study,</a> they expanded the experiment to 2,400 American Airlines flights from the US to Europe. They placed ~50% of planes in a treatment group, where flight dispatchers were given two choices: a standard flight plan and an alternative contrail-avoidance one. Their decision for which to recommend was voluntary.</p></li><li><p>For flights in this intervention group, contrails fell by 12% compared to a control group with no contrail-avoidance plan. Importantly, the contrail-avoidance routes also did not lead to a significant increase in fuel use. At first glance, these results seem positive, but modest. Digging into the results highlights the challenge of getting useful AI deployed at scale.</p></li><li><p>In particular, dispatchers who received contrail avoidance plans only recommended them to pilots 15% of the time. Even then, the avoidance plan was only <em>successfully</em> flown in 60% of flights. For planes that did successfully follow the avoidance plan, contrails fell by more than 60%, a much larger reduction. So the tech worked, but was often not used.</p></li><li><p>Why? Dispatchers are busy<strong> </strong>and must often deal with other priorities, like bad weather and turbulence. To avoid contrails, planes also need to climb and descend mid-flight. This is safe, but creates more work for pilots and air traffic controllers. As it was voluntary, the incentive to change to a contrail-avoidance plan was weak.</p></li><li><p>The way that the dispatchers received the information also meant that they didn&#8217;t fully understand <em>why</em> the suggested up and down changes were necessary. Happily, the authors feel that most of these obstacles are addressable, with a combination of a better user interface, some automation, and more incentives.</p></li><li><p>In addition to its immediate usefulness, the study is a rare real-world attempt to quantify the benefits of AI to tackling global warming. At the moment, the AI and climate change policy discussion is often negative and focuses on the emissions that may result from building and operating data centres (and other devices) to train and run AI models. This is important, but there are reasons to think that these emissions will be <a href="https://blog.andymasley.com/p/individual-ai-use-is-not-bad-for?open=false#%C2%A7emissions">relatively low</a>, or at least lower than many assume. In contrast, AI could potentially reduce emissions and warming by far larger amounts, for example by accelerating research on solar and fusion power, or making buildings and energy grids more efficient. But these benefits are typically more speculative, harder to quantify, or in the case of contrails, more <em>contingent </em>on human behaviour.</p></li><li><p>This experiment demonstrates that the benefits of AI to tackling global warming are real, but also points to the interventions that will be needed to push them to their full potential.  The study is also timely, given that governments <a href="https://assets.publishing.service.gov.uk/media/69b83baacf4af9cad362b4e7/jet-zero-taskforce-contrail-impact-mitigation-task-and-finish-group-a-strategic-framework-for-uk-contrail-impact-mitigation.pdf">are focussing</a> on contrail avoidance and some policy action may be required, for example to help standardise and mandate contrail prediction software or to generate high-resolution humidity data.</p><p></p></li></ul><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/ai-policy-primer-24?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/ai-policy-primer-24?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p>]]></content:encoded></item><item><title><![CDATA[How Tech Changed Chess]]></title><description><![CDATA[And why AI won&#8217;t end our games]]></description><link>https://www.aipolicyperspectives.com/p/how-tech-changed-chess</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/how-tech-changed-chess</guid><dc:creator><![CDATA[AI Policy Perspectives]]></dc:creator><pubDate>Wed, 25 Mar 2026 10:22:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CdTX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CdTX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CdTX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 424w, https://substackcdn.com/image/fetch/$s_!CdTX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 848w, https://substackcdn.com/image/fetch/$s_!CdTX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 1272w, https://substackcdn.com/image/fetch/$s_!CdTX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CdTX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png" width="1024" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f42add45-d564-4031-a602-e342e4b5c090_1024x572.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CdTX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 424w, https://substackcdn.com/image/fetch/$s_!CdTX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 848w, https://substackcdn.com/image/fetch/$s_!CdTX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 1272w, https://substackcdn.com/image/fetch/$s_!CdTX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff42add45-d564-4031-a602-e342e4b5c090_1024x572.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Credit: Gemini</figcaption></figure></div><p><em>From childhood upwards, we play games as a safe (and strangely joyful) way to battle, strategize, even lose without it coming to fisticuffs. Artificial intelligence grew up playing games too, with developers using the structured rules, scoring systems, and win/loss outcomes to train machines to learn, to improve, even to beat us.</em></p><p><em>In chess, bots have been bettering humans for years now. Yet our &#8220;loser&#8221; species still gathers at sunny park tables, in dank school gyms, and online in droves, all in hopes of crying, &#8220;Checkmate!&#8221; The resilience of chess is commonly cited as evidence that&#8212;even if AI surpasses us in various pursuits&#8212;humans won&#8217;t just give up.</em></p><p><em>However, there&#8217;s more to say about the intersection of technology and chess, in particular how the game has evolved with technology, including AI. Thankfully, the broadcaster and writer <a href="https://www.aipolicyperspectives.com/p/whats-it-like-to-be-a-bot">David Edmonds</a>&#8212;co-author of </em>Bobby Fischer Goes to War <em>(2004) and editor of the essay collection </em>AI Morality<em> (2024)&#8212;has spent decades observing this, both as a spectator and behind the board himself.</em></p><p style="text-align: right;"><em>&#8212;Tom Rachman, </em>AI Policy Perspectives</p><div><hr></div><p><strong>By DAVID EDMONDS</strong></p><p><strong>Among thousands of tournament games cited in the Batsford book of chess openings, tucked into the top right-hand column of Page 235, is an example of how white should </strong><em><strong>not </strong></em><strong>play.</strong></p><p>Explaining the Closed Sicilian Defense opening, the authors (former world champion Garry Kasparov and the British grandmaster and chess columnist Raymond Keene) spotlight a game in which black is already ahead as early as move 11. Indeed, the player with the white pieces ended up losing. I remember because that player was me.</p><p>That is my humiliating contribution to chess theory: what not to do. The book was published in 1982, and I&#8217;ve barely picked up a pawn in anger in the intervening four decades. But I still follow the chess world, and if there&#8217;s a tournament in London, I&#8217;ll go to watch, spending hours absorbed in the intricacies of the 64 squares.</p><p>As the digital revolution and AI juggernaut move through our lives, we may wonder whether there will still be domains in which humans can continue to find enjoyment and meaning. Chess offers a hopeful case study.</p><p>Chess and AI have had a long relationship. The great forefather of artificial intelligence Alan Turing wrote the <a href="https://www.chess.com/blog/the_real_greco/the-original-chess-engine-alan-turings-turochamp">first chess algorithm</a> in 1948. The following year, another seminal figure, Claude Shannon, distinguished two ways that a computer could play chess: by brute force, calculating every possible move; or by selective search, like a human.</p><p>Chess also proved a <a href="https://www.researchgate.net/publication/224834166_Is_chess_the_drosophila_artificial_intelligence_A_social_history_of_an_algorithm">favourite way</a> to evaluate AI advancement, both because many key innovators were keen players but also because the game&#8217;s mathematical structure and its win/loss conditions created benchmarks for comparing machine progress to human performance.</p><p>A longstanding goal&#8212;seemingly impossible at first&#8212;was to outclass the best humans in a game that has near-infinite <a href="https://en.wikipedia.org/wiki/Shannon_number">permutations</a>. Defeating humans at chess became the programmers&#8217; ultimate challenge, like runners seeking to break the four-minute mile or climbers reaching the summit of Mount Everest, both of which proved easier. Finally, in 1997, IBM&#8217;s Deep Blue vanquished Kasparov, the then-reigning world champion. A dejected Kasparov insinuated that there had been human intervention.</p><p>For a while, chess players comforted themselves with the thought that a hybrid combination of human and machine could outwit machine alone. That period has long passed. Today&#8217;s best player, Magnus Carlsen, would be trounced were he to compete in a series of games with my mobile phone.</p><p>In 2017, DeepMind&#8217;s <a href="https://deepmind.google/blog/alphazero-shedding-new-light-on-chess-shogi-and-go/">AlphaZero</a> took machine chess to the next level. While Deep Blue had relied on brute strength with some input from strong humans, AlphaZero was simply programmed with the basic rules, and then trained itself through reinforcement learning. In its learning phase, it played tens of millions of games against itself in just a few hours, then crushed the chess engine Stockfish. (Stockfish adapted its methods accordingly, and is now the leading chess engine.)</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p>World chess champions of the past exuded an aura. Their talents seemed mysterious, supernatural.  In part, that&#8217;s because few people, then and now, can comprehend the depth of thought that elite players achieve at the board. When it comes to music, we may never compose like Mahler, but we can appreciate Mahler&#8217;s symphonies. By contrast, we can neither play like Magnus Carlsen nor fully appreciate his games. It&#8217;s for this reason that the Armenian-born grandmaster, Lev Aronian, once <a href="https://www.prospectmagazine.co.uk/essays/53494/the-lion-and-the-tiger">confessed to me</a> that being one of the world&#8217;s top players was desperately lonely.</p><p>Carlsen has achieved the highest rating of any human in history. And, no surprise, he strikes a confident pose. Yet his strut no longer carries complete conviction. To spectators armed with portable chess engines, the chess gods have been humbled.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JeWo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JeWo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 424w, https://substackcdn.com/image/fetch/$s_!JeWo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 848w, https://substackcdn.com/image/fetch/$s_!JeWo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 1272w, https://substackcdn.com/image/fetch/$s_!JeWo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JeWo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png" width="728" height="555" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:555,&quot;width&quot;:728,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JeWo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 424w, https://substackcdn.com/image/fetch/$s_!JeWo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 848w, https://substackcdn.com/image/fetch/$s_!JeWo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 1272w, https://substackcdn.com/image/fetch/$s_!JeWo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81aa1b0b-fd78-4ab5-abda-a14f2c0fd43d_728x555.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The chess prodigy Samuel Reshevsky playing simultaneous games in 1920 against a variety of whiskery Parisians. Aged 8, he beat them all. (Credit: Creative Commons)</figcaption></figure></div><p>Even so, chess has not dwindled in popularity. On the contrary, more people are playing it than ever. The game received a boost during Covid, when we all hunkered down in our homes, connected by the Internet. Another boost came from the hit Netflix drama, <em><a href="https://en.wikipedia.org/wiki/The_Queen%27s_Gambit_(miniseries)">The Queen&#8217;s Gambit</a></em>. Meanwhile, a younger generation of telegenic chess masters has gained avid YouTube followings, turning commentary and stunts into short-clip entertainment.</p><p>Here are 11 ways that technology has changed chess. The 11th is the most interesting:</p><ol><li><p><strong>Opening Preparation.</strong> The systematic study of chess openings goes back a couple of centuries or more. Sequences of opening moves were mapped out&#8212;as in that 1982 book that included my embarrassing loss. But chess engines allow for a depth of opening analysis that was inconceivable in 1982. This means that 25 moves may pass before grandmasters find themselves in unfamiliar territory nowadays. Some openings have also been resurrected because engines have shown the positions to be more survivable than previously recognized.</p></li><li><p><strong>Opponent Preparation</strong>. Even in amateur tournaments, players routinely prepare for opponents in an individually tailored way. This is made possible because the games of each opponent are available online.</p></li><li><p><strong>Connectivity.</strong> Fancy a game? There are endless online adversaries willing to take you on, day and night, from India to Iceland, Cape Town to Chicago.</p></li><li><p><strong>No More Correspondence Chess.</strong> There was once a thriving chess scene in which games were played remotely over a long time period&#8212;months, sometimes years&#8212;with moves typically sent by post. How quaint.</p></li><li><p><strong>No More Adjournments</strong>. Historically, world championship games would sometimes stop after five hours to resume later. That can&#8217;t happen anymore, since players might simply identify the optimal continuation with the help of an engine. Time limits now ensure games finish within a single session.</p></li><li><p><strong>Shorter Games</strong>. Many in the online chess audience don&#8217;t have patience for lengthy games. For them, quicker time controls&#8212;Rapid (less than an hour); Blitz (3-5 minutes); or Bullet (under 3 minutes)&#8212;are more thrilling.</p></li><li><p><strong>Different Formats.</strong> Now that computers have shown with such depth which opening sequences are optimal, the early part of a game has been transformed into a feat of memory rather than creativity. As a result, Fischer Random (advocated early on by the ex-American world champion Bobby Fischer) has become increasingly popular. In Fischer Random, the starting position of the major pieces behind the pawns is randomized, making opening homework effectively impossible. It&#8217;s sometimes called Freestyle Chess, or Chess960 because there are 960 possible ways for the pieces to be shuffled.</p></li><li><p><strong>Job Generation.</strong> With a potential global audience, some players can now earn a decent living live-streaming their games, or offering online training.</p></li><li><p><strong>Roasting of Champions</strong>. This is an irksome development. Since chess engines assign an instant numerical evaluation of the position after each move (e.g. +1 means white is better by roughly one pawn), any patzer can see when a grandmaster has blundered, and is free to abuse them in online comments.</p></li><li><p><strong>Cheating</strong>. There have always been cheating accusations in chess. In 1978, the Soviet dissident Viktor Korchnoi claimed that the aides of his opponent, Anatoly Karpov, were using the flavour of the <a href="https://www.bbc.co.uk/sounds/play/w3cszmwf">yogurt</a> handed to Karpov to secretly convey messages. More recently, suspicion (tongue-in-cheek, but taken seriously by online trolls) has been raised of illicit advice being transmitted via <a href="https://www.bbc.co.uk/news/world-us-canada-66921563">vibrating sex toys</a>. In elite tournaments, grandmasters are now searched before they enter the playing arena, even accompanied to the toilet. Spectators, meanwhile, are prohibited from carrying phones, to prevent them signalling the best continuation. But in online games, cheating is almost impossible to prevent. Platforms try to detect cheats by comparing human moves to the recommendations of top engines. But if savvy cheaters consult an engine just once or twice in a game, they may win without being detected.</p></li></ol><p>And so to <strong>the 11th effect on chess: the expansion of human imagination</strong>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/how-tech-changed-chess?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/how-tech-changed-chess?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>In the last few years, there has been a slight but detectable shift in grandmaster play, as humans learn from machines, both through gameplay against bots and by using machine insights to prepare for human competition.</p><p>People who don&#8217;t play chess may imagine that what distinguishes strong from weak players is calculating power. And it&#8217;s true that top grandmasters can analyse many moves in advance. But their edge is tougher to articulate. It involves superior pattern recognition, with an intuitive sense for where their pieces should be placed and how a position should advance. Likewise, Mozart <em>felt</em> how a composition ought to develop; his instincts about building tension and creating contrasts were the product in part of having internalized countless musical patterns.</p><p>For chess players, some moves seem ugly. It might feel wrong to shunt a knight to the edge of the board, to break up a pawn structure, or to expose the king. But computers don&#8217;t <em>feel</em> anything. In chess, they care about patterns and the interplay between pieces only to the extent that they&#8217;re relevant to the ultimate objective: victory.</p><p>However, bots don&#8217;t necessarily play robotically. They produce moves that astonish and inspire human players, even make them <a href="https://youtu.be/CdFLEfRr3Qk?t=199">laugh</a> with surprise. One famous case of AI invention across the board came in another game, Go, when the <a href="https://deepmind.google/research/alphago/">AlphaGo</a> program was facing a top human player, and produced a move that caused professionals to gasp. &#8220;Move 37&#8221; is still cited with awe, as something a person would never have done, but that worked sublimely.</p><p>Likewise, chess engines regularly expand the imagination of human chess players, pushing beyond the habitual &#8220;correct&#8221; move they&#8217;ve seen many times before or have learned from books of chess theory. AI has even <a href="https://arxiv.org/abs/2510.23772">dabbled</a> in the art form of creating beautiful chess puzzles. And empirical studies <a href="https://www.pnas.org/doi/10.1073/pnas.2406675122">indicate</a> that leading players may pick up new ideas and strategies from machines.</p><p>Machines, in other words, can make humans more resourceful and inventive, breaking down rigid modes of thinking. The implausible becomes plausible. The readily dismissed becomes the carefully considered. This evolution of chess illustrates a broader idea in the development of AI that may prove immensely valuable in science and elsewhere in human endeavour: that how AIs think may help human experts learn <a href="https://arxiv.org/pdf/2502.07586">new ideas</a> themselves.</p><p>In his book <em>The Silicon Road to Chess Improvement</em>, the grandmaster Matthew Sadler argues that chess engines can improve every player, and he documents some of the counterintuitive patterns that humans could pick up from AI. By way of illustration, during a top tournament this January, the Indian grandmaster Arjun Erigaisi (playing against Vladimir Fedoseev of Russia) advanced his pawns in a way that looked reckless. In fact, computer analysis indicated he was still ahead after 28 moves. However, he blundered and lost. The danger of learning from a computer is that success may require you to proceed with computer-level accuracy.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p>As AI undertakes more activities formerly done only by people, it&#8217;s worth asking why human chess persists&#8212;and will likely continue to do so.</p><p>A Canadian philosopher, Bernard Suits, pointed out in his 1978 book <em><a href="https://books.google.co.uk/books/about/The_Grasshopper.html?id=1LmESO3NBuoC&amp;redir_esc=y">The Grasshopper: Games, Life and Utopia</a></em> that what defines &#8220;games&#8221; is that they involve the voluntary attempt to overcome unnecessary obstacles. Therein lies a defence against AI encroachment. In a market economy, companies aim to remove or overcome obstacles in the pursuit of profit. In games, obstacles have been deliberately inserted as an indispensable feature. What we enjoy in playing chess is testing our cognitive abilities. What we enjoy in watching chess is two humans pitting their wits against each other in a socially constructed activity where difficulty enhances enjoyment and satisfaction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b-Lb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b-Lb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 424w, https://substackcdn.com/image/fetch/$s_!b-Lb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 848w, https://substackcdn.com/image/fetch/$s_!b-Lb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 1272w, https://substackcdn.com/image/fetch/$s_!b-Lb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b-Lb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png" width="1200" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!b-Lb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 424w, https://substackcdn.com/image/fetch/$s_!b-Lb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 848w, https://substackcdn.com/image/fetch/$s_!b-Lb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 1272w, https://substackcdn.com/image/fetch/$s_!b-Lb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3aad9998-ae91-49a8-8154-a80018c3408d_1200x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Watching the big game. (Credit: Creative Commons)</figcaption></figure></div><p>There&#8217;s also a narrative element to caring about games. The contest&#8212;whether intellectual or physical&#8212;is absorbing precisely because it involves conscious creatures. In elite chess, there&#8217;s the backstory: the players&#8217; rise, their subsequent ups and downs, their history with specific opponents.</p><p>But watch an engine-against-engine tournament like TCEC (the <a href="https://en.wikipedia.org/wiki/Top_Chess_Engine_Championship">Top Chess Engine Championship</a>), and you&#8217;ll soon fall asleep. Computers aren&#8217;t competing after a divorce, or an illness, or the loss of a parent. Humans have character traits that spill onto the board, such as aggression (or passivity); patience (or impatience); equanimity (or volatility); and resilience (or fragility). Winning and losing have emotional resonance for a human&#8212;but not for AlphaZero.</p><p>It&#8217;s these qualities that guard against AI advance. AI might gobble up some of our jobs; even human-authored articles like this one may become rarer. But AI won&#8217;t take our chess.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/how-tech-changed-chess?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">It&#8217;s your move. Send this article to someone. </p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/how-tech-changed-chess?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/how-tech-changed-chess?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p></p>]]></content:encoded></item><item><title><![CDATA[The Past and Future of AI Standards]]></title><description><![CDATA[Lessons from history]]></description><link>https://www.aipolicyperspectives.com/p/the-past-and-future-of-ai-standards</link><guid isPermaLink="false">https://www.aipolicyperspectives.com/p/the-past-and-future-of-ai-standards</guid><dc:creator><![CDATA[Conor Griffin]]></dc:creator><pubDate>Tue, 17 Mar 2026 10:17:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!sMbF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sMbF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sMbF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!sMbF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!sMbF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!sMbF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sMbF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sMbF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!sMbF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!sMbF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!sMbF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F172dd689-5753-481c-8307-76bc1ce3a7db_1408x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini </figcaption></figure></div><p><strong>By Conor Griffin, Joslyn Barnhart &amp; Owen Larter</strong></p><p>In 1971, the marine archeologist Honor Frost heard news of wood protruding from the sea floor. Off the western coast of Sicily, she and her team donned scuba gear, and splashed into the shallow coastal waters. Wind whipped the surface, causing the underwater sand to swirl confusingly. But even in murk, they couldn&#8217;t miss it.</p><p>&#8220;A large timber (such as I had never seen before) emerged,&#8221; she <a href="https://artsandculture.google.com/story/the-discovery-of-the-marsala-punic-ship-honor-frost-foundation/2QWhIN7Uu9SK-Q?hl=en">recalled</a>, &#8220;like the head of a primeval animal crowned with weed; the presence of a buried wreck was evident.&#8221;</p><p>They excavated for months, gradually exposing the remains of a Carthaginian warship sunk more than 2,000 years before. Somehow, saltwater hadn&#8217;t eaten away letters painted on the wreckage, revealing a humble system that links antiquity to tomorrow.</p><p>Those shipwrights&#8217; marks told workers in ancient Carthage how to put together a vessel&#8212;akin to flat-pack furniture from IKEA, with numbered and lettered pieces. They were among the earliest surviving examples of a simple but potent tool in human progress: <strong>the technological standard</strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ydHP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ydHP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 424w, https://substackcdn.com/image/fetch/$s_!ydHP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 848w, https://substackcdn.com/image/fetch/$s_!ydHP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 1272w, https://substackcdn.com/image/fetch/$s_!ydHP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ydHP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png" width="1456" height="817" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:817,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ydHP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 424w, https://substackcdn.com/image/fetch/$s_!ydHP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 848w, https://substackcdn.com/image/fetch/$s_!ydHP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 1272w, https://substackcdn.com/image/fetch/$s_!ydHP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68982780-63f6-4c22-a3f8-f4486f3ef21b_1600x898.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: The Honor Frost Archive (MS439), University of Southampton.</figcaption></figure></div><p>Many of history&#8217;s grand projects have benefited from standards, from Egypt&#8217;s pyramids, to Europe&#8217;s cathedrals, to Gutenberg&#8217;s press, to everyone&#8217;s Internet. You can even thank standards for the development of beer.</p><p>Underpinning technological standards is a plain truth: people thrive when able to cooperate, not when we must keep negotiating the basics, whether it&#8217;s a matter of nuclear safety, or a phone-charger cord, or who goes next at the intersection. So, the goal is order. And the benefits are that innovators can proceed without excessive obstacles, while everyone else is treated fairly and kept safe.</p><p>But what should standards mean for artificial intelligence? In particular, how can they guide the most advanced large language models and AI agents that could transform society?</p><p>Venture around the AI frontier today, and you&#8217;ll find ambition to accelerate AI for economic growth and transformative science alongside concern that AI could clatter into what humans cherish most. What few dispute is this: standards will help set the path.</p><p>Standards have critics too. One criticism is that companies dominate the process, prioritizing their own products or miming security without truly ensuring it. Besides this, standards can stir geopolitical tensions, as when Western countries fear China&#8217;s influence in laying the path to tomorrow, while smaller nations worry that standards may be set without considering them at all.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pzRg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pzRg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 424w, https://substackcdn.com/image/fetch/$s_!pzRg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 848w, https://substackcdn.com/image/fetch/$s_!pzRg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 1272w, https://substackcdn.com/image/fetch/$s_!pzRg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pzRg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png" width="1024" height="702" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:702,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pzRg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 424w, https://substackcdn.com/image/fetch/$s_!pzRg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 848w, https://substackcdn.com/image/fetch/$s_!pzRg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 1272w, https://substackcdn.com/image/fetch/$s_!pzRg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a502385-eeee-4a5c-8a1a-beff6c058f82_1024x702.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Library of Congress</figcaption></figure></div><p>So, as we&#8217;ll keep insisting, standards matter! Only, there&#8217;s a problem.</p><p>For some, the mere mention of &#8220;standards&#8221; prompts slumber. And even those determined to stay awake may find themselves puzzled, gazing at the alphabet soup of standards organizations and committee meetings.</p><p>Part of the problem is that standards are often technical, such as efforts to standardize the protocols needed for AI agents to communicate. Or they are bureaucratic, negotiated out of public view, with dense, jargon-filled documents that are often behind a paywall.</p><p>Complicating matters even more, artificial intelligence is a general-purpose technology less akin to a hammer than to electricity. This will lead to standards (plus standards initiatives that don&#8217;t take) on everything from AI agents, to AI cybersecurity, to AI content provenance, to product-specific standards for AI-as-a-medical-device, and so on. And that&#8217;s not even mentioning standards for future AI applications that nobody has yet considered.</p><p>In short, standards will be immense. Standards will be tough to comprehend. But standards will also be vastly important.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><h3>CAN SOMEONE DEFINE STANDARDS, PLEASE?</h3><p>Standards are a diabolical blend: intricate, vague, and slippery.</p><p>They&#8217;re the invisible infrastructure of the modern world, according to Laurie Locascio, head of the American National Standards Institute, ANSI. She <a href="https://issues.org/who-sets-the-standard/?utm_campaign=34324386-Issues.org%20Newsletter&amp;utm_medium=email&amp;_hsenc=p2ANqtz-8wSj3J1qBqFPUcl9Fm0w3x0q01NijGR1b6Q05pK0a-5thOpaTOJHNf0vNbWaeBtAj68a6crdWY36mRoUGczN_ZIIpEaw&amp;_hsmi=404588886&amp;utm_content=404588886&amp;utm_source=hs_email">recounts</a> hearing an official at Boeing describe the airplane itself as &#8220;thousands of standards taking flight.&#8221; Standards are &#8220;the things you don&#8217;t think about,&#8221; Locascio says. &#8220;But oh, my God, you&#8217;re so glad they&#8217;re there.&#8221;</p><p>Expressed broadly, a standard defines the <em>how</em> of tech, whether it&#8217;s the default <em>product </em>specs that allow compatibility among manufacturers, or the formally endorsed risk management <em>processes</em> that encourage industry to act responsibly.</p><p>As technology evolves, standards do too. A leading scholar, Ken Krechmer, once <a href="https://web.njit.edu/~bieber/WWW-Standards-F01/krechmer96.pdf">noted</a> that standards initially defined how physical objects fit together (as with those markings on the Carthaginian longship). Over time, standards came to define the relationship <em>between</em> technological objects (as with internet protocols).</p><p>A standard also builds on other forms of guidance, such as norms, principles and industry best practices. Unlike norms, standards should be explicit. Unlike aspirational principles, a standard should be specific enough for performance against it to be judged. Unlike early best practices, a standard should have clear buy-in.</p><p>Developing a standard can be a protracted endeavor. In some cases, it might start in a researcher&#8217;s notebook, evolving into a product or a practice that gains traction in the marketplace. At other times, institutions set standards via years of deliberations and meticulous documents. Most often, it&#8217;s a messy back-and-forth between standards that emerge <em>in practice</em> and <em>on paper</em>. This makes standards a source of tension among companies, governments, and independent advocates, all trying to set the technological future they consider best.</p><p>Some presume that laws should be how we define permitted behavior. But high-quality legislation can struggle to keep up with the frantic speed of AI progress. And when laws are passed, they may rely on standards for implementation, as with the EU AI Act.</p><p>So how to persuade everyone to care when encountering standards, rather than just to snore or sob? How to get policy leaders to ponder the <em>entirety</em> of frontier-AI standards and align on where action is most needed?</p><p>Our answer is storytelling: to pluck forth tales about past standards, illustrating what this technological shaping can achieve, where it goes wrong, and how we might help cultivate standards wise enough to manage the breadth and speed of AI.</p><p>Our first stop? A battlefield of centuries ago.</p><h3>A &#8216;STANDARD&#8217; HISTORY</h3><p>Horrors encircled the boy soldier: swords clanging under the rain, excruciating howls of the wounded, the fast-approaching bellows of men hurtling across the bog to murder him. In wet turf, he shivered from knees to chattering teeth, his mouth parched, his gaze searching for any escape.</p><p>Up there?</p><p>On a hill, a flag rippled, where his legion had marked its territory. The Old French word for that banner was &#8220;<em>estandart</em>&#8221;: a sign of firmness and stability, a marker of where to go next, a statement of order amid chaos. To such banners, we owe the word &#8220;<a href="https://www.oed.com/dictionary/standard_n?tab=factsheet">standard</a>.&#8221;</p><p>More than a few historical standards emerged from war, where disorder could mean one&#8217;s brethren murdered, while coordination could mean an empire.</p><p><strong>~225 BCE to the dawn of mass production</strong></p><p>China&#8217;s first emperor, Qin Shi Huang, led an extensive <a href="https://www.google.com/books/edition/_/1OiMzAEACAAJ?hl=en&amp;kptab=overview">standardization process </a>that included mass-produced crossbow parts. If parts of a soldier&#8217;s weapon broke in the midst of battle, he could grab spares, and swap them in.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!37af!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!37af!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 424w, https://substackcdn.com/image/fetch/$s_!37af!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 848w, https://substackcdn.com/image/fetch/$s_!37af!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!37af!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!37af!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg" width="1456" height="958" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:958,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!37af!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 424w, https://substackcdn.com/image/fetch/$s_!37af!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 848w, https://substackcdn.com/image/fetch/$s_!37af!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!37af!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aef643-6da8-428b-a796-780ae6faa80d_1600x1053.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A Qin crossbow, displayed at Shaanxi History Museum, Xi&#8217;an. (Credit: WorldHistoryPics.com)</figcaption></figure></div><p>When Ancient Rome fought for primacy in the Mediterranean, its forces copied Carthaginian ship designs, eventually triumphing with <a href="https://books.google.com/books/about/The_Fall_of_Carthage.html?id=u684AgAAQBAJ&amp;source=kp_book_description&amp;redir_esc=y">standardized ships</a> of their own, along with standardized tools and <a href="https://artsandculture.google.com/story/roman-engineering/hgXhQHkIAE5q2g?hl=en">camp layouts</a>, all of which simplified maintenance and large-scale coordination.</p><p>Another advance in ancient times came from standardized measurements for length, volume, and weight. Previously, cultures often had distinct units; you can imagine the squabbling. But as trade expanded, standards prevailed, making cross-cultural exchange possible. In ancient Egypt, one of the earliest and most influential standards was the <a href="https://onlinelibrary.wiley.com/doi/10.1155/2014/489757">cubit</a>, a unit of length used to coordinate the building of the pyramids.</p><p>In Europe&#8217;s medieval period, <a href="https://books.google.com/books/about/The_European_Guilds.html?id=BrEPEAAAQBAJ&amp;source=kp_book_description&amp;redir_esc=y">guilds</a> established standards for quality control, so that weavers might set the necessary thread count or width of cloth, preventing low-quality products from undermining a craft&#8217;s reputation. Guilds also played a protectionist role, with licensing standards imposing strict controls on who could become a member.</p><p>The consumer might benefit from standards too, with measures such as England&#8217;s <a href="https://ifst.onlinelibrary.wiley.com/doi/10.1002/fsat.3801_5.x">Assize of Bread and Ale of 1266</a> establishing the acceptable quality, quantity, and price of baked goods and beer. Later, Gutenberg&#8217;s <a href="https://hob.gseis.ucla.edu/HoBCoursebook_Ch_5.html">standardized press</a> led to mass-produced books that spread ideas across the Continent.</p><p>However, technological standards reached new heights of utility during the Industrial Revolution, which set the foundations for many of today&#8217;s technologies.</p><p><strong>1760-1840: The First Industrial Revolution &#8212; The rise of engineers</strong></p><p>As ancient Chinese and Carthaginians had discovered long before, the Industrial Revolution&#8217;s manufacturers found that interchangeable parts offered transformative efficiency. Before, if you hand-built a musket, or a clock, or a steam engine, you might craft each screw, each<em> </em>bolt, each<em> </em>gear to fit. By contrast, interchangeability allowed for mass production, cutting costs, reducing errors, and establishing the basis for modern industry.</p><p>Screw threads are a classic example. Before standards, manufacturers used various designs, making repairs nightmarish. If you had one company&#8217;s bolt but another company&#8217;s nut, you were out of luck. In the 1800s, engineers built the first practical screw-cutting machines, allowing factories to produce <a href="https://en.wikisource.org/wiki/Miscellaneous_Papers_on_Mechanical_Subjects/A_Paper_on_an_Uniform_System_of_Screw_Threads">uniform threads</a> and a consistent system of measurement. The British Standard Whitworth became the first such standard in the world.</p><p>Screw-thread standards may not quicken your pulse. But their effects might. They played a part in British imperial ambitions, contributing to the expansion and maintenance of the British Empire through <a href="https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1095-9270.2004.00028.x">military</a> mobilization.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TPXx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TPXx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TPXx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TPXx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TPXx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TPXx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TPXx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TPXx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TPXx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TPXx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ba5c329-6ce2-4034-9876-a1c5279b4276_1600x873.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: NotebookLM</figcaption></figure></div><p><strong>1870-1914:</strong> <strong>The Second Industrial Revolution &#8212; National coordination &amp; path dependence</strong></p><p>The emergence of electricity, steel, and advanced machinery led to vast interconnected systems, including power grids, railways, and telegraph networks. Coordination wasn&#8217;t merely better; it was essential. To coordinate across a nation&#8212;and eventually across borders&#8212;the ambitious country needed technology standards. Two famed cases illustrate this, one successful, one bungled.</p><p>The success regards the quintessential technology of the times: railroads. By the 1870s, the U.S. rail system was a mess, with more than 20 different track gauges. When a train reached a section built to a different track-gauge width, everything&#8212;each passenger, piece of luggage, every single crate&#8212;had to be unloaded, and transferred to a new train.</p><p>By the 1880s, matters had become slightly less chaotic, with either a southern gauge or northern &#8220;standard&#8221; gauge used across most of the country. Yet this still divided national transport until, in  1886, rail companies pulled off a <a href="https://dash.harvard.edu/entities/publication/73120379-10c3-6bd4-e053-0100007fdf3b">remarkable feat</a>. Over two days, they converted 13,000 miles<em> </em>(that&#8217;s 21,000 kilometers) of southern U.S. track to the northern standard, integrating the national transportation network. When trains rolled out on June 2, 1886, they were able to travel seamlessly across the United States for the first time in history.</p><p>A second case illustrates bungled standards. In the 1880s, the rival inventors Thomas Edison and Nikola Tesla found themselves at the center of &#8220;<a href="https://books.google.com/books?id=2_58p3Z69bIC&amp;source=gbs_book_other_versions&amp;redir_esc=y">the War of the Currents.</a>&#8221; Edison championed direct current (DC), a one-directional flow of electricity that had been the early U.S. standard. Tesla, backed by entrepreneur industrialist George Westinghouse, advocated alternating current (AC), or electricity that reverses direction many times per second, and can be stepped up or down in voltage with a transformer.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!assm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!assm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 424w, https://substackcdn.com/image/fetch/$s_!assm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 848w, https://substackcdn.com/image/fetch/$s_!assm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 1272w, https://substackcdn.com/image/fetch/$s_!assm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!assm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png" width="1024" height="535" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:535,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!assm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 424w, https://substackcdn.com/image/fetch/$s_!assm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 848w, https://substackcdn.com/image/fetch/$s_!assm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 1272w, https://substackcdn.com/image/fetch/$s_!assm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf132576-b74c-4a94-b82b-c2e7158aa434_1024x535.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Gemini</figcaption></figure></div><p>From an engineering standpoint, AC had a decisive advantage: it could transmit power over long distances cheaply and efficiently, while DC could not. AC eventually won out. But by the time it had emerged as the superior solution, the world had already built electrical systems without any coordinated technical governance. As there was no international authority harmonizing electrical standards, the United States went with 120 volts at 60 hertz (a legacy of Edison&#8217;s early low-voltage DC networks). Much of the rest of the world adopted 230 volts at 50 hertz.</p><p>Once wires had been laid and appliances built, the world was locked into two incompatible systems. To this day, we&#8217;re burning out hair dryers bought in America but used in Paris, or realizing too late that we don&#8217;t have the right <a href="https://www.iec.ch/world-plugs">plug</a> for our laptops. If it&#8217;s irksome for the average user, it&#8217;s more burdensome for manufacturers, obliging them to build different versions for different countries.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><p>Another classic tale of path dependence is under our fingertips as we type: the QWERTY keyboard. Why <em>does </em>the top row spell QWERTYUIOP? One account goes like this: In the mid-to-late 1800s, early typewriters jammed each time the user struck neighboring keys in rapid succession. So, designers produced a <a href="https://patents.google.com/patent/US182511A/en">layout</a> that deliberately distanced many common letter pairs. Remington purchased this QWERTY design, and began mass-producing typewriters.</p><p>Before long, typing schools had trained the future secretarial workforce on QWERTY, while firms wanting fleet-fingered staff had to buy those machines. Manufacturers subsequently  resolved the key-jamming problem and other keyboards <a href="https://www.smithsonianmag.com/history/the-qwerty-keyboard-will-never-die-where-did-the-150-year-old-design-come-from-49863249/">tried to</a> depose QWERTY, some <a href="https://fbaum.unc.edu/teaching/articles/David_AER_1985.pdf">claiming</a> to quicken typing by as much as 40%. But QWERTY had become a de facto standard. (Scholars continue to <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1069950">debate</a> the specifics, with some arguing that QWERTY works just fine.)</p><p>In any case, the indisputable lesson is to watch for <a href="https://www-2.rotman.utoronto.ca/insightshub/behavioural-economics-marketing/beware-path-dependence">path dependence</a>. The standards we establish for frontier AI today&#8212;or fail to establish&#8212;may determine future efficiency or future failure.</p><h3><strong>1914-1964: Standard Development Organizations &amp; Digital Technology</strong></h3><p>In 1918, engineering societies joined with the U.S. government to establish a standards committee that developed into <a href="https://www.ansi.org/about/history">ANSI</a>, the American National Standards Institute. Today, ANSI provides the &#8220;stamp of approval&#8221; for many U.S. standards organizations, including those working on AI. In subsequent decades, standardization went global. While the United Nations was founded as a governmental venue for diplomacy, the International Organization for Standardization, <a href="https://www.iso.org/news/2017/02/Ref2163.html">ISO</a>, emerged as a non-governmental body for peaceful technical coordination across borders. Bit by bit, additional standards bodies formed, cooking up the alphabet soup of acronyms&#8212;each a different org, subgroup, or committee&#8212;that lies before us today.</p><p>Soon, another transformation for standards was taking shape in the form of digital tech. Back then, computers filled entire rooms of universities, and each manufacturer built hardware and software within its own format. Computers could not run programs written for other systems, and accessories like printers or storage devices were incompatible.</p><p>A turning point came in 1964, with <a href="https://www.ibm.com/history/system-360">IBM&#8217;s System/360</a>. Software on one model could more easily run on another; accessories like printers worked across IBM models. You could upgrade and expand computer systems with relative ease.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zNAA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zNAA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 424w, https://substackcdn.com/image/fetch/$s_!zNAA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 848w, https://substackcdn.com/image/fetch/$s_!zNAA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 1272w, https://substackcdn.com/image/fetch/$s_!zNAA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zNAA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png" width="1024" height="806" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:806,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zNAA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 424w, https://substackcdn.com/image/fetch/$s_!zNAA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 848w, https://substackcdn.com/image/fetch/$s_!zNAA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 1272w, https://substackcdn.com/image/fetch/$s_!zNAA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88d7ee18-5b8e-4d84-9aad-39d17a40347d_1024x806.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: U.S. National Archives and Records Administration</figcaption></figure></div><p><strong>1969-today: Talking Machines</strong></p><p>The year that hippies grooved at Woodstock and astronauts walked on the Moon, the U.S. Defense Department was testing a project that must have seemed minor by comparison: connecting research institutions and government agencies. Yet the Advanced Research Projects Agency Network, ARPANET, which sent its first message in 1969, was the precursor to our transformed world.</p><p>Before ARPANET, <a href="https://artsandculture.google.com/story/from-punch-cards-to-the-cloud-museum-for-communication-frankfurt/VQXBq16p7orTYw?hl=en">moving information</a> from one computer to another was a struggle, with researchers forced to carry magnetic tapes or punched cards between locations, while those working far apart had to rely on snail-mail.</p><p>To convey information between independent systems, ARPANET adopted packet switching, breaking data into small units that could travel independently and reassemble at their destination. Extending this, Robert Kahn and Vint Cerf began designing a <a href="https://ieeexplore.ieee.org/document/1092259">universal communication framework</a> in 1973 for different types of networks to connect. Their collaboration ultimately produced <a href="https://cloud.google.com/blog/topics/public-sector/50-years-internet-celebrating-vision-vint-cerf-and-bob-kahn-and-exploring-future-connectivity-and-innovation">TCP/IP</a>, the Transmission Control Protocol and Internet Protocol that underpins today&#8217;s online communication.</p><p>A key effect of the TCP/IP standard was decentralization: no single authority could control the flow of data, and any network that adhered to the protocol could connect without permission from central authorities.</p><p>In 1989, a British scientist at CERN, Tim Berners-Lee, <a href="https://www.w3.org/History/1989/proposal.html">proposed</a> another transformation that developed into a project called &#8220;<a href="https://docdrop.org/download_annotation_doc/Tim-Berners-Lee---Weaving-the-Web_-The-Original-Design-and-U-88myd.pdf">WorldWideWeb</a>,&#8221; which envisioned a global <a href="https://home.cern/science/computing/birth-web/short-history-web">network</a> of documents accessible through software, operating on <a href="https://timeline.web.cern.ch/cern-puts-world-wide-web-public-domain">open standards</a> that nobody could lock it into a proprietary system. Two standards organizations, the Internet Engineering Task Force and the World Wide Web Consortium, helped to formalize the vision, crafting standards for structuring content (HTML), transferring data (HTTP), identifying resources (URI), and more.</p><p>But while standards help spread technology, this diffusion can also lead to greater harm. The expansion of railroads led to more wrecks, forcing uptake of safety  standards for signaling, brakes and more. When electricity was first installed in the White House in the late 19th century, President Benjamin Harrison and his wife Caroline <a href="https://www.energy.gov/articles/history-electricity-white-house">were so afraid</a> of shocks that they refused to turn the lights off. Such fears&#8212;often well justified&#8212;led to the standardization of building and electrical codes. When it came to digital technology, the risks extended beyond immediate physical safety into areas like data theft. This demanded standards such as <a href="https://www.ssl.com/article/what-is-ssl-tls-an-in-depth-guide/">SSL/TLS</a> to provide security for data sent over computer networks.</p><p>A <a href="https://en.wikipedia.org/wiki/Collingridge_dilemma#:~:text=The%20Collingridge%20dilemma%20is%20a,extensively%20developed%20and%20widely%20used.">recurrent challenge</a> with frontier tech is that experts struggle to predict how exactly it will affect society. But once it is widely used, it can be sticky and hard to change. The effects of powerful technologies can also be subtle, indirect and slow-burning, for example if they change how we access and consume information. In the digital era, this has shifted technological standards from periodic safety checks of products towards ongoing <em>processes</em> that organizations can use to identify, evaluate and mitigate a growing suite of risks.</p><p>By way of example, the U.S. government&#8217;s National Institute of Standards and Technology, NIST, introduced the voluntary <a href="https://www.nist.gov/itl/ai-risk-management-framework">AI Risk Management Framework</a> in 2023, building on its earlier framework for managing cybersecurity risks. Likewise, the <a href="https://www.iso.org/committee/6794475.html">ISO/IEC committee on AI</a> that is considering <a href="https://www.safer-ai.org/an-overview-of-existing-and-potential-future-genai-gpai-standards">standards</a> on everything from red-teaming to LLM interoperability also passed the first official international AI management standard, <a href="https://www.iso.org/standard/42001">ISO/IEC 42001</a>, which organizations can use to demonstrate that they are responsibly integrating AI into their operations. </p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/subscribe?"><span>Subscribe now</span></a></p><p></p><h3>5 LESSONS FROM HISTORY</h3><p>Studying the past, you see how often standards&#8212;by design or bumbling&#8212;have shaped the technological present. But what about our technological future?</p><p>To develop good standards for general-purpose AI models and agents, we&#8217;ll need inputs from a range of groups, from scientists with know-how to institutions who can convene. Below [see infographic], we have identified five groups who&#8217;ll perform key roles.</p><p>What we mapped includes more than just official standards development organizations. We also want to capture the early spaces where standards emerge <em>in practice </em>before they are formalized <em>on paper. </em>How this works is closer to a swirl of inputs than a steady procession. Sometimes, the same organization or individual may operate in several groups at the same time. Ideas and efforts may also originate in one group, then migrate to another, with different groups offering varying degrees of speed, flexibility, expertise, and perceived neutrality.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Ue2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Ue2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-Ue2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-Ue2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-Ue2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Ue2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-Ue2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-Ue2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-Ue2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-Ue2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d1469b4-2b50-4ec6-a9ad-5e12e38246fa_1600x893.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">NotebookLM</figcaption></figure></div><p>For these groups&#8212;and the policymakers, business leaders, and advocates who shape their work&#8212;what lessons can history teach about our AI future? Here are five:</p><ol><li><p><strong>Standards matter! </strong>At best, technological standards chart a wise path; at worst, they fill the path with potholes. Consider the bulky electrical converters that one still needs when traveling&#8212;it didn&#8217;t have to be that way. On the other hand, when we get it right, the benefits of technology spread faster, more inclusively, and more securely.</p><p></p></li><li><p><strong>The standards process needs to speed up. </strong>ISO says that the <a href="https://www.iso.org/developing-standards.html">average time</a> to develop one of its standards is three years, and ISO is not an outlier. Given the pace of change in AI, we need to speed up. For priority goals, like finding secure ways for agents to operate and interact,  which the US Center for AI Standards and Innovation <a href="https://www.nist.gov/caisi/ai-agent-standards-initiative">is working on</a>, we need to find ways to accelerate that don&#8217;t jeopardize the overall quality and integrity of the process. This may mean looking across the many groups now focusing on AI standards and finding ways to collaborate early, rather than duplicate. It may mean focusing more on technical protocols and <a href="https://scc-ccn.ca/standards/flexible-standards-based-solutions/publicly-available-specification">specifications documents</a> that are quicker to develop. It may also mean using AI to <a href="https://www.w3.org/community/aiwss/">help deliberate on and write standards</a>, and moving to more <a href="https://www.iso.org/smart">nimble digital formats</a> that are easier to update and use. </p></li></ol><ol start="3"><li><p><strong>We need more efficient ways to input on standards. </strong>All standards, from those underpinning steam engines to the Internet, had to chart a unified path through diverging viewpoints, with an end result that did not please everyone. For AI, the challenge will be far greater. It is more akin to 1,000 technologies, and will affect different groups in different ways. This means that any broad directive&#8212; say, to &#8220;develop standards that make AI fair&#8221;&#8212;risks an <a href="https://www.nytimes.com/2023/04/02/opinion/democrats-liberalism.html">everything-bagel solution</a>. Many groups would rightly be heard, but the output would be too vague to provide the &#8220;how&#8221; that justifies a standard, leading to confusion, a stifling of innovation or the standard being ignored. This suggests that most standards should be precise in scope, targeting specific components of AI systems or specific concerns, from certifying <a href="https://spec.c2pa.org/specifications/specifications/2.3/index.html">the source and history of online content</a> to combating the leaking of confidential data. More precise standards will make it easier to identify a wider range of relevant voices and incorporate their input.</p></li></ol><ol start="4"><li><p><strong>Frontier AI standards should focus on large-scale risks. </strong>Historically, standards have accelerated the diffusion of technology, amplifying its benefits but also, in places, its negative impacts. For AI, foresight and risk management standards will be critical to getting ahead of future risks and speeding adoption. But with a technology as general-purpose, fast-improving, and poorly understood as AI, perfect foresight is impossible. Standards move at a human pace and cannot standardize a future that we cannot perfectly see. As a result, the focus should be on developing scientifically robust standards to address the most consequential or large-scale risks, such as those targeted by labs&#8217; <a href="https://deepmind.google/blog/strengthening-our-frontier-safety-framework/">Frontier Safety Frameworks</a>.</p></li><li><p><strong>Wrong paths are inevitable, so we should catch them early. </strong>Now and then, technology stumbles into a poor standard, and it&#8217;s onerous to go back. But not necessarily impossible. Especially if we act if we catch it early. Consider the U.S. railroads taking action to unify their systems through a mighty coordinated effort. Groups working on AI standards devote much time to building consensus about new initiatives. They should also use the processes available to them to review and withdraw standards, where needed, to avoid sub-optimal lock-in. This also means giving third parties more opportunities to access, understand and constructively critique early AI standards. And designing standards and protocols that are modular, and can be swapped out, or updated, without major downstream consequences.</p></li></ol><div><hr></div><h2>QUESTIONS FOR YOU</h2><ol><li><p>Where do you feel most hope for frontier-AI standards?</p></li><li><p>Where do you worry about a lack of progress on frontier AI standards?</p></li><li><p>When you imagine a missing standard for frontier AI, what is that? A technical protocol specified in code? Or a fuzzier process-standard?</p></li><li><p>Might your standard become politicized? Is it something that hinges on values? Or might most governments in the world support its adoption?</p></li><li><p>What&#8217;s a scenario in which your proposed standard goes awry? How could you detect and mitigate that?</p></li><li><p>What would be the primary role of government in your standard? Supplying technical expertise? Convening authorities and experts? Incentivizing your standard via public procurement, regulation, or other methods?</p><p></p></li></ol><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.aipolicyperspectives.com/p/the-past-and-future-of-ai-standards?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.aipolicyperspectives.com/p/the-past-and-future-of-ai-standards?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><em>Thank you to Shaked Karabelnicoff, Tom Rachman and Bruno Galizzi for support with research and review. As with all pieces you read here, this is written in a personal capacity. All opinions and any mistakes belong to the authors.</em> </p>]]></content:encoded></item></channel></rss>