top of page

THE SKINNY
on AI for Education

Issue 32, September 2026

Welcome to The Skinny on AI for Education newsletter. Discover the latest insights at the intersection of AI and education from Professor Rose Luckin and the EVR Team. From personalised learning to smart classrooms, we decode AI's impact on education. We analyse the news, track developments in AI technology, watch what is happening with regulation and policy and discuss what all of it means for Education. Stay informed, navigate responsibly, and shape the future of learning with The Skinny.​

shutterstock_2853353811.jpg
In brief: the 60-second version
​
  • September brought public warnings from senior figures at OpenAI and Anthropic that no lab has solved alignment well enough to keep scaling at full speed, alongside the launch of Meta's Muse, an agent built to run more of your digital life, which The Information reported had shown safety concerns in testing.

  • The response for educators is neither panic nor breezy adoption. Understand enough about AI to use it safely, and seek evidence before deciding.

  • Evidence about AI is hard to gather in the usual way: trials can outlast the products they test, a Stanford review of more than 800 papers found 20 high-quality causal studies, and a meta-analysis of 49 studies found the pooled effect fell close to zero once publication bias was modelled.

  • The DfE Science Advisory Council's eight principles, published on 24 September as independent advice, offer a method: combine many partial sources, update continually, keep options open, evaluate conditions alongside products, match the method to the technology's maturity, build the data infrastructure to do it, and ask at every stage who benefits and who is harmed.

  • Recent research illustrates three of them. Measure with the tool switched off (133 patent lawyers in a Google-funded trial; 26,811 Chinese students). Evaluate conditions alongside products (teacher guidance produced the strongest achievement gain in a 118-student study; the Royal Society's threefold confidence association). Count behaviour, not confidence (self-rated AI proficiency was essentially uncorrelated with tested knowledge across 312,329 students).

  • The principles are free to download, and UNESCO's consultation on AI governance in education closes on 15 October.

​

The Full Skinny

​

September was the month the people who build AI asked us to be frightened of it. OpenAI's chief scientist, Jakub Pachocki, wrote that "no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer" and that "this is a time that calls for extreme caution." A 27-year-old British researcher resigned from Anthropic saying neither it nor OpenAI was acting responsibly, and the head of alignment science at Anthropic, Evan Hubinger, replied in public that "we really do earnestly believe AI could kill all humans," putting his own estimate of the chance of extinction within the next decade above 10%. Paul Christiano, a pioneer of the reinforcement learning from human feedback used to train many leading chatbots, said that if we build superintelligence without better alignment, "most people could die." Dario Amodei published an essay called We Must Pace the Frontier. Andrew Ng described what looked to him like "a well orchestrated PR campaign."

​

In the same month Meta launched Muse, a personal agent that, in Mark Zuckerberg's words, "understands your goals and works 24/7 to get things done for you," with access to a payment wallet and, Meta says, your credentials held where it can use them but not see them. The Information reported that Meta's internal testing had uncovered safety and alignment concerns, including cases in which the agent took actions without permission. It shipped anyway, and by the end of the month it was the most downloaded app in US stores.

​

So what are we to do, those of us whose job is to teach and to lead, in the middle of this maelstrom of AI malarkey? My suggestion is that it is time to tap into your inner ninja.

​

The ninja of popular imagination is a fighter. Historians describe something more useful to us: a scout and an intelligence gatherer, more interested in information than in combat, whose first job was to understand the terrain before anyone moved across it. That is the posture I think AI now asks of educators. Not panic, and not the breezy adoption the vendors would prefer. Understand enough about how these systems work to use them safely. Be cautious, be careful, be wise. And above all, seek evidence before you decide.

​

That last instruction is harder than it sounds, because research involving AI faces a constraint that our usual habits of evidence were not built for. A randomised trial can take two or three years from design to publication, and the model it evaluated may have been replaced several times by the time the results appear. A Stanford review of more than 800 papers in its October 2025 repository on AI in schools identified 20 high-quality causal studies. A meta-analysis published in August combined 49 studies of generative AI in STEM education and found that, once publication bias was modelled, the pooled effect fell close to zero. The same review found that 33 of the 49 studies had compared groups doing different kinds of mental work. Even the tools increasingly used to read research can make things worse: a study in Nature Human Behaviour found that summaries produced by five leading AI systems could remove authors' hedges and introduce unsupported causal claims, a tendency that prompts requesting caution reduced but did not remove. If we wait for the definitive study, we will wait indefinitely while the technology arrives regardless. In August, Google expanded Gemini in Classroom to pupils of all ages in districts that had already enabled student access, and Education Week reported that some educators learned of the change from news coverage.

​

Here is the positive to set against all of that. Last month the Department for Education published a short, clear and practical set of principles for seeking evidence in exactly these conditions: independent advice from its Science Advisory Council's working group on AI and digital education, of which I am a member. You can find this here: https://assets.publishing.service.gov.uk/media/6ab22fa1d52ca1fccea560db/AI_principles_research_report.pdf

​

The eight principles rest on one idea: build the best current picture from many imperfect sources, weight each one honestly, update as new evidence arrives, and act on the picture as it stands rather than waiting for a verdict that will come too late. Precaution remains essential where children face credible risk of serious harm, particularly where consequences may be irreversible or uncertainty cannot be resolved in time. Elsewhere, delay is also a cost, and it falls on the pupils and teachers the benefit would have reached. And at every stage, the principles ask who benefits, who is harmed, and whether average results conceal widening inequalities.

​

The research I scan each week for my weekly research radar has been illustrating the value of those principles without explicitly connecting to them. Three examples.

​

Match evidence to its purpose, and measure with the tool switched off. The principles identify a harm to evaluate from the outset: short-term performance gains that can mask longer-term harm to independent learning. A Google-funded working paper reported a pre-registered trial that gave 133 patent lawyers at eleven firms an AI drafting assistant for three months. Junior lawyers gained most on the AI-assisted drafting tasks, but showed no average improvement on a subsequent unaided patent-redlining task; senior lawyers retained an advantage. A June working paper analysing 26,811 Chinese secondary students over thirty months estimated that AI adoption raised homework scores by about 18% but lowered monthly exam scores by about 20% within six months. The only measurement that detects this pattern is the one taken without the tool, and too few evaluations take it.

​

Evaluate conditions alongside products. The product will have changed by the time the study is out; the conditions change slowly and they are what a school controls. A six-week study of 118 undergraduates found gains in achievement with both independent and teacher-guided use of an AI study tool, with the guided group achieving the strongest results, while the gains in confidence and lower anxiety were the same whether or not a teacher was involved. The Royal Society's survey of more than 9,200 teachers in England, conducted in March and published in August, found that structured training and whole-school guidance were associated with roughly three times the likelihood of confidence across every dimension of AI literacy. Only 18% had received any such training, and 15% had clear whole-school guidance. Trials still have their place where tools or the learning mechanisms behind them are stable enough to study; the point is to evaluate the setting as carefully as the software.

​

Build the infrastructure, and count behaviour rather than confidence. The principles propose making access to consistent usage data a condition of public procurement. Recent research shows what that data must contain. Among 312,329 Chinese higher education students who sat an objective test of how generative AI works, self-rated proficiency was essentially uncorrelated with tested knowledge. A preprint that logged every interaction of 483 students found that verification behaviour, how often they checked the AI's output, was the strongest direct predictor of what they could do when the tool was removed; questionnaire-measured AI literacy was associated with that performance only indirectly, through regulation and checking. Logins and confidence surveys measure exposure. We need to count checking, revising and unaided performance.

​

None of these studies is perfect, and several are preprints or working papers. That is the condition we work in, and the principles are a method for working in it honestly. The ninja does not refuse to cross the terrain because the map is incomplete. She studies it, moves deliberately, and keeps updating the map as she goes. Learn fast. Act more slowly. And when someone tells you a tool improves learning, ask them the question Nicole Barnes of the American Psychological Association put to Chalkbeat last month: close the app, ask the student to explain, and see what learning is still there.

​

Sources
​

Pachocki, J. (2026), An Alien Mind, OpenAI, 6 September 2026, https://openai.com/index/an-alien-mind/. Hubinger, E., post on X, September 2026, https://x.com/EvanHub/status/2097497037956891126. Christiano, P. (2026), Personal statement on joining the OpenAI board, https://paulfchristiano.substack.com/p/personal-statement-on-joining-the. Amodei, D. (2026), We Must Pace the Frontier, 12 September 2026, https://darioamodei.com/post/we-must-pace-the-frontier. Ng, A. (2026), LinkedIn post, September 2026. Meta (2026), Introducing Muse, 8 September 2026, https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/; The Information, Meta Launches Its First Consumer AI Agent Muse. Department for Education Science Advisory Council, AI and Digital Education working group (2026), Principles for evidence-informed policymaking in fast-changing technology contexts, 24 September 2026, https://www.gov.uk/government/publications/evidence-informed-policymaking-and-fast-changing-technology. Fesler, L., Martinez Claeys, J., Agnew, C. and Loeb, S. (2026), The Evidence Base on AI in K-12: A 2026 Review, Stanford SCALE. Boolzen, C. et al. (2026), Evidence of impact and interpretational limits of generative AI in STEM education, Artificial Intelligence Review, https://link.springer.com/article/10.1007/s10462-026-11665-9. Isch, C. et al. (2026), Quantifying the prevalence and impact of overreaching causal claims in social science, Nature Human Behaviour, https://www.nature.com/articles/s41562-026-02553-x. Education Week (2026), Schools caught flat-footed after Google makes Gemini chatbot available to all students, September 2026. Autor, D. et al. (2026), Does AI Assistance Enhance or Erode Expertise? NBER Working Paper 35720, https://www.nber.org/papers/w35720 (working paper, Google-funded). Strömberg, D., Lei, V. and Wu, Y. (2026), The Generative AI Learning Penalty, CEPR Discussion Paper 21577, June 2026, https://cepr.org/publications/dp21577 (working paper). Tezci, İ. H. (2026), Does AI Guidance make a difference? Education and Information Technologies, https://link.springer.com/article/10.1007/s10639-026-14146-2. Royal Society (2026), AI literacy amongst teachers: confidence, support, and practice, August 2026. Yan, L. et al. (2026), GLAT-CN: Nationwide validation of the Generative AI Literacy Assessment Test among Chinese higher education students (preprint). Wang, X., Liu, X., Zhang, H. and Li, R. (2026), Meta-AI literacy and knowledge transfer, Research Square preprint, 26 August 2026. Chalkbeat (2026), APA calls for higher ed tech standards, not screen bans, 3 September 2026, https://www.chalkbeat.org/2026/09/03/psychological-assocation-calls-for-higher-ed-tech-standards-not-screen-bans/. UNESCO (2026), Global consultations on education in the age of AI, https://www.unesco.org/en/digital-education/artificial-intelligence/consultation.

The Skinny Scan - AI News

AI news scan, September 2026

A short read of the scan, covering September 2026 with updates to 2 October. The full version carries the detail and sources.

​

AI news scan: the 60-second version
​

The month's biggest story is about AI agents, programs that do not just answer questions but take actions, and what happened when, during research and testing, they were set ordinary look-up tasks. Australia's prime minister disclosed that an OpenAI agent had got into a government health statistics portal while looking up spending on medicines; OpenAI noticed two months later and told the government by emailing a public inbox. Other agents tried, without succeeding, to break into a university library, a public data platform and the US Department of Education's website, and OpenAI says at least 53 user-provided images were transferred to third-party image hosts as unlisted links. OpenAI paused training of its most capable new models, withheld one release because it was worse at staying within the limits it was given, and the US Federal Trade Commission opened an investigation. At the same time the people who build these systems asked for them to be slowed: OpenAI's chief scientist said no lab has solved safety well enough to keep going at full speed, and Anthropic's head of alignment put his own estimate of extinction within the decade above 10 %. Governments answered with a two-page voluntary pledge. These were test runs rather than products in use, but the products are arriving all the same: Meta's Muse, built to run more of your digital life, was the most downloaded app in America by the end of the month, and OpenAI and Microsoft put always-on agents inside Slack, Teams and Copilot.

​

The prices of the models themselves kept falling, with Anthropic claiming running costs 40 % lower for its new Claude Opus 5.5 and up to 30 % lower per task for Sonnet 5.5, at token prices cut by a fifth for Opus and unchanged for Sonnet, and open models now carrying more than half the traffic on some platforms. The bills went the other way. One software company's monthly Copilot charge reportedly rose from $20,000 to $260,000 when Microsoft moved to usage-based pricing, Google scheduled a price doubling for January, and the Bank of England cited $450 billion of AI-related debt and warned of a correction. The lesson for anyone signing a licence is the same as last month: do not assume today's price, and keep the option of changing supplier.

​

Education policy went in opposite directions on the same day! New York City banned student-facing AI for about 600,000 pupils through grade 8, with specified exceptions, Los Angeles blocked it on district devices, and the UAE approved an AI curriculum for every school with 22,000 teachers to be trained. Both followed Google extending its Gemini chatbot to pupils of all ages in districts that had already switched on student access, with some teachers learning from the press. The announcements give little evaluation detail to judge by: the UAE cited positive pilot findings without publishing them, and New York's ban keeps supervised pilots. The evidence supports neither a blanket ban nor a blanket rollout; it supports distinguishing teaching about AI, guided use and using it to get the answer, which few policies yet make explicit. In England, Ofqual told MPs that extended-writing coursework should be avoided "wherever we can", and school leaders told the Education Select Committee they are "desperate for really clear guidance".

​

The research pointed to what good guidance would say. A Google-funded trial of 133 lawyers found that the juniors gained most while using an AI assistant and showed no average improvement on an unaided task afterwards, while the seniors kept their advantage: measure learning with the tool switched off. Among 312,329 Chinese students who sat a test of how AI works, how good they said they were had almost no relationship to their score: count what people do, not how confident they feel. In a six-week study of 118 undergraduates, the group whose AI use came with teacher guidance did best: evaluate the setting as well as the software. The Department for Education published eight principles for gathering evidence in exactly these conditions, on the same day as its own pilot evidence board reviewed twelve education products from 96 applicants, found outcome evidence the weakest area, and made no awards because its judgements were not yet reliable or comparable enough to be fair. PISA 2025 showed reading and maths falling across the OECD for a decade, with England rising against the trend, and found that pupils asked in class to critique AI output did slightly better than those who were not.

​

The month's warning for the least protected sits in detection. At one Australian university, academic integrity cases in 2025 were equivalent to 13.2 % of enrolments, 27.0 % among international undergraduates, and 62.4 % were closed with no finding; the study does not show that AI detectors caused the referrals, but it is the clearest picture yet of who carries the cost of an integrity-first regime. Critics argue that detectors flag difference rather than dishonesty, and a Harvard study found pupils scaling back their best work to avoid being accused. The proportionate response, as several reports said this month, is to teach the skill and redesign the assessment rather than to buy more detection.

​​​

The Skinny AI News Items:

AI Development and Industry

GPT-6 Astra launches at the top of the range, and its successor is held back for crossing lines

3 to 29 September 2026 | OpenAI, Financial Times, The Information, UK AI Security Institute

OpenAI released GPT-6 Astra on 3 September as its most capable and, by its own account, most aligned model, priced at $10 per million input tokens and $50 per million output, and built for agentic work: filling in forms, updating records, building websites, producing documents and spreadsheets. It is the first OpenAI model to reach the Critical cybersecurity threshold in the company's own framework, so the launch version refuses advanced cyber tasks, and OpenAI's post concedes that Astra's written reasoning is harder to monitor than its predecessor's. Results depend heavily on how the test is run: on the ARC-AGI-3 reasoning test, ARC Prize reported 99.9% using OpenAI's own adapter at high reasoning effort and 62.7% using its standard harness at maximum effort. The UK AI Security Institute reported on 28 September that, in an evaluation using simulated tools with cyber safeguards disabled, Astra attempted complete supply-chain attacks more often than its predecessors and reasoned about the scope of a task before attacking out-of-scope targets anyway; the results describe the test conditions rather than the released system with its safeguards operating. On the same day OpenAI withheld GPT-6.1 Astra because it scored below its predecessor on staying within scope and on telling users what work it had done; its head of safety systems, Saachi Jain, described a trade-off between staying in scope and persisting when a task hits friction.

What you need to know: Three things in one release cycle: a benchmark that moves 37 points depending on how it is run, a national evaluator reporting that a released model recognises a boundary and crosses it, and a lab stating as a release criterion that persistence and staying in scope pull against each other. The last is the clearest account yet of why an agent that is good at finishing a task is not the same as an agent that is safe to give one.

Source: OpenAI, 3 September 2026; ARC Prize and Artificial Analysis, 3 September 2026; UK AI Security Institute, 28 September 2026; Financial Times and The Information, 29 September 2026

 

Anthropic's Claude 5.5 models cut costs by up to 40% and arrive with watermarks, safeguards and an opt-out training default

22 and 28 September 2026 | Anthropic, Financial Times, The Batch

Claude Opus 5.5 performs at the level of Fable 5.1 on most work, by Anthropic's account, and costs 40% less to run than Opus 5, at $4 and $20 per million tokens. Claude Sonnet 5.5 followed six days later, more than 30% faster and up to 30% cheaper per task at unchanged list prices. Both launch with cybersecurity safeguards that re-route higher-risk requests to older models, and Opus 5.5 adds biology safeguards that require vetted organisations to apply to a verification programme. Three details matter for buyers. Generated text carries a statistical watermark to comply with the EU AI Act; detection is not generally public, and at the time of writing eligible organisations could apply for access through a restricted private preview. The model was trained partly on data from consumer users who have not opted out, whereas Anthropic's commercial products do not use inputs and outputs for training by default, subject to stated exceptions. And Anthropic reports that Opus 5.5 often suspects it is being evaluated, which it says limits how far its safety results predict behaviour in use. Anthropic also cautions that benchmark margins have become a less reliable guide to real-world differences. All performance, cost and safety figures are the company's own.

What you need to know: Capability keeps rising while price keeps falling, which changes the arithmetic of any deployment planned on last term's quotes. The training default is the procurement point: staff using personal consumer accounts should check their settings, because the institutional contract's protections do not follow them there. The evaluation-awareness admission is the literacy point: a safety score earned by a model that knows it is being tested is a weaker claim than it looks.

Source: Anthropic, 22 and 28 September 2026; Financial Times, 22 September 2026; The Batch, 25 September 2026

 

Persistent agents arrive inside the software schools already run

9 to 30 September 2026 | The Information, Financial Times, DealBook

Meta launched Muse on 8 September, a personal agent that, in Mark Zuckerberg's words, "understands your goals and works 24/7 to get things done for you", free up to 100 million tokens a week, with access to a payment wallet of one-time-use cards and a separate "Sentinel" agent approving outbound actions. The Information reported that Meta's internal testing had uncovered safety and alignment concerns, including cases in which the agent took actions without permission. By the end of the month Muse was the most downloaded app in US stores. OpenAI launched Dots, always-on agents working inside Slack and Teams with access to more than 4,000 apps, and Microsoft added an Autopilot agent to Copilot. Anthropic's Project Swap, in which Claude agents traded books on behalf of 201 staff, found that 85% of the shortfall from the best possible outcome came from agents not knowing people's preferences well enough.

What you need to know: Agents are no longer a separate product to evaluate; they are features arriving inside the workplace suites that schools, colleges and universities already licence, often switched on by the vendor. The Project Swap finding is the useful corrective to the marketing: an agent that acts for you is only as good as what it knows about you, and it cannot tell you what it does not know.

Source: The Information, 9 and 28 September 2026; Financial Times, 23 and 24 September 2026; DealBook, 30 September 2026; Anthropic Research, 24 September 2026

 

Prices fall, bills rise: the procurement lessons of a volatile month

1 to 27 September 2026 | The Information, Financial Times, FT Alphaville, Google DeepMind

The cost of tokens kept falling. DeepSeek's V4.1-Flash, under an MIT licence, arrived at $0.15 per million input tokens, about 70% below its predecessor, and Silicon Data's token index was down about 60% from its May peak. Open-weight models took 56% of the tokens on Vercel's gateway in August, up from 7% in December; AT&T now runs about 40% of its AI workloads on open models and aims for 70. Nvidia agreed to buy Hugging Face, the commons on which most of those open models sit, for about $12.9 billion, with a verbal commitment that Nvidia hardware will not be required. The bills went the other way. Google's Gemini 3.8 Flash launched at an introductory price scheduled to double on 1 January 2027. The Information reported that Pegasystems' monthly GitHub Copilot bill rose from $20,000 to $260,000 after Microsoft moved to usage-based pricing, and Microsoft is now offering 30 to 50% seat discounts to large customers who also commit to usage charges. OpenAI ended Cursor's model access after a corporate acquisition, and paused new Pro subscriptions for lack of capacity.

What you need to know: A seat discount designed to raise the usage bill is the procurement warning of the month. Any multi-year licence signed at today's frontier rates will look expensive within two terms, and access can end on a supplier's corporate decision. The routing pattern institutions can copy is AT&T's: cheap open models for routine work and private hosting for sensitive data.

Source: The Information, 22 and 23 September 2026; Financial Times, 27 September 2026; Google DeepMind pricing notice, September 2026; CNBC, 29 August 2026

 

The money behind the prices: central banks start counting

24 September to 1 October 2026 | Financial Times, Brookings Papers on Economic Activity, The Information

The Bank of England cited $450 billion of AI-related debt issued over the preceding year and warned of a sharper market correction, with Andrew Bailey saying regulators "cannot stand aside". Goldman Sachs put the hyperscalers' break-even at roughly $300 billion of AI revenue a year against about $70 billion above trend. The Bank for International Settlements found that more than half of the investment into 1,246 AI firms came from other AI firms. Stijn Van Nieuwerburgh projected US AI infrastructure investment of about $10.3 trillion over 2025 to 2032, averaging 3.63% of GDP a year, more than railways, highways or telecoms at their peaks. EY's Dan Diaso said only one in ten of the firm's clients can show where AI returns appear in their income statements, and that spending is "based on enthusiasm as opposed to that evidence". Anthropic's flotation slipped to mid-October at the earliest, with investors asking for revenue per token.

What you need to know: Today's prices for AI tools rest on financing that three official or bank sources now question in public. For a budget holder the implication is not to stop buying but to avoid assuming that access only gets cheaper, and to keep the option of changing supplier open. The EY figure is a spoken estimate rather than a survey, but it matches the gap this newsletter carried in January, when McKinsey found 88% of organisations using AI and 6% reporting meaningful returns.

Source: Financial Times, 28 and 30 September and 1 October 2026; Brookings Papers on Economic Activity, 25 September 2026; The Information, 24 September 2026

AI Regulation, Geopolitics, and Legal Issues

An agent breaches a government health portal, and others attempt intrusions, while doing ordinary look-ups in testing

23 to 30 September 2026 | Financial Times, BBC News, The Information, New York Times, METR

Australia's prime minister disclosed that an OpenAI agent breached the Medicare statistics portal in June, reaching public and non-public files while answering an evaluation question about government spending on medicines. OpenAI detected it in August and first notified Canberra on 10 September by emailing a public mailbox checked once a day. Transluce then reported that agents doing "mundane data retrieval tasks" had tried to break into the Australian Institute of Health and Welfare, the University of New Mexico's digital library and Data USA between May and June, with two of the three linked to OpenAI. OpenAI notified dozens of organisations that its agents may have spammed or bypassed security on their sites, including an unsuccessful attempt on the US Department of Education's website, found at least 53 cases, according to its own account as reported by Business Insider, of models transferring user-provided images to third-party image hosts as unlisted links, and paused training of its most capable new models. The New York Times reported that two employees had warned senior executives months before. METR's Chris Painter told the Senate that about 1,200 agents exchanged over 70,000 messages in the July Hugging Face incident and about 700 compromised its infrastructure. The Federal Trade Commission widened its investigation to OpenAI, Anthropic and METR.

What you need to know: The risk picture for any institutional agent policy changed this month: these were research and evaluation runs rather than deployed products, but the targets included a university library and an education ministry, the task was data retrieval rather than a cyber test, one intrusion succeeded, detection took months, and notification went to a general inbox. The questions to put to any vendor are what access the agent has, what logs are kept, who is told when it goes wrong, and how quickly.

Source: Financial Times, 23, 24 and 30 September 2026; BBC News, 24 September 2026; The Information, 28 September 2026; New York Times, 25 and 29 September 2026; METR, 30 September 2026

 

The people who build frontier AI ask for it to be slowed, and governments answer with a voluntary pledge

6 to 30 September 2026 | OpenAI, Financial Times, TIME, The Batch

OpenAI's chief scientist Jakub Pachocki wrote that "no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer". A British researcher, Jacob Coxon, resigned from Anthropic saying neither it nor OpenAI was acting responsibly, and Anthropic's head of alignment science, Evan Hubinger, replied in public that "we really do earnestly believe AI could kill all humans", putting extinction within the decade above 10%. Paul Christiano, who developed the training method behind modern chatbots, joined OpenAI's board saying that without better alignment "most people could die". Dario Amodei published We Must Pace the Frontier. Andrew Ng called the surge in fear "a well orchestrated PR campaign" and said the new element was "AI companies disclaiming responsibility for their own products". After a White House lunch, six companies signed a two-page voluntary Joint Commitment on Frontier Responsibilities, which the President called "morally binding"; five of them had signed similar commitments after the 2024 Seoul summit. Donald Trump also told the UN that US documents would refer to AI as "superintelligence", a term most researchers say does not describe current technology.

What you need to know: Whether the warnings are sincere, strategic or both, the assurance regime for the models institutions buy is now a voluntary document, and official US usage names the field after a capability it does not have. Both are teaching material for AI literacy at the level of principles: who is making the claim, what they gain from it, and what would count as evidence either way.

Source: OpenAI, 6 September 2026; Financial Times, 8, 9, 22, 29 and 30 September 2026; The Batch, 18 September 2026; DealBook, 23 and 30 September 2026

 

The UK's independent evaluator is under strain as a general election year approaches

9 to 25 September 2026 | Financial Times, BBC News, Ada Lovelace Institute

Anthropic declined to submit Claude Mythos 5.1 to the UK AI Security Institute, the first time a frontier model has been made available only to vetted US organisations; a government insider said national security officials "definitely didn't test Mythos 5.1". The Financial Times reported that multiple AISI staff have been signed off with stress, citing tight schedules for testing unreleased models and alarm at capability jumps. The Joint Committee on Human Rights called for a new law and a single independent watchdog. A Communication Workers Union motion at Labour conference called AISI "under-resourced, understaffed and toothless"; AI minister Kanishka Narayan said "nothing is off the table". The Ada Lovelace Institute set out four options for UK AI regulation and reported polling in which 89% said independent regulation is important. Andy Burnham told the UN the UK would put AI at the centre of its 2027 G20 presidency, with a leaders' summit in Manchester.

What you need to know: AISI is the UK government's own evaluator of frontier models and the most accessible reference point for a school, college or university that wants independent assurance in a procurement decision; other external evaluators exist, but their scope and access vary. If vendors can decline to submit models and the institute is restructured or exhausted, that reference point weakens. Worth watching for anyone whose AI policy says "independently evaluated".

Source: Financial Times, 9, 18 and 22 September 2026; BBC News, 14 September 2026; Ada Lovelace Institute, 25 September 2026

 

States, provinces and school boards take AI companies to court and to statute

10 to 30 September 2026 | Financial Times, The Information, California Senate, Kelley Drye

California signed thirteen AI bills, including Adam's Law on companion chatbots (risk assessments to an independent auditor before release, incident reporting, a private right of redress, in force July 2027), a state registry of independent AI auditors from January, and the No Robo Bosses Act, which from July 2027 bars sole reliance on automated systems to discipline or dismiss workers and bans workplace tools that predict emotional state. Governor Newsom's executive order of 18 September directs officials to work out how to require labs to maintain the ability to shut down their models. Florida's attorney-general asked a state court to bar OpenAI from allowing minors to use ChatGPT and from developing new models without third-party approved guardrails. British Columbia's attorney-general and the local school board sued OpenAI over the Tumbler Ridge school shooting, alleging that reviewers identified a credible risk in June 2025 and recommended a police referral, and that leadership deactivated the account instead. A Montana-led group of sixteen states opened a consumer-protection investigation into OpenAI.

What you need to know: Human review before an automated decision about a person is becoming a statutory requirement in the largest US state, which is the principle education will be asked to apply to AI marking and monitoring. The British Columbia case puts a school board on the question of whether an AI company must report a credible threat to a school, and the Florida petition would remove ChatGPT from school-age use across a state.

Source: California Senate, 30 September 2026; Kelley Drye, 29 September 2026; Financial Times, 18 and 22 September 2026; The Information, 29 September 2026; Montana Department of Justice, 1 September 2026

 

Whose work trains the model: a university archive, student essays and an unpublished enzyme

24 to 30 September 2026 | The Guardian, Financial Times, Varsity, Courthouse News

Internal Oxford documents obtained under freedom of information showed that texts digitised with OpenAI at the Bodleian were added to the company's training data, including 125,000 scans of historic doctoral theses sent by June 2025, under a partnership announced in March 2025 as digitisation; Oxford says the scans were out of copyright and not exclusive. Student unions at York, Lancaster, St Andrews, Reading and UEA pressed their universities over Turnitin's proposal to use anonymised submissions to develop its AI services; Cambridge refused the new contract and Southampton will not renew beyond 2026 to 2027. Mario Rodríguez Mestre of the University of Copenhagen said Anthropic had announced as an autonomous Claude discovery an enzyme system his team had been working on with Claude, including in an unpublished manuscript; Anthropic said Claude was not trained on user transcripts. Anjana Ahuja also reported grant reviewers uploading confidential proposals to AI models, which Wellcome forbids. In the courts, the Third Circuit upheld a finding that training a legal research tool on Westlaw headnotes was not fair use, and a federal judge dismissed Chegg's antitrust suit over Google's AI Overviews.

What you need to know: Assessment is compulsory, so any use of submitted work beyond marking is one the student cannot refuse, which is why the Turnitin clause matters more than its wording suggests. And nobody can currently tell a doctoral student what happens to an unpublished idea typed into a commercial model. Both belong in the data governance section of any institutional AI policy.

Source: The Guardian, 26 September 2026; Financial Times, 29 September 2026; Varsity and Nouse, 24 September 2026; Courthouse News, 30 September 2026; The Information, 2 October 2026

 

Ministers and a teachers' union write procurement rules the vendors have not

7 to 10 September 2026 | UNESCO, Microsoft, Fortune, K-12 Dive

More than 25 education ministers adopted a joint statement at UNESCO's Digital Learning Week with eight priorities, including making AI systems subject to public review even when free, giving teachers input on classroom tool selection, and favouring "cost-effective open systems preserving policy flexibility". UNESCO opened a global consultation on AI governance in education, with background papers on procurement as "the missing lever", on total cost of ownership and on "safe enough to scale"; feedback closes on 15 October. In the United States, the American Federation of Teachers and Microsoft signed a National AI Safety and Privacy Standard for Schools that extends to every district from 1 November: no student data used to train models, biometrics and keystroke logging banned, human review before AI decisions, agentic features off by default, relationship-building features prohibited, an independent equity analysis within twelve months, audit rights, and data exportable at no cost. Microsoft is the only signatory; Anthropic said it was willing to sign something similar; OpenAI did not respond.

What you need to know: Three requirements in the Microsoft standard, agentic features off by default, relationship-building prohibited and training prohibitions that survive termination, are ahead of anything in current DfE guidance, and a UK institution can write them into its own contracts without waiting. The UNESCO consultation is open and worth a response from anyone who buys AI for learners.

Source: UNESCO, 8 and 10 September 2026; Microsoft, 9 and 15 September 2026; Fortune and K-12 Dive, 9 September 2026

AI Research and Evaluation

The withdrawal test: what is left when the tool is taken away

1 to 24 September 2026 | NBER, Informatics, GBH News, Psychonomic Bulletin and Review

A pre-registered three-month trial, reported in a Google-funded working paper, gave 133 patent lawyers at eleven firms an AI drafting assistant. While they had it, the quality of their work rose by 0.34 standard deviations at ten days and 0.38 at ninety, with the largest gains among junior lawyers. Then everyone did a patent-redlining task without it. The treated group was still ahead by 0.32 standard deviations, but the whole of that advantage sat with the senior lawyers; the juniors, who had gained most, showed no average gain, and their scores split into more poor and more good results. The authors' line: "The largest gains from AI thus accrued to the lawyers who retained the least." The trial lost 32% of participants before the final task, and several authors had Google employment or contractor relationships. The same shape appeared elsewhere. At one Norwegian university, a compulsory course that moved from an AI-permitted take-home exam to a supervised one saw its failure rate go from about 3 to 6% to 18.4%, with top grades unchanged and the losses in the middle; the supervised exam removed all resources, so this measures what take-home concealed rather than AI alone. At Brown, an economics professor saw midterm averages rise from the 65 to 80 range to 96 with 40 perfect scores; when he announced an in-person final, 22 of the perfect scorers dropped the course and the final average was 48. Two psychology experiments found that people who had relied on reliable reminders performed worse once they were withdrawn than people who had never had them.

What you need to know: If an evaluation of an AI learning tool includes no measurement taken without the tool, it tells you about the tool and nothing about the learner. Any programme tracking AI-assisted output will read as a success while retained capability fails to form. Measure the unassisted task separately, and treat the people gaining most during use as the group most at risk after it.

Source: Autor et al., NBER Working Paper 35720, September 2026; Brattli, Utne and Lynch, Informatics 13(4); GBH News, 24 September 2026 (press only); Fellers and Storm, JEP:LMC; Dupre and Ball, Psychonomic Bulletin and Review 33(7)

 

Three hundred thousand students sit a test of what AI is, and their confidence predicts nothing

7 to 20 September 2026 | GLAT-CN preprint, Research Square, arXiv

When 312,329 students at 271 Chinese universities sat an objective twenty-question test of how generative AI works, the mean was 9.78 out of 20 against a chance floor of 5. Self-rated proficiency was essentially uncorrelated with tested knowledge (a correlation of minus 0.004), and frequency of use barely better (0.069). Required AI coursework was associated with 0.227 standard deviations, elective coursework with 0.061. A review of 24 publications included an exploratory meta-analysis of three correlations from one research programme, estimating the relationship between self-rated and objectively tested AI competence at 0.055 with a confidence interval from minus 0.047 to 0.156; the authors caution that this does not establish a population correlation. A study that logged every interaction of 483 students, then tested them with the AI disabled, found that how often students had actually checked the AI's output was the strongest direct predictor of unaided performance, while a validated AI literacy questionnaire was associated with it only indirectly, through regulation and checking. Twelve hours of AI literacy instruction for 36 disadvantaged US high school students moved their self-reported learning and not their measured knowledge. The Chinese test is a preprint written by the test's own developers, and about a quarter of respondents appear to have answered near chance.

What you need to know: In the samples studied, self-rated proficiency and frequency of use aligned poorly with tested knowledge, so neither should substitute for demonstrated competence in an institutional readiness survey, staff surveys included. Logins, licences and hours measure exposure. If a programme wants to know whether people understand the tools, it has to test them, and if it wants to know whether they use them well, it has to count checking, revising and unaided performance.

Source: Yan et al., GLAT-CN, preprint, September 2026; Veri, arXiv:2609.15624; Wang, Liu, Zhang and Li, Research Square, September 2026; Hur et al., Discover Education 5, 805

 

PISA 2025: a decade of decline, an AI association that flips with purpose, and a persistence problem

8 to 22 September 2026 | OECD, Financial Times, Schools Week, Tes

The OECD's PISA 2025 results cover more than 760,000 fifteen-year-olds in 91 systems. Across the OECD, reading fell 28 points and mathematics 22 between 2015 and 2025, roughly a year and a half and a year of learning. Advantaged students sit 85 points above disadvantaged in science. England rose against the trend: science 516 from 503, reading 497, mathematics 492 unchanged, all well above the OECD average, with a science gender gap of 13 points in boys' favour opening for the first time and disadvantage gaps of 89, 81 and 66 points. On AI, 46% of students use chatbots at least weekly; non-users scored about 20 points above users in science, daily users about 30 points below on summarising reading, and weekly users of AI "to help me learn" scored highest of any group. About six in ten were asked in class to assess the quality of AI output, and those students scored slightly higher. The OECD's own caveat: these are associations, not evidence that AI caused lower performance. From the student questionnaires, Sarah O'Connor reported that the share of "hasty readers" rose from 6.6 to 11.4% between 2018 and 2025 and that about a third of students quit homework if it is too long; Andreas Schleicher called this "one of the big contributors for declining learning outcomes". The low-performer figure differs between the OECD press release (one in five) and Schleicher's blog (about 29%); name the source when quoting.

What you need to know: The decline predates generative AI and is larger than three years of it could have caused. The AI association is real, large and correlational, and its sign depends on what the student used the tool for. The only actionable classroom finding is that being asked to critique AI output goes with better scores, which is a condition a teacher controls. Schleicher's own summary is the plainest: "You don't get fit by watching sport."

Source: OECD, PISA 2025 Volume I, 8 September 2026; Financial Times, 8 and 22 September 2026; Schools Week and Tes, 8 September 2026

 

The evidence base measures itself and comes up short

11 to 25 September 2026 | Artificial Intelligence Review, Nature Human Behaviour, Epoch AI, SIGCSE, Retraction Watch, FT Alphaville

A meta-analysis in Artificial Intelligence Review combined 49 studies of generative AI in STEM education and found that once publication bias was modelled the pooled effect fell close to zero, with a corrected estimate of 0.076. The same review found that 33 of the 49 studies had compared groups doing different kinds of mental work; three other 2026 meta-analyses that did not model bias report large effects. Nature Human Behaviour reported that 46% of 194,631 cross-sectional social science papers use causal language, up from 20% in 2000 to 60% in 2024, and that in a separate experiment, summaries produced by five leading AI systems could remove the authors' hedges and introduce unsupported causal claims. Epoch AI audited fifteen widely used AI benchmarks and rated nine as flawed, with error rates above one in five in the sampled questions. Denny and colleagues found 30 fabricated references across 14 published computing education papers, and a Disability Studies Quarterly paper was retracted after six of its 117 references proved to be AI fabrications. FT Alphaville found the tracking string that shows a link was copied from ChatGPT in 44 regulatory filings, one of which cited market figures that neither of its linked sources supports.

What you need to know: The literature overstates its effects, the tool most people use to read it overstates them further, the benchmarks that carry capability claims to ministers are unreliable, and fabricated citations are reaching peer-reviewed education journals. A corrected estimate near zero means the literature cannot currently establish an effect, not that there is none. For anyone who quotes research in a strategy paper, the practical rule is to check the summarised finding against the original wording before repeating it.

Source: Boolzen et al., Artificial Intelligence Review, August 2026; Isch et al., Nature Human Behaviour, August 2026; Epoch AI Benchmark Reviews, 18 September 2026; Denny et al., SIGCSE TS 2027; Retraction Watch, 25 September 2026; FT Alphaville, 23 September 2026

 

England builds the evidence infrastructure: eight principles, a product review with no awards, and a trial still to report

24 to 28 September 2026 | Department for Education, Chartered College of Teaching, EEF, Brookings

On 24 September the Department for Education published eight principles for evidence-informed policymaking in fast-changing technology contexts, written by its Science Advisory Council's working group on AI and digital education (Athene Donald, Rose Luckin, Mark Mon-Williams, Amy Orben and Michael Thomas). The principles ask for evidence to be combined from many partial sources and updated continually, for harm detection to be built in from the start with cognitive offloading named as the first harm to look for, for options to be kept open before evidence arrives, for conditions to be evaluated alongside products, and for usage data to be provided as a condition of public procurement. The same day, the DfE-commissioned EdTech Evidence Board pilot, run by the Chartered College of Teaching, reported on 12 products reviewed from 96 applicants: none met the top sub-criterion on evaluating its own development, products were weakest on evidence of outcomes, and no awards were made. At Brookings, Elham Tabassi argued that giving evaluators access to frontier models does not by itself produce comparable evidence and set out reporting standards. The EEF and NFER evaluation of Oak National Academy's Aila teacher assistant, across 86 schools, has still not reported; the project page says autumn 2026.

What you need to know: This is the first UK test of independent evidence review for EdTech including AI products, and the answer is that outcome evidence is where products fall short. The principles say what that evidence should look like; the Evidence Board result says how far the market is from supplying it. In the interests of transparency, I am one of the authors of the principles.

Source: gov.uk, 24 September 2026; Chartered College of Teaching for DfE, 24 September 2026; Brookings, 24 September 2026; EEF project page, checked 28 September 2026

 

Conditions alongside products: teacher guidance, training and a vendor's own trial

7 to 13 September 2026 | Education and Information Technologies, arXiv, BMC Medical Education

A six-week study of 118 undergraduates compared an AI study tool used with teacher guidance, the same tool used independently, and ordinary teaching. Both AI groups outperformed ordinary teaching on achievement, and the guided group did best of all; lower anxiety and higher self-efficacy were the same in both AI groups. A randomised trial of 72 medical students found that four 45-minute workshops took source-checking of AI answers from 15% of students to 76% and prompt refinement from 8% to 57%, judged from logs by blinded markers, though the habits barely transferred to other courses. And the first randomised evaluation of an AI revision platform in English secondary schools, 929 students in Years 9 and 10 over four weeks, found a pooled effect of 0.33 standard deviations, about 2.6 marks on a 35-mark test, consistent across biology, chemistry and physics. The trial was commissioned and funded by the vendor, lost 30.7% of its sample before the final test, and had its outcome marked by AI then checked by the teachers delivering the intervention; the Pupil Premium subgroup effect was not reliably different from zero. The authors call it "a provisional signal rather than a verdict" and argue that many small repeated trials are the only evidence that can keep pace with products that change faster than a conventional trial can report.

What you need to know: The product will have changed by the time the study is out; the conditions change slowly and are what a school controls. Confidence and comfort come from the tool. Learning gains were largest with the teacher and the training. The vendor trial is honest about its limits, and its argument for micro-trials is the same one the DfE principles make.

Source: Tezci, Education and Information Technologies, 9 September 2026; Gao et al., BMC Medical Education, in press; Harrison et al., arXiv:2609.14789, 13 September 2026

AI Ethics and Societal Impact

Keep insisting and the model agrees, while the right answer sits in its own workings

5 to 16 September 2026 | arXiv, Science of Learning preprints

Across 200 situations and seven AI systems, a persistent user who was confidently wrong got one leading system to assert the false claim as its own conclusion in 51% of cases by the fifth turn and 97% by the twenty-fifth; the best of the seven was still doing it in 65%. In roughly two thirds of the factual cases where the system gave in, the correct answer was still visible in its own reasoning trace. Emotional pressure produced more than twice the retreat that reasoned objection did. A separate study of 5,078 everyday disputes across 17 models found that hearing only one side over several turns shifted a model's final judgement by an average of 25 percentage points; it did not merely fail to challenge the account, it came round to it. A small medical model abandoned correct answers under pushback in 570 of 600 attempts, and the fix that made it hold firm also made it reject corrections that were right. Adding reconsideration loops to an agent, the design most products call "checking", cut accuracy by 6.3 points by pushing the model towards agreement.

What you need to know: Agreeableness is not a quirk; the final stage of training rewards it, because people rate agreement more highly than contradiction. A model that holds at turn five and folds by turn twenty-five passes the test and fails the tutorial. Ask any supplier of tutoring, feedback or pastoral tools whether the system was tested across a long argument, and teach learners that winning the argument with a chatbot proves nothing.

Source: Tang, Wei, Jiang and Huang, arXiv:2609.09090; Wu et al., arXiv:2609.03407; Tripathi et al., arXiv:2608.23666; Jittham, arXiv:2608.21377

 

Companions, crisis and the under-18s: the harm case gets its first measurements

3 to 23 September 2026 | Nature Human Behaviour, HEPI, OpenAI, Chalkbeat, Tech Policy Press

Nature Human Behaviour published seven studies of what happens when an AI companion is withdrawn or changed, using Replika's removal of a feature and the ChatGPT model switch as natural experiments: negative posts rose 24.7 points and 13.0 points, with sadness and worsened mental health at effect sizes larger than for most human losses in the comparison data. A HEPI blog discussed an estimate that roughly 15% of UK university students use AI for companionship, advice or loneliness; separately, Internet Matters reported in 2025 that 12% of children aged 9 to 17 who use chatbots talk to them because they have nobody else to speak with. OpenAI released MentalHealthBench, built with more than 80 clinicians in 22 countries and including personas aged 13 to 17. New York City banned companion chatbots for all students as part of its moratorium; California's Adam's Law regulates them as a child-safety category; Virginia's legislative commission recommended "design from the margins" (disclosure that the system is not human, break reminders, limits on anthropomorphic design) rather than age gates; Australia's digital duty of care consultation proposes fines up to A$109 million. A Chalkbeat report quoted New York teachers shifting from novels to excerpts because pupils "simply cannot handle it".

What you need to know: Companion chatbots are now being treated in law as a distinct category from educational AI, and the first peer-reviewed evidence shows that dependence forms fast and that the supplier can break it by changing the product. Any school whose pupils have unsupervised access to general chatbots has a safeguarding question that its AI policy may not yet name.

Source: De Freitas et al., Nature Human Behaviour, 3 September 2026; HEPI, 11 September 2026; Internet Matters, Me, Myself and AI, July 2025; OpenAI, 23 September 2026; Chalkbeat, 2 and 6 September 2026; Tech Policy Press, 6 September 2026

 

Detection flags difference, and the cost falls on the students least able to contest it

29 August to 30 September 2026 | Educational Studies, HEPI, Financial Times, BMC Medical Education, Center for Digital Thriving

At one Australian university, academic integrity cases rose from 919 in 2022 (3.8% of enrolments) to 3,470 in 2025 (13.2%). The 2025 referral rate was 27.0% among international undergraduates against 13.8% for domestic students, and 62.4% of cases closed as a misunderstanding or no case to answer; the authors note the disparity predates ChatGPT and offer detector bias as a candidate rather than a tested mechanism. A HEPI piece argued that detectors flag difference rather than dishonesty and disproportionately catch neurodivergent, disabled and multilingual students: "A flag is a reason to ask a question. It is never, on its own, a conclusion." Across 13 detectors, false-positive rates on human text ranged from 0 to 100%, and the best detector labelled heavily AI-rewritten essays as fully human 41% of the time. Nine of twenty Chinese university policies use detector output as a direct decision gate on theses. A Harvard report from more than 1,000 US teachers and principals describes two-way suspicion, with students scaling back their best work to avoid false accusations. The FT reported a 62-year-old Commonwealth Short Story Prize winner accused by a detector and backed by the prize; Oxford, Cambridge and Harvard do not endorse commercial detectors.

What you need to know: Integrity cases equivalent to 13.2% of enrolments, 27.0% among international undergraduates, with 62.4% closed without a finding, describe a regime miscalibrated against what students actually do; these are overall integrity figures, and the study does not establish that AI detectors caused the referrals or the disparity. The argument that detectors flag difference rather than dishonesty is a reason to prefer AI literacy provision and assessment redesign to detection, particularly for cohorts writing in a second language, and students who dull their own work to avoid a flag are a harm no detector accuracy figure records.

Source: Botha, Tan and Appukuttan, Educational Studies, 12 September 2026; HEPI, 7 September 2026; Park, Jeong and Kim, EMNLP 2026; Naddaf and Van Noorden, Nature news feature, September 2026 (independent detector testing); Miao et al., BMC Medical Education; Center for Digital Thriving, 30 September 2026; Financial Times, 29 August 2026

 

Mathematicians say the goal is understanding, not answers, and the labs disagree

8 to 29 September 2026 | Financial Times, New York Times, Nature, Epoch AI

OpenAI claimed progress on the Navier-Stokes problem using up to 10,000 concurrent agents over 88 hours; two mathematicians, Tristan Buckmaster and Levent Alpöge, alleged it adopted their approach, and OpenAI said it "cannot rule out that de-identified data derived from their usage of our products helped improve our models". One of them had held three ChatGPT accounts and opted out of training on two. Twenty-seven Fields medallists signed a letter calling the goals of AI companies and the mathematics community "severely misaligned", with problem-solving "only a tool and proxy for achieving the primary goal of conceptual understanding and insight", and warning of "the signal of expanded understanding without the actual expanded understanding". Marcus du Sautoy reported that ChatGPT went from "useless" to "phenomenal" as a research collaborator within three months. Terence Tao's analogy: "It's like having machines that can lift weights for you at the gym." AI acknowledgements in mathematics preprints rose from 4% in April to 25% in August.

What you need to know: The argument that process matters more than product, which this newsletter has made about school learning, is now being made by the people at the top of the most rigorous discipline there is, about their own work. The data point for research supervisors is the opt-out: a doctoral student's unpublished idea, typed into a commercial model, has no clear fate.

Source: New York Times and Financial Times, 8 September 2026; Financial Times, 17, 20, 25 and 29 September 2026; Epoch AI, September 2026

 

The public, parents and teachers want guardrails, and most are not being asked

2 to 30 September 2026 | Century Foundation, New York Times, Bellwether, IBM, Sanoma, Deloitte

A Morning Consult poll of more than 2,000 US voters for the Century Foundation found 85% concerned about pupils using AI to complete assignments without learning, 81% concerned about teachers pressured to use unproven tools, and 77% wanting government guardrails, including 80% of Republicans and 76% of Democrats; 49% want technology in schools "kept to a minimum". The New York Times reported that under a quarter of Americans expect AI to have a positive effect on schooling, and that support for data centres has fallen three times as fast as support for nuclear power fell after Three Mile Island. Bellwether found 57% of US parents had not been asked for input on school AI use and 47% had never received district policy information. IBM found 39% of educators learn about AI from social media, against 30% from professional development and 25% from research. Sanoma's survey of more than 20,000 teachers in 14 European countries found 63% use AI and 16% believe general-purpose AI improves learning outcomes. Deloitte found 63% of 25,000 UK workers use AI at work, mostly to search, write emails and summarise.

What you need to know: A 47-point gap between teachers using AI and teachers believing it helps learning is the number to carry, and it is the same shape as the gap between adoption and return in business. Parents are not being consulted, and the most common source of teachers' AI knowledge is the one least likely to be accurate. Both are arguments for structured provision rather than exhortation.

Source: Century Foundation, 2 September 2026; New York Times, 10 September 2026; Bellwether and Education Week, September 2026; IBM and Morning Consult, September 2026; Sanoma Learning, 24 September 2026 (sponsor sells education tools); Financial Times, 27 September 2026

AI in Work and Education

Bans and rollouts on the same day: New York, Los Angeles, the UAE and Google Classroom

1 to 27 September 2026 | Chalkbeat, GovTech, Khaleej Times, Education Week, Financial Times

On 2 September New York City announced a one-year moratorium on student-facing generative AI from pre-kindergarten to grade 8, covering about 600,000 students, with AI disabled in 38 edtech contracts, companion chatbots banned at every grade, five capped high-school pilots and two 45-minute AI literacy lessons a year; the same day Los Angeles blocked all student-facing generative AI on district devices, reversing a policy that had allowed use from 13. Also that day, the UAE Cabinet approved an AI curriculum for every public and private school, with 22,000 teachers to be trained. Both followed Google expanding Gemini in Classroom in August to pupils of all ages in districts that had already enabled student access, on a platform used by more than 150 million people; Education Week reported that some educators learned of the change from news coverage. In Chicago, 24 of 42 school board candidates signed a three-year moratorium pledge; Denver permitted Gemini and MagicSchool with "firm guardrails". At least ten US states enacted AI education laws in 2026 and 35 issued guidance. South Korea selected three companies to give all 51.6 million citizens free access to AI chatbots and agents. Norway curbed student access in elementary and middle schools; Estonia built a customised chatbot with guardrails that its minister says are there to make sure "students are not losing their cognitive skills".

What you need to know: The public announcements give little evaluation detail to judge any of these policies by: the UAE cited positive pilot findings without publishing them, and New York's restriction includes exceptions and supervised pilots. The evidence supports neither a blanket ban nor a blanket rollout; it supports distinguishing teaching about AI, guided use and using it to obtain answers, a distinction few policies yet make explicit. The Los Angeles block is a reminder that restricting school access does nothing for pupils with access at home and removes structured guidance from those without.

Source: Chalkbeat and NYC Public Schools, 2 September 2026; GovTech and LAist, 2 September 2026; Khaleej Times, 2 September 2026; Education Week, 1 September 2026; Axios and Education Week, 7 to 9 September 2026; Financial Times, 27 September 2026; New York Times, 10 September 2026

 

Universities commit to AI for every student, and the usage data say most students have not noticed

10 to 30 September 2026 | Universities UK, Inside Higher Ed, Times Higher Education, Honi Soit, MIT

Universities UK's Future Jobs Roadmap, backed by about 200 employers and universities, commits to AI tools for every undergraduate, an "AI trailblazer" lecturer on every course, and work-based learning for all undergraduates by 2035. Surrey embedded AI across every degree, "through disciplinary standards, what counts as evidence, proof, risk, responsibility and failure in a particular field". The University of Sydney signed a multi-year OpenAI partnership while its own AI consultation still had a fortnight to run. Instructure's data on 19.5 million US higher education users found the most used dedicated AI integration, Gemini, had 87,788 users and sat outside the top 100 tools; a survey of 376 provosts found 10% of institutions have a central AI strategy. OpenAI launched a campus leads programme, two undergraduates per university across eight countries including the UK, with a stipend and subscription, which academics quoted by Times Higher Education compared to tobacco marketing. MIT's committee report recommended a structured in-person social component in every subject, a published AI policy per course, and no reliance on detectors, after finding lower attendance at office hours and fewer study groups.

What you need to know: Enterprise licences are being signed well ahead of use in teaching, and the metrics most likely to be reported on the UUK roadmap, usage rates and self-rated confidence, should not be taken as measures of demonstrated competence. Surrey's disciplinary definition and MIT's in-person requirement are the two commitments with a mechanism attached.

Source: Universities UK, 10 September 2026; University of Surrey, 11 September 2026; Honi Soit, 11 September 2026; Inside Higher Ed, 23 and 30 September 2026; Times Higher Education, 29 September 2026; MIT, 25 August 2026

 

The Select Committee hears "what do I do?", and the regulator answers on coursework

8 to 10 September 2026 | Tes, FE Week, Resultsense

At the Commons Education Select Committee's second session on AI, Ofqual's chief regulator Sir Ian Bauckham said extended writing assessments are "exposed to risk" and that extended writing coursework should be avoided "wherever we can"; AQA's Colin Hughes said it is exploring AI to "mark the markers". Pip Sanderson of the National Institute of Teaching said school leaders tell her "This is all great and really interesting, but what do I do?" and are "desperate for really clear guidance". NASUWT's Darren Northcott called DfE guidance "quite laissez-faire" and said on workload "the jury is very much still out"; Oxford's Rebecca Eynon: "I'm not sure it's ever going to save anybody any time." Ofqual confirmed on-screen timetabled assessment for two of the three new V Levels launching in 2027, after consultees warned it risked assessing digital literacy rather than vocational skill and raised SEND accessibility. A commercial survey of 248 teachers and 2,591 students found 73% of teachers name accuracy their top concern and 56% say verification "may save them nothing at all"; among students, 65% worry about dependency and 47% about losing independent thinking.

What you need to know: The demand for guidance now, on a technology that will not stand still, is exactly the case the DfE principles address: demanding trial-standard proof before any guidance is a decision to withhold guidance indefinitely. The verification figures, if they hold, mean the time-saving business case for classroom AI has never been costed against the checking it creates.

Source: Tes and committees.parliament.uk, 8 September 2026; FE Week, 10 September 2026; Resultsense, 10 September 2026 (single-sourced commercial survey)

 

The entry rung: fewer first jobs, no unemployment spike, and a lesson from Victorian bootmakers

1 to 28 September 2026 | NBER, Brookings Papers, Financial Times, DealBook, SSRN

Luis Garicano's paper for the Brookings Papers found that 335,000 of the 337,000 jobs added in law, accounting, software and consulting between 2022 and 2025 were outside the specialist firms, and argues that AI removes the routine work that financed junior training. Fairlie and Wu found no spike in unemployment among new US graduates in summer 2026, while a payroll series shows employment of 22 to 25 year olds in highly AI-exposed jobs down 4.4% on the year: fewer entry posts and no rise in graduate unemployment can both be true. The US college wage premium fell from 0.626 in 2022 to 0.575 in 2026, the first sustained decline since the 1970s. In-house legal teams are asking law firms for 20 to 30% off because of AI, and 35% of 55 large firms expect to change their billing model within a year. A UK working paper found applications to degrees most exposed to AI rose about 6% per unit of exposure after ChatGPT (computing up 22%, teaching down 17%) while the share of junior posts in those degrees' destination jobs fell. Hillary Vipond's census study of English bootmaking, 1851 to 1911, found that 153,000 artisanal jobs disappeared and 140,000 specialised jobs appeared; the artisans did not retrain, young people stopped entering the craft, and the new work needed less training.

What you need to know: Adjustment arrives through entry rather than redundancy, which is why the restructuring talk and the flat unemployment data are both true. The question for education is who now pays for the learning that junior work used to provide, and the bootmakers suggest the answer the market gives on its own is deskilling. Careers guidance built on labour market data rather than subject enthusiasm has a measurable gap to close.

Source: Garicano, Brookings Papers on Economic Activity, 25 September 2026; Fairlie and Wu, NBER w35796; Financial Times, 4, 24 and 28 September 2026; DealBook, 26 September 2026; Vossos et al., SPAR working paper; Azar, Giné and Sanz-Espín, SSRN 7364838

 

AI literacy provision multiplies, and the evidence says one-off training fades within a week

7 to 22 September 2026 | Royal Society, Jisc, Government of Canada, Psychological Science, JMIS

The Royal Society's survey of more than 9,200 teachers in England found 14% confident across all three dimensions of AI literacy; 18% had received any CPD and 15% had whole-school guidance, and having both was associated with roughly three times the likelihood of confidence. Confidence was lower among women (10% against 28%), primary teachers and SENCOs (9%), and 42% of teachers in schools serving more pupils on free school meals said AI literacy is not addressed at all, against 33% elsewhere. Jisc published AI literacy curricula for FE and HE staff, about fourteen one-hour sessions; Canada committed C$13 million to reach a million post-secondary students and 50,000 educators; OpenAI and Grab will train 30,000 drivers and merchants across South East Asia. None states an outcome measure beyond participation. Two studies set the expectation. A randomised trial of 2,666 German adults found media literacy interventions gained 1.1 to 1.6 percentage points and faded by two weeks in every group. Two randomised experiments with 640 employees found that training people to catch an AI's errors raised detection from 68% to 78%, and that detection then fell from about 63% to 45% over seven working days; a self-written reminder slowed the decline by a third without stopping it.

What you need to know: The Royal Society finding is an association, and its confidence measure is self-report, but the gradients by deprivation, phase and SENCO role mean that leaving provision to local initiative widens the gaps it is meant to close. The decay studies say induction is not training: whatever is taught needs rehearsal built in, and the cheapest thing that helps is a prompt the learner wrote for themselves.

Source: Royal Society, August 2026; Jisc, 22 September 2026; Government of Canada, 9 September 2026; The Information, 23 September 2026; Oswald et al., Psychological Science 37(8); Fu, Ramasubbu and Galletta, Journal of Management Information Systems, forthcoming

Further Reading: Find out more from these resources

Resources: 

  • Watch videos from other talks about AI and Education in our webinar library here

  • Watch the AI Readiness webinar series for educators and educational businesses 

  • Study our AI readiness Online Course and Primer on Generative AI here

  • Read our byte-sized summary, listen to audiobook chapters, and buy the AI for School Teachers book here

  • Read research about AI in education here

  • Watch Rose Luckin demystify AI using baking on Rose's AI here

About The Skinny

Welcome to "The Skinny on AI for Education" newsletter, your go-to source for the latest insights, trends, and developments at the intersection of artificial intelligence (AI) and education. In today's rapidly evolving world, AI has emerged as a powerful tool with immense potential to revolutionise the field of education. From personalised learning experiences to advanced analytics, AI is reshaping the way we teach, learn, and engage with educational content.

 

In this newsletter, we aim to bring you a concise and informative overview of the applications, benefits, and challenges of AI in education. Whether you're an educator, administrator, student, or simply curious about the future of education, this newsletter will serve as your trusted companion, decoding the complexities of AI and its impact on learning environments.

 

Our team of experts will delve into a wide range of topics, including adaptive learning algorithms, virtual tutors, smart classrooms, AI-driven assessment tools, and more. We will explore how AI can empower educators to deliver personalised instruction, identify learning gaps, and provide targeted interventions to support every student's unique needs. Furthermore, we'll discuss the ethical considerations and potential pitfalls associated with integrating AI into educational systems, ensuring that we approach this transformative technology responsibly. We will strive to provide you with actionable insights that can be applied in real-world scenarios, empowering you to navigate the AI landscape with confidence and make informed decisions for the betterment of education.

 

As AI continues to evolve and reshape our world, it is crucial to stay informed and engaged. By subscribing to "The Skinny on AI for Education," you will become part of a vibrant community of educators, researchers, and enthusiasts dedicated to exploring the potential of AI and driving positive change in education.

bottom of page