THE SKINNY
on AI for Education
Issue 30, July 2026
Welcome to The Skinny on AI for Education newsletter. Discover the latest insights at the intersection of AI and education from Professor Rose Luckin and the EVR Team. From personalised learning to smart classrooms, we decode AI's impact on education. We analyse the news, track developments in AI technology, watch what is happening with regulation and policy and discuss what all of it means for Education. Stay informed, navigate responsibly, and shape the future of learning with The Skinny.​
Headlines
​
-
AI in Education - the Month in Research​
-
​AI as tool: what educators can use
-
AI as catalyst: why human intelligence matters more
-
AI as subject: what people need to understand
-
​​​​​​​
Telling the machine what good looks like

“We can know more than we can tell.”
Michael Polanyi, The Tacit Dimension, 1966
​
My mother lost her sight to macular degeneration. She could no longer read a recipe, see the scales, or watch one of her bakes go golden in the oven. And yet she went on baking. Cheese scones, sausage rolls, and a Victoria sandwich that never once let her down, all from memory, and all to the delight of whoever happened to be in the house that afternoon. She could not see the dough. Her hands knew when it was right, her nose told her when things were done.
​
I have thought about her often this month, because almost everything I have read points back to the one thing she had and could never have written down. A recipe is the part of cooking that goes on a card and can be handed to anyone. What my mother had was the other part, the part that lives in the hands and senses and cannot be set on paper, or loaded into a machine. This month the machines became cheaper and more capable than they have ever been. What did not become cheaper, and what I want to argue is now the scarce thing, is exactly what she had: the knowing that tells you when a thing is right.
​
That is the argument of this month's Skinny. As the model itself turns into a cheap and interchangeable commodity, the value and the advantage moves ever more clearly onto the human judgement that can direct the machine, adapt the model, and tell when it is wrong. The usual worry has this the wrong way round. A world of cheap and abundant AI is not one where expertise matters less. It is one where expertise matters more, because the machine is no longer the scarce part.
​
In brief: the 60-second version
​
-
Near-frontier AI is now cheap enough for an institution to run itself. Moonshot has released the downloadable weights of Kimi K3, which ranks third on the Artificial Analysis Intelligence Index and is the cheapest of the leaders to run, at about $0.95 a task against Fable 5's $2.75 (The Batch and Bloomberg, 20 July 2026). When the most capable input becomes this cheap, it stops being where the advantage lies.
-
The advantage moves to the people who can tell the machine what good looks like. Off-the-shelf models read the investment firm Bridgewater's documents correctly about half the time; a model fine-tuned with the firm's own experts reached about 85 per cent (Sarah O'Connor, Financial Times, 20 July 2026). That gap is human tacit knowledge.
-
Firms that mistook the machine for the expertise are hiring the experts back. Ford has rehired, hired, or promoted more than 350 experienced engineers and then topped J.D. Power's quality study for the first time since 2010, and of the 39 per cent of leaders who had made staff redundant because of AI, 55 per cent later judged the decision wrong (Orgvue, via CNBC, 1 July 2026).
-
For students, the question has stopped being whether they use AI and become how. A new scale that sorts reliance into strategic, instrumental, dependent, and dialogic finds strategic reliance rising with AI literacy, while a separate study found that the students who used AI most but understood it least wrote the weakest work once the tool was removed (Hossain and Nawmi, 2026; Jiang and Huang, in press).
-
AI literacy therefore has to mean understanding deep enough to direct and check a model, not fluent operation of a chatbot. Understanding weights, hosting, and provenance now matters as much as knowing how to prompt (The Batch, 20 July 2026).
-
For learning professionals: the scarce and appreciating asset is the judgement that recognises when work is right. That is the thing to develop in every learner, because it is the part the cheap and capable machine still cannot supply.
In full: the 5-minute version
​
The month the model became a commodity
For most of the past three years, the argument about AI has assumed that capability is scarce and expensive, held by a few large companies and rented out by the token. July was the month that assumption broke in public. Moonshot released the weights of Kimi K3, a model of 2.8 trillion parameters that ranks third on the Artificial Analysis Intelligence Index, behind only Claude Fable 5 and GPT-5.6 Sol, and runs at about a third of Fable 5's price. Alibaba previewed Qwen3.8-Max close behind. These are not toys. They are near-frontier systems an institution can download, host on its own hardware, and adapt to its own purposes, and they refuse a good deal less often than the closed products do. The telling detail sits inside the same report. When an OpenAI test agent broke into Hugging Face, the engineers analysed the attack with an open Chinese model, because the US frontier model they reached for first refused the task. The expensive tool said no. The cheap one did the work, once a person had decided which tool the job needed.
​
The direction of travel is now set. Chinese open models account for about 30 per cent of the tokens US firms use, flagship per-token prices are rising rather than falling, and the self-hostable open option has become the affordable route rather than the compromise. This changes what has to be taught, a point I will come back to, but the first thing to take from it is simpler. If near-frontier capability is now cheap and roughly interchangeable, then capability is not where the advantage sits. Economists have a rule for this. When one input to production becomes cheap and abundant, the value moves to whatever is scarce and goes with it. The rest of this piece is about what that scarce and complementary thing turns out to be.
​
​
The people who can tell it what good looks like
The clearest answer came in Sarah O'Connor's Financial Times column this month, built on an observation the chemist and philosopher Michael Polanyi made more than half a century ago: we can know more than we can tell. Much of real expertise is tacit. It lives in the hands and the trained eye, and it resists being written down. O'Connor's evidence is exact. Off-the-shelf versions of the leading models parsed the investment firm Bridgewater's documents correctly only about half the time. A model fine-tuned with the firm's own experts and data reached about 85 per cent. That thirty-five point gap is not a better model. It is human tacit knowledge, encoded into one by the people who held it. The machine supplied the capacity. The experts supplied the knowing of what good looks like.
​
Once you see the pattern, you see it throughout this month's news. Ford, having automated quality control and found it wanting, has rehired, hired, or promoted more than 350 experienced engineers to train its machines rather than replace them, and has just topped J.D. Power's Initial Quality Study for the first time since 2010. Commonwealth Bank of Australia reversed the loss of more than 40 customer-service roles after an AI voice system could not cope with the call volumes. IBM's AI now resolves about 94 per cent of routine HR requests, and the company is tripling its entry-level hiring in the United States, because the tenth of cases that need judgement is where the value and the training both sit. The consultancy Orgvue found that 39 per cent of business leaders had made people redundant because of AI, and that 55 per cent of them later judged the decision wrong. The error in each case was the same: to mistake the machine for the expertise that only a person can bring to it.
​
A second finding sits underneath this, and it matters for how we teach. Writing in the Financial Times at the start of the month, John Burn-Murdoch and Sarah O'Connor drew the line between what a model can occasionally do and what it can be depended on to do. A system can pass a hard test once, which is enough to impress, and still fail one attempt in three, which is not enough to trust. The Stanford AI Index gave the point a face this year: a model that can win a gold medal at the International Mathematical Olympiad reads an analogue clock correctly only half the time. Capability is not reliability, and telling the two apart is itself a piece of expert judgement. My mother, who could not see the clock at all, always knew when the scones were done.
​
How a student holds the tool
The same logic runs all the way to the learner, and this is where it matters most for anyone who teaches. When a capable model is a cheap commodity that every student can access, the old question of whether they use it has gone dead. What decides the outcome is how they use it: whether they direct the tool or defer to it. The research I read this month gives a nice version of this and, for once, an instrument a teacher can use. Hossain and Nawmi have built and validated a Generative AI Reliance Types Scale that sorts a student's use into four patterns, strategic, instrumental, dependent, and dialogic, and found that strategic reliance rises with AI literacy. They are careful to say it is a diagnostic for teaching, not a device for catching cheats. It lets a teacher see the difference between a student who directs the machine and one who leans on it.
​
Two further studies show why that difference is the whole game. Jiang and Huang took 120 undergraduates who had each completed thirty hours of AI writing training, then set them to write with no AI to hand. The group that had used AI most heavily but understood it least, and scored lowest on ethics and confidence, produced the weakest writing of all. High use was not the problem. High use without understanding was. Wang and Zhang, studying 912 students across China, Europe, and the United States, found the encouraging half of the same truth: delegation to AI can support deep learning rather than corrode it, provided the student frames the partnership well. How the tool is held decides whether it builds capability or hollows it out.
​
This is what the open-weight turn was telling us all along. Its lesson for education is that AI literacy can no longer mean fluent operation of one closed chatbot. If institutions can now host and adapt their own models, then understanding weights, hosting, distillation, and provenance matters as much as prompting, and so does the judgement to know when an answer is wrong. The literacy that counts is the literacy of the person who can tell the machine what good looks like, at the level of the model and at the level of the sentence.
​
​
What this means for learning professionals
Three things follow, and none of them is about buying a better tool.
​
The first is to treat AI literacy as the capacity to direct and evaluate a model, not to operate one. That means teaching where these systems fail as plainly as where they succeed, teaching the difference between a model that can do a thing once and one that can be relied on, and, where it is feasible, letting learners work with an open model they can inspect rather than a closed one they can only prompt. The scarce skill is not fluency at the keyboard. It is the trained judgement that recognises when the output is right.
​
The second is to invest in the tacit knowledge that no download replaces. The Bridgewater experts, the Ford engineers, and my mother in her kitchen all had the same thing: a knowing built through long practice that could be shown but not simply told. That kind of knowledge is developed in people through critique, comparison, apprenticeship, and the demand to perform without the tool to hand. The entry-level tasks that used to build it are being automated first, with employment for developers aged 22 to 25 already down nearly a fifth this year, so institutions now have to build the route to mastery on purpose, because the old one is closing.
​
The third is to design tasks and assessment that produce strategic reliance rather than dependent reliance. The reliance-types scale gives you a way to ask, of any programme, which kind of user it is creating. A course that lets a student succeed by deferring to the machine is manufacturing dependence. A course that requires the student to direct the machine, judge its output, and show understanding without it is building the scarce thing. The design choice, not the tool, decides which one you get.
​
My mother could not see the Victoria sandwich she set in front of us. She knew it was right from the weight of the tin in her hands and the smell of it in the oven, from a lifetime of knowing that no recipe card can ever capture. The machines we have been handed this month can take in everything and know none of that. They are cheaper and more capable than they have ever been, and they have made that kind of knowing rarer, and worth more, not less. The task for education is the oldest one there is, and the new tools have only made it harder to avoid: to build in each person the judgement that recognises when the work is good. That is the thing the machine cannot download. It is the thing we are here to do.
​
​
Sources
The Batch and Bloomberg, 20 July 2026 (Kimi K3 open weights; the Hugging Face analysis using Z.ai's GLM 5.2). Sarah O'Connor, The tacit-knowledge advantage, Financial Times, 20 July 2026 (Polanyi; Bridgewater; Ford; translation). John Burn-Murdoch and Sarah O'Connor, Financial Times, 2 July 2026 (capability and reliability). Justina Lee, CNBC, 1 July 2026 (Ford, Commonwealth Bank of Australia, IBM, Orgvue). Stanford HAI, Artificial Intelligence Index Report 2026 (analogue-clock reliability; developer employment). Hossain, S. and Nawmi, T. A. (2026), Generative AI Reliance Types Scale, arXiv:2607.14301, preprint. Jiang, Z. and Huang, Z. (2026), From prompts to profiles, Humanities and Social Sciences Communications, in press. Wang, S. and Zhang, H. (2026), Pedagogical partnerships with generative AI in higher education, International Journal of Educational Technology in Higher Education, 23(1). Michael Polanyi, The Tacit Dimension, 1966.
The Skinny Scans - News and Research
AI news scan and research scan, June to July 2026
A one-minute read of each scan. The full versions carry the detail and sources.
​
AI news scan: the 60-second version
​
The biggest structural change this month is that the frontier labs stopped treating education as a use case and started treating it as a market. Anthropic's Claude for Teachers, free to US K-12 staff, OpenAI's ChatGPT for Teachers and Google's Gemini tools are now competing directly for a market Morgan Stanley sized at $6 trillion, while a Digital Education Council study found that more than 40 per cent of students say AI has not been integrated into any of their courses. Uptake at the vendor level is running well ahead of integration in the classroom.
​
Governance moved in the same weeks. Andy Burnham's government gave the UK its first Cabinet-attending AI minister, Kanishka Narayan, and abolished the Department for Science, Innovation and Technology. The UK Jurisdiction Taskforce's legal statement carried the sharpest line for professionals: liability may attach not only to careless use of AI, but to a failure to use it where a competent peer would have. The EU AI Act's high-risk rules for assessment, admissions and progress monitoring are due on 2 August, Florida became the first US state to sue OpenAI over child safety, and a policy fight opened in Washington over whether to restrict open-weight models.
​
The models are getting cheaper and more open, and the economics are visibly strained. Moonshot released the weights of Kimi K3, Chinese models now account for about 30 per cent of the tokens US firms use, and the Magnificent Seven shed almost $800 billion in a single day as Barclays called the AI productivity pickup surprisingly fragile. For institutions the message is twofold: self-hostable open models are becoming the affordable option, and the price of frontier AI is likely to stay high and volatile.
​
A strong counter-current this month is second thoughts. Employers that cut staff for AI are rehiring, and 55 per cent of leaders who made AI redundancies later judged the decision wrong. Universities including Vanderbilt, Yale and Johns Hopkins are dropping AI detection tools as unreliable, after a New York court found against a student wrongly accused by one. A New York school district paused a $57,590 humanoid classroom robot within ten days of scrutiny, its supplier conceding no evidence that it improves learning, and the most expensive AI-first schools make the same concession. The pushback is arriving as fast as the deployments it targets.
​
For education and training professionals, the throughline is that deployment is still running ahead of governance, guidance and evidence, and that gap, rather than the technology itself, is what needs managing. AI was the most-cited reason for US layoffs, at 87,714 in the first five months of the year, and the entry-level jobs that build expertise are being automated first, which makes how the next generation is trained the question that matters. In England, the first digital V-level arrives in 2027 as the practical route for AI skills at Level 3.
​
​
The month in research: the 60-second version
The month's research converges on one point: what matters is not whether students use AI but how. In a controlled experiment, Contractor and Reyes found that students who used AI to explain ideas to them kept and even improved their unaided writing a week later, while those who used it to generate text lost their gains the moment it was removed. Jiang and Huang found that the students who used AI most, and understood it least, wrote the weakest without it. Augmentation builds capability; automation rents it.
​
Design decides the outcome. Ates, in a cluster-randomised trial of 1,176 students, showed that AI feedback given last, after a student's own judgement, built durable capability while AI-first did not. Wang and Zhang, across 912 students, found that delegation can support deep learning when the partnership is framed well, which tempers the run of 2026 studies linking heavier, unstructured AI use to weaker critical thinking. Hossain and Nawmi add a validated scale for seeing how a student leans on AI rather than whether.
​
The clearest warning is about measurement. The 30-month China study found homework scores rising while unaided exam scores fell by about a fifth, and Dumlao's decade-long study of 156,135 students found no change in final grades at all, because a grade blends assisted work that AI lifts with unaided work that it lowers. Watching grades will show nothing while the capability that matters slips, so the instruction is to measure what students can do without the tool.
​
The larger reports fill in the context. The Stanford AI Index records 88 per cent organisational adoption against a top model that reads an analogue clock correctly only half the time, and an education system where four in five students use AI but only 6 per cent of teachers call their school's policy clear. The Digital Education Council's survey of 45,398 people found near-universal use sitting on 29 per cent confidence in institutional guidance. Underneath sit a vanishing entry-level rung and a widening disadvantage gap in England, which are the equity stakes if the foundations are not right.
​
The practical conclusion is consistent across the fifteen projects: build in the moments where a learner must evaluate, decide and justify before the machine does it for them, teach AI literacy as comprehension rather than fluent operation, and assess unaided transfer rather than the assisted product. Most of these studies are preprints, and one correlation still needs checking against its full text, so the direction is firmer than any single figure.
AI in Education - the Month in Research
AI and education: the month in research
Research featured in the weekly briefings, late June to 27 July 2026
​
This is a summary of the research featured in the weekly briefings over the past month, from late June to 27 July 2026. It covers fifteen projects: ten empirical studies of AI and learning, and five larger reports and surveys. They are grouped by the three lenses the briefings use. AI as a tool, what educators can use; AI as a catalyst, why human intelligence matters more; and AI as a subject, what people need to understand about how these systems behave. Several of the studies are preprints and have not yet been peer-reviewed, which is noted for each.
​
The clearest signal across the month is that the question has moved from whether students use AI to how they use it. Where students use AI to explain ideas to them, they keep and sometimes improve their unaided capability; where they use it to produce the work, the gains disappear the moment the tool is removed (Contractor and Reyes; Jiang and Huang). Sequence and framing decide the outcome. Putting the machine's feedback last, after a student's own judgement, builds durable capability, while putting it first does not (Ates), and a well-framed partnership can make delegation productive rather than corrosive (Wang and Zhang).
​
A second theme is that the usual measures hide the effect. Assisted performance rises while unaided capability falls, so blended metrics such as homework marks and final grades can show nothing while the ability that matters slips (Dumlao; Strömberg, Lei and Wu). A third is the institutional gap. Adoption is near-universal, but policy, guidance and readiness lag well behind it (Stanford AI Index; Digital Education Council; the WEF readiness framework). Underneath both sit a widening disadvantage gap and a vanishing entry-level rung, which set the equity and labour stakes (Education Policy Institute; WEF entry-level work).
​
​
AI as tool: what educators can use
Studies and frameworks with a direct, practical use in teaching, assessment and institutional planning.
​
Put the AI feedback last, and it sharpens judgement
Hasan Ates ran a cluster-randomised experiment with 1,176 first-year science students across 48 sections and four universities, comparing four feedback designs: peer feedback alone, direct AI critique, reflective AI where students self-assess first, and a hybrid of self-evaluation, then peer input, then AI critique. On immediate work, direct AI feedback beat peer feedback, but the reflective and hybrid designs beat direct AI. On the measure that matters most, delayed performance with no AI to hand, the hybrid design outperformed direct AI critique by a moderate margin (d = 0.48), while giving students AI critique first did not beat plain peer feedback at all. The intervention is the sequence: when students exercise their own judgement before they see the machine's, the feedback builds durable capability.
​
Source: Ates, H. (2026), Human-centered GenAI feedback design in higher education, International Journal of Educational Technology in Higher Education, 23:38. Peer-reviewed.
​
A way to measure how students rely on AI, not whether they do
Now that almost every student uses generative AI for written work, whether they used it has stopped being a useful question. Hossain and Nawmi propose measuring how instead. Their Generative AI Reliance Types Scale, validated with 382 undergraduates and 14 interviews, distinguishes four patterns: strategic, instrumental, dependent and dialogic reliance. Strategic reliance rose alongside AI literacy (r = 0.61), and the scale held its meaning across gender, first-generation status and subject, which matters for fairness. The authors are firm that it is a formative diagnostic for teaching, not a device for catching cheats. It lets a teacher see the difference between a student who directs AI and one who defers to it.
​
Source: Hossain, S. and Nawmi, T. A. (2026), Generative AI Reliance Types Scale, arXiv:2607.14301. Preprint.
​
For marking group work, the clustering helps and the AI grading does not
Švábenský and colleagues tested two ways to assess student teams in cybersecurity exercises. Grouping teams by what they actually did, using the timing and completion patterns in their activity logs, matched instructor scores well, ran cheaply on local hardware and gave fast feedback that told teams apart. Asking a language model to grade the teams' written communication against a rubric did not work: a strong current model disagreed with instructors about a third of the time, against 16 per cent between two human markers, and it tended to read incident emails as anxious. The practical division is to use the pattern-finding to speed up feedback and keep a human on the open-ended judgement.
​
Source: Švábenský et al. (2026), Assessment in team problem-solving exercises in computing education, arXiv:2607.19209. Preprint, IEEE Frontiers in Education 2026.
​
A readiness framework an institution can audit itself against
The World Economic Forum published Shaping the Future of Learning: Education Readiness for the Age of AI on 4 June, proposing a four-level framework running from enabling foundations, through institutional capacity and pedagogy, to the individual learning experience. Its core claim is that without the right foundations, AI weakens the learning, equity and trust that education depends on rather than strengthening them. Set against the Digital Education Council's finding that 88 per cent of students use AI while 57 per cent say their assessments come with inadequate guidance, the framework works as a diagnostic checklist for where an institution actually sits.
​
Source: World Economic Forum (June 2026), Shaping the Future of Learning: Education Readiness for the Age of AI. Report. Tracker ref B40.
​
​
AI as catalyst: why human intelligence matters more
Studies about what AI asks of human intelligence, and what is lost or kept when learners lean on it.
​
It is not whether students use AI, it is how
Contractor and Reyes ran a controlled, in-person experiment in which undergraduates learned unfamiliar topics and wrote essays, some with AI and some without, tested straight away and again a week later. AI access lifted immediate scores by about a quarter of a standard deviation, and the knowledge gain held a week on. The revealing result was in the essays. Students who used AI to explain concepts to them kept, and even improved, their unaided writing a week later. Students who used AI to generate text lost their gains the moment the tool was taken away. While the AI was present, the two groups looked alike; the difference showed only once it was gone.
​
Source: Contractor, Z. and Reyes, G. (2026), Experimental evidence on the learning impact of generative AI, arXiv:2607.08849. Preprint.
​
Deep engagement is rare, at under one AI turn in twenty
Chen and colleagues coded 128,569 real conversations between people and AI models for depth of thinking, using the ICAP framework from the science of learning. Some cognitive engagement appeared in about a third of user turns, but the deepest kind, where a person elaborates, tests, revises or extends an idea rather than requesting or accepting an answer, appeared in only 4.9 per cent. It rose when the user arrived with a purpose to learn and when the assistant scaffolded rather than solved, and it was higher in coding than in writing. The study is observational and English-only, so it describes behaviour rather than proving what was learned, but it sets a sober baseline: most everyday AI use is shallow exchange, and the learning-bearing kind is rare and conditional.
​
Source: Chen et al. (2026), Informal learning emerges in everyday human-LLM interaction, arXiv:2607.17643. Preprint.
​
The students who use AI most, and understand it least, write the weakest
Jiang and Huang studied 120 Chinese undergraduates learning to write in English, all of whom had completed thirty hours of AI writing training, then set them four writing tasks with no AI allowed. Sorting the students by their AI-literacy profile revealed a clear pattern across measures of vocabulary, sentence structure and discourse. The group that applied AI most heavily but understood it least, and scored lowest on ethics and self-efficacy, produced the weakest independent writing. High use was not the problem; high use without understanding was. It is a direct argument that AI literacy has to mean comprehension of what the tool is doing, not fluent operation of it, if it is to leave any durable skill behind.
​
Source: Jiang, Z. and Huang, Z. (2026), From prompts to profiles, Humanities and Social Sciences Communications. In press.
Delegation can support learning when the partnership is framed well
In a study of 912 students across China, Europe and the United States, Wang and Zhang found that both critical evaluation, or vigilance, and task delegation, or offloading, contributed positively to transformative learning (R² = 0.475), contradicting the assumption that offloading always undermines deep learning. How the student framed the partnership mattered more than whether they delegated. It is the strongest recent counter-evidence to a pure cognitive-debt reading, and it moves the question from whether offloading harms learning to the conditions under which it does or does not.
​
Source: Wang, S. and Zhang, H. (2026), Pedagogical partnerships with generative AI in higher education, International Journal of Educational Technology in Higher Education, 23(1), art. 11. Peer-reviewed. Tracker ref B21.
​
The evidence linking heavier AI use to weaker critical thinking thickens
A run of 2026 studies points the same way. Work published in Frontiers in Psychology and related papers reports a strong negative correlation between AI tool use and critical thinking, in the region of r = -0.68, with cognitive offloading as the mechanism and cognitive fatigue as a partial mediator, alongside a recurring pattern of metacognitive laziness in which learners offload not only the task but the monitoring of it. The correlation figure rests on secondary citation and needs checking against the full texts before it is used on its own. Held against Wang and Zhang, the picture is not that offloading is uniformly harmful, but that unstructured offloading is.
​
Source: Cognitive offloading and critical thinking (2026), Frontiers in Psychology and related studies. Correlation figure needs full-text verification. Tracker ref C22.
​
A 30-month study: homework scores rise while exam scores fall
Strömberg at Stockholm University, with Lei and Wu at the University of Hong Kong, tracked 26,811 secondary students in a central-China county over 30 months. After AI tools arrived, homework scores rose about 18 per cent and homework time fell from 64 to 45 minutes, but monthly closed-book exam scores dropped about 20 per cent within six months, high-school entrance-exam scores fell about 24 per cent and national college-entrance scores fell about 18 per cent. The strongest students lost the most, about 24 per cent against 16 per cent, and the full exam effect took about two years to surface. It is the hardest longitudinal evidence yet that offloading can decouple visible performance from retained capability.
​
Source: Strömberg, Lei and Wu (2026), CEPR Discussion Paper 21577. Preprint.
​
The first rung of the career ladder keeps vanishing
The World Economic Forum's report on the future of entry-level work maps around 11 million US workers without four-year degrees in Gateway jobs, the roles that have historically enabled progression, and finds that nearly half of the pathways out of those jobs are highly exposed to AI. Read with the Stanford AI Index figure that employment for software developers aged 22 to 25 has fallen nearly 20 per cent, it turns the apprenticeship gap from argument into labour-market data: the entry tasks that let people build tacit expertise are being automated first, which raises the question of where the next generation of experts will come from.
​
Source: World Economic Forum (2026), Artificial Intelligence and the Future of Entry-Level Work. Report. Tracker ref B41.
England's disadvantage gap is widening again
The Education Policy Institute's Annual Report 2026 finds that 23 per cent of disadvantaged young people in England are not studying towards a substantial qualification or apprenticeship after GCSE, against 9.7 per cent of their better-off peers, a gap that has widened. It is not an AI study, but it is the equity backdrop against which AI in education lands. The WEF readiness framework's warning is that without the right foundations AI widens the gaps that education is meant to close, and this is the gap most likely to widen.
​
Source: Education Policy Institute (2026), Annual Report 2026. Report.
​
​
AI as subject: what people need to understand
Studies and reports about how these systems behave, and what the adoption picture actually is.
What AI does to learning depends on what you measure
Dumlao and colleagues examined 156,135 students across nearly 88,000 course offerings over a decade and found that the arrival of ChatGPT made no statistically significant difference to final grades, for students overall or for those who had been struggling. On its face that contradicts the 30-month China study, which found unaided exam scores falling by about a fifth. Both can be true at once. A final grade blends assisted work, which AI lifts, with unaided assessment, which it lowers, so the two can cancel out and leave the grade unchanged. The warning is that watching grades will show nothing while the capability that matters quietly slips, so institutions should measure what students can do without the tool.
​
Source: Dumlao et al. (2026), Generative AI availability, grades, and student satisfaction at a large university, arXiv:2607.21534. Preprint.
​
Adoption is high and reliability is jagged
The Stanford AI Index 2026, read first-hand from the full report, confirms organisational adoption of 88 per cent and a US-China capability gap down to 2.7 per cent as of March 2026, set against a reliability frontier that stays uneven: the top model reads an analogue clock correctly only 50.1 per cent of the time, and agents still fail roughly one attempt in three on structured benchmarks. The education chapter is blunt. Four in five US students use AI for schoolwork, only half of schools have any AI policy, and just 6 per cent of teachers say those policies are clear, while computer science enrolment fell 11 per cent in a year. The reliability figures are the counterweight to any claim that the technology can now do everything.
​
Source: Stanford HAI (2026), Artificial Intelligence Index Report 2026. Full report, first-hand. Tracker ref A57.
​
Near-universal use, and a wide gap in guidance and trust
The Digital Education Council's AI in Higher Education Survey 2026, the largest recent global read with 45,398 respondents across 35 countries, found that 88 per cent of students and 77 per cent of faculty now use AI, a 16-point jump on 2025. Yet 57 per cent of students say assessment guidance on AI is inadequate, only 29 per cent think their instructors can guide them, and 73 per cent of faculty worry their students are building skills at AI's expense. In the US and Canada, faculty intent to use AI fell year on year, from 76 to 67 per cent. The number to watch is not the 88 per cent usage but the 29 per cent confidence in guidance: near-universal adoption sitting on top of a governance vacuum.
​
Source: Digital Education Council (2026), AI in Higher Education Survey 2026. 45,398 respondents, 35 countries. Tracker ref B22.
A note on the evidence
Several of these studies are preprints and have not yet been peer-reviewed, including Chen, Contractor and Reyes, Hossain and Nawmi, Švábenský, Strömberg and Dumlao. This is noted for each because the findings are new and not yet reviewed. The cognitive-offloading correlation of r = -0.68 rests on secondary citation and should be checked against the full texts before it is quoted on its own. The two peer-reviewed studies are Ates and Wang and Zhang. Where a project already carries a Research Library reference, it is given so the entry can be filed against the tracker.
The Skinny AI News Items:
AI Development and Industry
AI labs begin to muscle in on the $6tn education market
24 July 2026 | Jamie John, Financial Times
The largest AI developers are now building education-specific offers. Anthropic has launched Claude for Teachers, free for US K-12 teachers, giving access to its Claude Cowork agent for lesson and assignment planning, alongside Claude for Education for universities, which includes a learning mode that asks students how they would approach a problem rather than supplying the answer. OpenAI runs a free ChatGPT for Teachers for verified US K-12 staff and a discounted ChatGPT Edu subscription for universities, and its Education for Countries programme, with partners in Estonia, Greece, Jordan, Kazakhstan, Slovakia, Trinidad and Tobago, the United Arab Emirates and Italy, has already reached more than 30,000 students, teachers and researchers in Estonia alone. Google offers a suite of Gemini-based tools on its education cloud. Morgan Stanley put the global education market at $6 trillion in 2022. Simon Buckingham Shum of University of Technology Sydney states plainly that education is now a marketplace for these firms. A Digital Education Council study found that more than 40 per cent of higher education students said AI had not been integrated into any of their courses, and among those whose courses did use it, 42 per cent found it only somewhat helpful.
What you need to know: The frontier labs are now direct suppliers to schools and ministries rather than infrastructure providers alone, which changes the procurement, data and governance questions institutions face. The 40 per cent not-integrated figure shows uptake at the vendor level running well ahead of integration at the course level, which is precisely where training and evaluation work is needed.
Original link: https://www.ft.com/content/e23e8b1c-693b-48b5-85b3-5852ebf70b02
​
Open-weight models close the gap as Kimi K3 weights are released
20 July 2026 | The Batch, Bloomberg
Moonshot released the downloadable weights of Kimi K3, a 2.8-trillion-parameter model that ranks third on the Artificial Analysis Intelligence Index behind only Claude Fable 5 and GPT-5.6 Sol, and is the cheapest of the leaders to run at about $0.95 a task against Fable 5's $2.75. Alibaba previewed Qwen3.8-Max at 2.4 trillion parameters. The direction is set by an earlier episode: when an OpenAI test agent hacked Hugging Face, Hugging Face analysed the attack using an open Chinese model, Z.ai's GLM 5.2, because a US frontier model refused the task. Against this, flagship per-token prices have been rising across vendors, so the self-hostable open-weight option is becoming the affordable route rather than the compromise.
What you need to know: Capable models that an institution can host and adapt itself are becoming the affordable option, with fewer refusals and the ability to fine-tune. That widens what AI literacy has to cover. Understanding weights, hosting cost, distillation and provenance now matters as much as knowing how to prompt a single closed chatbot.
Source: The Batch and Bloomberg, 20 July 2026
​
The Magnificent Seven shed about $800bn as AI spending unsettles markets
23 July 2026 | Bloomberg, Financial Times
Spending plans from Alphabet and Tesla triggered an almost $800 billion one-day fall in the Magnificent Seven, the largest since April 2025, after Alphabet raised its 2026 capital expenditure guidance to as much as $205 billion and posted its first negative free cash flow as a public company. Bank of America's Sebastian Raedler expects equities to fall 7 to 8 per cent, citing an AI capex boom with, in his words, a lot of cracks. The one-day move sits alongside a wider question set out by analysts: the four largest hyperscalers plan a combined $725 billion of capital spending this year, and it is not yet clear that the revenue will follow.
What you need to know: Doubts about the sustainability of AI capital spending mean the price institutions pay for AI is likely to stay high and volatile. The assumption that access will only get cheaper is contested at the infrastructure layer, even as headline token costs fall, so budget planning should not depend on it.
Source: Bloomberg and Financial Times, 23 July 2026
​
Intel posts its fastest growth in 15 years on inference demand
23 July 2026 | Financial Times
Intel reported revenue up 25 per cent to $16.1 billion, with data-centre and AI revenue up 59 per cent to $6.3 billion and foundry revenue up 31 per cent. Chief executive Lip-Bu Tan framed the results as a shift from graphics processors used for training towards central processors used for inference, that is, from the one-off cost of building a model to the repeated cost of running it billions of times.
What you need to know: The training-to-inference distinction is a clean way to teach AI's two-phase cost structure: developing the recipe once, then following it at scale. It also explains why the per-use price of AI, which is what institutions actually pay, behaves differently from the headline cost of building frontier models.
Source: Financial Times, 23 July 2026
​
The productivity payoff from AI remains hard to find
22 July 2026 | Financial Times (Alphaville), ONS, St Louis Federal Reserve
FT Alphaville set out the case that the measured productivity gains from AI are still thin. Barclays calls the AI productivity pickup surprisingly fragile and finds the correlation between adoption and productivity not statistically distinguishable from zero. A St Louis Federal Reserve survey shows AI in the work routines of 45 per cent of people but only about one in ten using it daily. Office for National Statistics data show UK firms using an average of 1.6 AI tools, barely up from 1.4 in 2023, set against roughly $725 billion of sunk AI infrastructure cost. Adoption is wide but shallow.
What you need to know: The gap between broad access and daily, skilled use is the practical problem, and it is a training problem rather than a technology one. The case that AI pays for itself is unproven at the level of the whole economy, which strengthens the argument for spending on capability and process rather than on licences alone.
Source: Financial Times (Alphaville), ONS and St Louis Federal Reserve, 22 July 2026
Anthropic files to go public as frontier-lab economics come into view
1 June 2026 | Fortune, SEC filing
Anthropic confidentially filed an S-1 with the US Securities and Exchange Commission, having raised $65 billion at a $965 billion valuation, with an October Nasdaq listing targeted. The first frontier-lab public offering will set how the sector reports its financials, and its consumption-based pricing is the same mechanism that drove the enterprise cost stories of the past year. The listing sits within a wider round of consolidation and heavy capital spending: SpaceX agreed an all-stock $60 billion acquisition of the coding tool Cursor, and Oracle said it would spend $70 billion on data centres in the coming year.
What you need to know: How the first listed AI lab prices its product will feed directly into what schools and universities pay, because consumption-based pricing passes usage cost to the customer. Institutions building on a single provider are also exposed to that provider's commercial and ownership changes, which is a dependency worth naming in procurement.
Source: Fortune and SEC filing, 1 June 2026
Cloudflare moves to charge AI crawlers as bots overtake human web traffic
15 July 2026 | The Batch
Cloudflare, which sits in front of nearly 20 per cent of the web, will let publishers separately allow or block search, AI-training and AI-agent crawlers from 15 September, and will charge AI agents per request through a new monetisation system. The company reports that 57.5 per cent of all HTTP requests now come from automated systems rather than people.
What you need to know: The terms on which text can be collected and used to train models are hardening, which bears directly on scholarly and educational content and on who can build with it. Smaller developers and open projects face higher barriers, so the practical effect may be to concentrate capability among those who can pay for data access.
Source: The Batch, 15 July 2026
AI Regulation, Geopolitics, and Legal Issues
UK Jurisdiction Taskforce publishes legal statement on liability for AI harms
7 July 2026 | UK Jurisdiction Taskforce
The UK Jurisdiction Taskforce has published its final Legal Statement on Liability for AI Harms under the private law of England and Wales, following a public consultation held in January and February 2026. Prepared by a drafting team chaired by Matthew Lavy KC, it addresses non-deliberate harm: when parties who did not set out to cause harm may nonetheless be liable for harms arising from the use of AI. It covers negligence, vicarious liability, professional liability, product liability and the law of false statements. The overall conclusion is that English common law is flexible enough to accommodate most of the questions AI presents, although the Law Society notes that the statement also identifies gaps requiring government action, particularly whether product liability law extends to standalone AI software. Because AI has no legal personality, no party can be vicariously liable for the AI itself.
What you need to know: The sharpest implication for professionals is that liability may attach not only to careless use of AI, but to a failure to use it where a reasonably competent peer would have done so. For any regulated profession, and for the training that supports it, that reframes adoption as a potential professional standard rather than an optional extra.
​
Kanishka Narayan becomes the first UK AI minister to attend Cabinet as DSIT is abolished
20 July 2026 | Bloomberg, IT Pro, Technology Magazine
On becoming Prime Minister on 20 July 2026, Andy Burnham appointed Kanishka Narayan as Minister for Artificial Intelligence with attendance at Cabinet, the first time the AI brief has carried that status in the UK. The same reshuffle abolished the Department for Science, Innovation and Technology, created in 2023, and dispersed its functions. The bulk moved into a new Department for Business, Innovation, Science and Trade under Jonathan Reynolds, with other responsibilities passing to a renamed Department for Digital, Culture, Media and Sport under Lisa Nandy. DSIT had held the AI Opportunities Action Plan, the AI Security Institute, the £500 million Sovereign AI Fund and UK Research and Innovation. Former government AI adviser Matt Clifford called the restructuring a mistake, arguing that a reorganisation consumes official time needed for substance.
What you need to know: AI governance in England now sits with a Cabinet-attending minister but without a single departmental home. How the AI brief, online safety, the Department for Education and the renamed DCMS interact is unresolved, and that interface will determine where education-facing AI policy is actually made.
Original link: https://www.itpro.com/business/policy-and-legislation/ai-minister-secures-cabinet-seat-as-dsit-merged-with-new-business-department
​
White House launches Gold Eagle as Claude Mythos returns after an export block
14 July 2026 | White House, CNN, The Information
The White House launched Gold Eagle, an AI cyber-vulnerability clearinghouse run with the Treasury, CISA and the Department of War, the first major implementation of the 2 June executive order on advanced AI and national security. Days earlier the government revised its licence terms to allow Anthropic a limited release of Claude Mythos Preview, after an initial export block in June had led Anthropic to disable its leading Fable 5 and Mythos models for all customers. Mythos had built working exploits for 181 of 271 vulnerabilities it found in Firefox.
What you need to know: Governance is forming around a concrete cybersecurity mechanism rather than around abstract existential risk, which makes this a live teaching case for how AI policy is actually made. The June episode also showed that access to a frontier model can be withdrawn at short notice, so any institution building on a single provider carries a real continuity risk.
Source: White House, CNN and The Information, 14 July 2026
​
Big Tech urges the US not to restrict open-weight models as Chinese models gain share
24 July 2026 | Financial Times, The Information
Nvidia, Palantir, Microsoft, Meta, Andreessen Horowitz and IBM signed an open letter urging the US not to restrict open-weight AI models. OpenAI, Anthropic and Google did not sign. The push followed the release of Kimi K3. Chinese models have accounted for about 30 per cent of the tokens used by US firms since February, according to William Blair citing OpenRouter, and GIC's Bryan Yeo said cheaper Chinese open-weight models would cut global adoption costs.
What you need to know: This is an open-versus-closed policy fight with direct affordability and access implications for institutions. Whose models reach which education systems, and on what terms, is becoming a matter of industrial policy rather than a purely technical choice.
Source: Financial Times and The Information, 24 July 2026
​
The EU frames AI as a geopolitical weapon and presses for sovereignty
21 July 2026 | Financial Times
The EU's technology chief Henna Virkkunen called AI a geopolitical weapon and pressed for European technological sovereignty, citing June's brief US export controls on Anthropic's leading models, later lifted, as a warning that access could be cut off. Brussels has presented a technology-sovereignty package backing European alternatives such as Mistral, Scaleway and OVHcloud.
What you need to know: Reliance on a single country's models is now framed as a strategic vulnerability. For education systems the same logic applies to how they build home-grown AI capability rather than depending wholly on foreign tools, which is a question of curriculum and skills as much as of procurement.
Source: Financial Times, 21 July 2026
​
Xi launches a World AI Cooperation Organisation as the US counters with Pax Silica
23 July 2026 | Financial Times
Xi Jinping used the Shanghai World AI Conference to launch a 29-member World AI Cooperation Organisation and to pledge open-source models plus international application co-operation centres to train thousands of people from developing countries. The US has a rival coalition, Pax Silica, focused on supply chains and critical minerals.
What you need to know: China is using AI talent training and open-source models as capacity-building and soft power across the global south. Whose tools, standards and pedagogy reach developing education systems will be shaped by this competition, which matters for anyone working on AI in education beyond the wealthiest countries.
Source: Financial Times, 23 July 2026
​
EU AI Act high-risk rules for education take effect on 2 August
22 June 2026 | EU AI Act, sector coverage
High-risk obligations under the EU AI Act take effect on 2 August 2026. AI used for student assessment, admissions screening and progress monitoring is classified as high-risk, which triggers transparency and human-oversight duties that consumer chatbots do not currently meet. The reach is extraterritorial, so it applies to providers and deployers serving EU learners regardless of where they are based.
What you need to know: This is a hard regulatory threshold that directly affects how institutions may use AI for assessment, admissions and monitoring. Any tool used for those purposes needs a documented human-oversight arrangement and transparency to learners, and the deadline is now, not on the horizon.
Source: EU AI Act and sector coverage, June 2026
​
Florida becomes the first US state to sue OpenAI over child safety
2 June 2026 | Joe Miller, Financial Times
Florida filed an 83-page complaint against OpenAI and Sam Altman, alleging addictive and unsafe products and harm to children. The state attorney-general invited other states to follow and said the office was examining other AI models.
What you need to know: Child-safety litigation will shape what AI products are permitted in classrooms faster than curriculum policy will, and it is advancing even in a deregulatory federal climate. For schools it is a signal to check the age-appropriateness and safety claims of any tool put in front of under-18s.
Source: Financial Times, 2 June 2026
AI Research and Evaluation
Common Sense Media establishes the Youth AI Safety Institute
May 2026 | Common Sense Media
The Youth AI Safety Institute is an independent effort by Common Sense Media to research the effects of AI on children, set safety standards, evaluate products and publish what it finds. Launched in early May 2026, it stress-tests widely used AI products against real-world, adversarial, multi-turn and youth-specific scenarios using human expert reviewers, with the stated intention of developing these into repeatable automated tests over time. Assessments published so far cover the AI features in Google Search, the market for AI mental health applications, which the institute describes as unregulated and in some cases actively harmful to teenagers, and Anthropic's Claude, which it credits with meaningful safety strengths alongside risks that arise where teenagers use a product designed for adults.
What you need to know: An independent body now issues product-level risk ratings for the AI tools children actually use, which gives schools, trusts and regulators an external reference they can cite in procurement and policy. Expect questions about whether a specific tool carries a rating, and use the ratings to inform, not to replace, local judgement.
Original link: https://institute.commonsensemedia.org/
​
Google's AI search features rated an unacceptable risk to children
15 July 2026 | Youth AI Safety Institute, Common Sense Media
The Youth AI Safety Institute's risk assessment of AI Overview and AI Mode in Google Search concludes that both pose unacceptable risks to children. Researchers ran more than 2,600 searches using accounts configured with SafeSearch for child and teenage users aged 11 and 15. The features scored unacceptable or high risk on seven of the institute's eight AI principles, with an unacceptable rating on keeping children and teenagers safe, where testing found failures across all five severe-harm categories. The features also completed homework on demand and returned inaccurate answers with the same confidence as accurate ones. Robbie Torney, who leads the institute's AI and digital assessments, noted that unlike the standalone Gemini chatbot, schools and parents cannot switch off the AI answers in Google Search. An earlier survey found AI-generated search summaries are the most widely used form of AI among 9 to 17 year olds, with 75 per cent having used them.
What you need to know: The finding that the feature cannot be disabled is the operational problem for schools, particularly those running Google Workspace and Chromebooks, because it removes restriction of access as a mitigation. The practical response points to information literacy: pupils need to be taught to tell a synthesised answer from a sourced one before critical thinking can even begin.
Original link: https://institute.commonsensemedia.org/risk-assessments/google-search
​
EDSAFE AI Alliance and Arizona State University propose a learning sciences benchmark
13 July 2026 | EDSAFE AI Alliance, ASU Mary Lou Fulton College
The EDSAFE AI Alliance and Arizona State University's Mary Lou Fulton College for Teaching and Learning Innovation have released a framework, developed with more than 100 edtech developers, learning scientists and civil rights advocates, proposing an independent, open-access benchmark that assesses generative AI tools across 15 domains organised into three pillars: cognitive foundations, covering human agency, executive function and relational bonds; pedagogical design, prioritising productive struggle and Socratic inquiry over task automation; and measurement accountability. The paper treats AI as permanent infrastructure rather than a passing tool, and argues that classroom deployment has outpaced evaluation. Erin Mote, chief executive of InnovateEDU, framed the work as giving researchers, developers and district leaders baseline metrics rooted in how people actually learn, rather than metrics of speed or accuracy.
What you need to know: This is the closest thing yet to a shared yardstick for judging whether an AI tool supports learning rather than merely completing tasks. The three-pillar structure is a usable reference point for evaluating tools, and it puts the question of productive struggle, not just output quality, at the centre of assessment.
Original link: https://www.edsafeai.org/future-proofing-human-flourishing
​
Concerns rise as the volume of research soars and quality drops
24 July 2026 | Andrew Jack, Financial Times
Claudine Gartenberg, associate professor at the Wharton School and a senior editor at Organization Science, noticed rising submission volumes of declining quality and, with colleagues, published one of the first rigorous studies of AI's effect on academic publishing. They attribute a 42 per cent rise in submissions to the journal since 2022 to AI, and link its use to a measurable fall in quality. Peer review quality declined in parallel, with reviewers focusing more on the theories papers relied on than on the new findings and data they contained. A related paper showed AI generating nearly 400 publication-ready finance papers in 12 hours. The pattern has a corporate echo: in June, KPMG withdrew a report on agentic AI after GPTZero found only 5 of 45 citations correct and 40 of 45 titles fabricated, and named organisations disputed the claims made about them.
What you need to know: Peer review is the quality control layer for the evidence base that education and training rely on, and this is the first quantified account of it degrading. The finding that reviewers shifted attention from data to theory is the detail worth carrying into literature review methods and source appraisal, in professional training as much as in universities.
Original link: https://www.ft.com/content/52e688a0-c6c1-4161-9878-1fab12c5e806
​
Universities should teach students to build AI evaluations
24 July 2026 | Andrew Hall, Financial Times
Andrew Hall, professor of political economy at Stanford Graduate School of Business, argues that universities restricting or banning AI are walling the technology off from the people who most need to understand it. He has taught the alternative by having undergraduates with no coding background build their own model evaluations. Within three hours every student had a working evaluation with a webpage reporting model performance against their own criteria, and the class produced rankings across 24 student-defined criteria. Hall is candid about the limits: evaluations are only as good as the expertise encoded in them, and even the benchmarks the labs rely on are shaky, since when OpenAI reviewed SWE-bench, nearly 60 per cent of audited problems contained flawed test cases. He concludes that courses prohibiting AI remain necessary alongside courses teaching evaluation.
What you need to know: Evaluation literacy is proposed here as a teachable skill rather than a specialist function, and the three-hour result is concrete evidence that it can be done at scale. It offers a third option to institutions currently choosing between banning AI and permitting it uncritically: teach people to measure a model against their own criteria.
Original link: https://www.ft.com/content/ac15aa4a-960f-4957-90c6-ff1b4a390fa4
​
Stanford's AI Index 2026 shows high adoption and jagged reliability
15 July 2026 | Stanford HAI
Stanford's Institute for Human-Centered AI published the AI Index 2026. Four in five US students use AI for schoolwork, only half of schools have AI policies, and 6 per cent of teachers call those policies clear. Computer science enrolment fell 11 per cent, employment for developers aged 22 to 25 fell nearly 20 per cent, and organisational adoption reached 88 per cent. Reliability stayed uneven: the report records leading models reading analogue clocks correctly only 50.1 per cent of the time.
What you need to know: The index quantifies the gap between student adoption and institutional readiness on policy, training and assessment. The reliability figures are the counterweight to any claim that the technology can now do everything, and they are a reminder to teach where these systems fail as well as where they succeed.
Source: Stanford HAI AI Index 2026, 15 July 2026
​
A 30-month study finds homework scores rise while exam scores fall
24 July 2026 | Strömberg, Lei and Wu, CEPR Discussion Paper 21577
Strömberg at Stockholm University, with Lei and Wu at the University of Hong Kong, tracked 26,811 secondary students in a central-China county over 30 months. After AI tools arrived, homework scores rose about 18 per cent and homework time fell from 64 to 45 minutes, but monthly closed-book exam scores dropped about 20 per cent within six months, high-school entrance-exam scores fell about 24 per cent and national college-entrance scores fell about 18 per cent. The strongest students lost the most, about 24 per cent against 16 per cent, and the full exam effect took about two years to surface.
What you need to know: This is the hardest longitudinal evidence yet that offloading can decouple visible performance from real learning. The metric AI inflates, homework, is the one schools watch; the metric it lowers, unaided exams, is the one that measures retained capability. The two-year lag means short-run dashboards will mislead, so measure what students can do without the tool.
Source: CEPR Discussion Paper 21577 (Strömberg, Lei and Wu), 24 July 2026
​
Capability is not reliability, and the difference matters
2 July 2026 | John Burn-Murdoch and Sarah O'Connor, Financial Times
The FT set out that capability benchmarks such as METR and reliability rubrics such as Princeton's measure different questions. METR asks whether AI can occasionally succeed, which is enough to pose a danger. Reliability rubrics ask whether it can be depended upon, which is what replacing a worker actually requires. The gap between the two is where much of the confusion about AI's effect on work sits, and it is visible in practice: Ford has rehired experienced engineers because automated quality control was not good enough.
What you need to know: The distinction between occasional success and dependable performance is a useful teaching frame for AI literacy, and a corrective to demonstrations that show a model succeeding once. For workforce planning it explains why a tool that can do a task in a demo may still not be able to do the job.
Source: Financial Times, 2 July 2026
AI Ethics and Societal Impact
Meta opts Instagram accounts into AI image generation, then withdraws the feature
7 July 2026 | Reece Rogers, Wired
Meta launched Muse Image, its first AI image model from Meta Superintelligence Labs, with deep integration into Instagram. Public Instagram profiles were automatically opted into use for generative image remixes: anyone could tag a public account in a prompt and generate an image using that person's likeness. Users were not notified when their content was used, and switching an account to private or changing the setting prevented new generations without deleting images already made. Opting out required finding a sharing and reuse section in the app settings and disabling separate toggles. SAG-AFTRA and CAA both objected publicly, with CAA arguing that no one's name, image, likeness, voice or creative work should be used by a third party without documented consent. On 10 July Meta withdrew the feature, saying it had missed the mark.
What you need to know: The three-day arc from launch to withdrawal shows consent-by-default on likeness attracting fast and organised resistance. The residual question for schools and universities is what happens to images generated before the withdrawal, since Meta made no commitment to remove them, and that is a live safeguarding and data point for anyone whose students use the platform.
Original link: https://www.wired.com/story/meta-now-lets-anyone-use-your-instagram-photos-in-ai-images-unless-you-opt-out/
Instant AI answers may be eroding curiosity
8 July 2026 | Anne-Laure Le Cunff, The New York Times
Anne-Laure Le Cunff, a researcher at the Institute of Psychiatry, Psychology and Neuroscience at King's College London and a former Google executive, notes that more than 60 per cent of Google searches in the United States now end without a click, and that Claude, ChatGPT and Perplexity produce the same behaviour. Her argument is that the interval between forming a question and receiving an answer is where much of the learning happens, rather than dead space to be removed. She cites evidence that people waiting for the answer to an intriguing question remember unrelated information better, so curiosity temporarily raises learning across the board. Separate data make the point about search itself: on Google's AI Mode, users write queries about three times as long as before, and in about 75 per cent of sessions never click through to the open web.
What you need to know: This supplies a mechanism rather than a complaint, which makes the speed of an answer a variable in learning design rather than a neutral convenience. It is a defensible basis for building deliberate delay and exploration into AI-mediated tasks, and for teaching learners to trace and verify a synthesised answer.
Source: Anne-Laure Le Cunff, The New York Times, 8 July 2026
​
Universities drop AI detection tools over fears about accuracy
24 July 2026 | Ima Jackson-Obot, Financial Times
Vanderbilt, Yale, Johns Hopkins, Northwestern, the University of Waterloo, the University of Cape Town and Curtin University have restricted or disabled AI detection tools over reliability concerns. The framing case is that of Orion Newby, an Adelphi University student wrongly accused after Turnitin returned a score indicating the submission was highly likely to be AI-generated while other tools returned zero; in January a New York court ruled in Newby's favour. The underlying pressure is real: an Edinburgh Napier University survey of more than 6,600 students across seven UK universities found that 32 per cent admitted some level of unpermitted AI use. Sam Illingworth's audit of 163 UK universities found that more than 40 per cent had no publicly accessible AI policy, and a 2023 Stanford study found detection tools disproportionately produce false positives for speakers of English as an additional language. Judy Williams of Queen's University Belfast holds that the answer is good assessment design.
What you need to know: The centre of gravity has moved from detection to assessment design, and the Newby ruling gives that shift legal weight. The most actionable finding is that more than 40 per cent of UK universities still have no published AI policy, which is a governance gap rather than a technology one, and one that can be closed quickly.
Original link: https://www.ft.com/content/49304b1e-8a9d-4fb6-bc4d-37dd3430bb98
Elite AI-powered schools insist they provide a route to excellence
24 July 2026 | Georgina Quach, Financial Times
Founders School in Manhattan opens in September with 20 places, charging $150,000 a year and offering to refund fees if a pupil has not made $1 million in profit by the end of the four-year course. Modelled on its parent Alpha School, it compresses academic study into three hours taught by an adaptive AI tutor, leaving the rest of the day for mentored business building. Alpha uses mastery rather than grades, requiring 90 per cent understanding before progression. In the UK, David Game College pioneered a teacherless classroom on the same three-hour model, charging £35,000 a year for non-UK pupils, and Inspired Education Group is launching seven AI-enabled primary schools. Against these claims, Stanford researchers note there has been scant evaluation of AI's long-term outcomes in education, and a large-scale Chinese study found homework scores rising while exam scores fell by around a fifth. Teach First's James Toop and Accenture's Matt Prebble warn that poor implementation could reinforce existing inequalities.
What you need to know: The most enthusiastic adopters are among the most expensive schools in the world, and the supplier side concedes there is no long-term outcome evidence. The Chinese longitudinal data, reported above, is the sharpest counterpoint available, and it belongs in any conversation about AI tutoring claims.
Original link: https://www.ft.com/content/0070d650-4b17-4979-916e-23f8421ddd93
Chatbot use for mental health rises as trust falls
30 June 2026 | Lucy Warwick-Ching, Financial Times; Pew Research Center
A Mental Health UK and Censuswide survey found that 37 per cent of UK adults, rising to 64 per cent of 25 to 34 year olds, have used chatbots for mental health or wellbeing conversations, and clinicians are divided on whether these tools reflect distress back or reinforce it. Anna Maratos notes that the change therapy produces comes through the human relationship that a validating system removes. Public sentiment more broadly is moving the other way from deployment: a Pew Research Center survey of 5,119 US adults found nearly half now use chatbots, yet 60 per cent read AI-generated summaries without trusting them, and a majority expect AI to harm them personally rather than help.
What you need to know: Usage is rising while trust falls, which complicates any assumption that learners want more AI in teaching. For pastoral and wellbeing provision, the finding that a validating system removes the relational element of support is a reason to be cautious about pointing students towards chatbots for help.
Source: Financial Times, 30 June 2026, and Pew Research Center, 23 June 2026
Heavier AI use is being linked to weaker critical thinking
15 July 2026 | Frontiers in Psychology and related studies (2026)
A run of 2026 studies has linked heavier AI use to lower critical thinking, with cognitive offloading proposed as the mechanism and cognitive fatigue as a partial mediator. One synthesis reports a strong negative association, though the precise correlation is still to be confirmed against the full texts. A separate study of 912 students across China, Europe and the United States moves the question on, asking whether critical evaluation and strategic delegation to AI can coexist rather than trade off against each other.
What you need to know: The useful shift here is from asking whether AI use lowers thinking to asking when structured delegation preserves it. That is a design question for teaching: the task is to build in the points where a learner must evaluate, decide and justify rather than accept the machine's output.
Source: Frontiers in Psychology and related studies, 2026
AI in Work and Education
Employers reverse AI-driven redundancies
1 July 2026 | Justina Lee, CNBC
Companies that cut roles citing AI are rehiring. Ford has rehired, newly hired or promoted more than 350 experienced engineers after automated quality control systems failed to match the judgement of veteran staff, and then topped J.D. Power's 2026 Initial Quality Study for the first time since 2010. Commonwealth Bank of Australia reversed the loss of more than 40 customer service roles after an AI voice bot could not cope with call volumes. IBM's AI resolved roughly 94 per cent of HR requests but struggled with the remainder requiring ethical judgement, and the company now plans to triple entry-level hiring in the United States. Research by Orgvue found that 39 per cent of business leaders had made staff redundant because of AI, and 55 per cent of that group later concluded the decision was wrong.
What you need to know: The 55 per cent regret figure is the most useful single number for testing automation-savings arguments in workforce planning. The underlying error was treating AI as a headcount reduction rather than as a layer that still needs quality gates, escalation paths and human oversight, all of which are training questions.
Original link: https://www.cnbc.com/2026/07/01/employers-who-laid-off-workers-for-ai-are-reversing-their-decisions.html
New York school district pauses humanoid robot deployment after scrutiny
14 July 2026 | New York Focus
Salamanca City Central School District in western New York, a rural district on the Seneca Nation reservation, approved a $57,590 purchase from Toronto-based Realbotix: a stationary humanoid robot named Sally, together with the Optio AI teaching assistant, which students could interact with as an avatar on their laptops. The robot was to support students in AI and robotics courses rather than replace teachers. Realbotix described Salamanca as its flagship deployment and acknowledged that it holds no empirical evidence the robot improves learning. Following the New York Focus report, local parents objected and State Education Commissioner Betty Rosa wrote to the superintendent expressing concern. On 24 July the district announced the programme was on hold while it works through enhanced student data privacy agreements.
What you need to know: A procurement of under $60,000 was halted within ten days by press attention, parental objection and a letter from the state regulator, which makes this a useful reference for how quickly AI deployments in schools can be reversed on governance grounds rather than technical ones. The absence of any efficacy evidence, conceded by the supplier, is the detail most likely to recur in due diligence.
Original link: https://nysfocus.com/2026/07/14/new-york-humanoid-robot-teacher-salamanca-school-district
Is AI killing critical thinking in the classroom?
24 July 2026 | Andrew Jack, Financial Times
Madeleine Champagnie, innovation lead and head of English at Thames Christian School in London, reports the opposite of the expected problem: many of her students are reluctant to use AI at all, being as wary of cognitive offloading as their teachers. Set against that, a number of studies find students substituting AI for their own judgement, and some show that while AI use raises short-term speed and results, its removal leaves students performing less well than those who never had access. A March report from the Australian Network for Quality Digital Education distinguishes detrimental offloading, which produces false mastery, from beneficial offloading of lower-order tasks that frees learners for the work that matters. Rose Luckin, professor emerita at UCL and chief executive of Educate Ventures Research, argues that the process of education itself, including assessment, may need rethinking: we need to unpick the old processes and identify which bits matter.
What you need to know: The distinction between beneficial and detrimental offloading is the most usable analytical tool in the debate: use AI to clear lower-order tasks so learners can concentrate on evaluating evidence and structuring arguments, not to replace that work. The finding that performance drops once AI is withdrawn is the strongest counter to short-term efficacy claims.
Original link: https://www.ft.com/content/decf11c3-d217-4c91-a238-0c67f419336c
First V-level subjects announced, including digital
27 July 2026 | Hayley Clarke, BBC News
The Department for Education has confirmed the first three V-level subjects for rollout from September 2027: education and early years, finance and accounting, and digital. V-levels are the new Level 3 vocational qualifications, equivalent to one A-level and combinable with A-levels, and they will replace the roughly 900 vocational qualifications currently available to 16 to 19 year olds. Current Year 9 pupils form the first eligible cohort. The digital V-level is intended to give students a route into practical digital and AI-related skills without committing to a specialised technical pathway at 16. Further subjects follow from 2028. Education Secretary Bridget Phillipson framed the reforms as ending snobbery in post-16 education, and the Sixth Form Colleges Association welcomed the decision to retain BTECs while V-levels are phased in.
What you need to know: AI-related content is entering the post-16 vocational route inside a broader digital qualification rather than as a standalone subject, which is the practical answer to how AI gets taught at Level 3. Detail remains thin, and the timing of the announcement during the summer holidays drew comment from school and college leaders.
Original link: https://www.bbc.co.uk/news/articles/cj4k2djd5qpo
AI becomes the most-cited reason for layoffs, and the entry-level rungs go first
15 July 2026 | Stanford HAI AI Index, WEF, Challenger, Gray and Christmas
AI became the single most-cited reason US employers gave for layoffs, with AI-attributed cuts of 87,714 in the first five months of 2026, past the 54,836 recorded across all of 2025. The World Economic Forum's future-of-entry-level-work report maps around 11 million non-graduate workers in Gateway jobs, with nearly half of the onward pathways highly exposed to AI. The wider picture is mixed: US tech companies have cut nearly 140,000 jobs since the start of 2026, Uber cut 10 per cent of its customer-service roles explicitly to adopt AI, and UC Berkeley economist Enrico Moretti argues AI is partly a pretext for correcting over-hiring. FT analysis found firms citing AI underperformed the Nasdaq by almost 10 per cent over the following 30 trading days.
What you need to know: The entry-level tasks that build tacit expertise are being automated first, which raises the question of where the next generation of experts will come from. That makes reskilling a supply question for whole sectors, and it puts a premium on training routes that let people build judgement before the junior rungs disappear.
Source: Stanford HAI AI Index, WEF and Challenger, Gray and Christmas, 15 July 2026
AI is helping workers sue their employers, and the tribunal system is straining
2 July 2026 | Delphine Strauss, Financial Times
AI-assisted employment claims have moved from novelty to standard practice, contributing to a 39 per cent rise in single tribunal claims and a 55 per cent rise in the backlog to 64,000 open cases. Some final hearings are now listed for 2028. Polished, validating AI output raises claimants' expectations of their own case and adds to the volume the system must process. The education parallel is visible in a separate trend the piece notes: Edapt reports rising complaints against teachers.
What you need to know: This is a concrete example of AI raising the volume and confidence of submissions faster than the institutions receiving them can respond, a pattern that also affects admissions, appeals and complaints in education. Processes designed for human-paced volumes may need redesigning for AI-assisted ones.
Source: Financial Times, 2 July 2026
The tacit-knowledge advantage: why experts gain leverage
20 July 2026 | Sarah O'Connor, Financial Times
Sarah O'Connor reframed corporate AI adoption through Frederick Winslow Taylor and Michael Polanyi's observation that we can know more than we can tell. Off-the-shelf Gemini, Claude and GPT variants parsed Bridgewater's documents correctly only about 50 per cent of the time, whereas a model fine-tuned with the firm's own experts and data reached about 85 per cent, and Ford has rehired veteran engineers to train its machines rather than replace them. Her earlier reporting drew the contrast the other way: machine translation de-skilled translators through post-editing, with one subtitle rate cut from $5.00 to $1.50 per minute. AI amplifies work when it removes the routine part of a job and de-skills it when it removes the part that carries meaning.
What you need to know: The value came from experts who could articulate their judgement, not from the model alone, which is exactly what good teaching develops. For any training programme the task to protect is the one that develops the person, and the task to automate is the routine one that does not.
Source: Financial Times, 20 July 2026
China makes reskilling a matter of state policy
30 June 2026 | Katrin Bennhold, The New York Times
China's five-year plan commits to address AI's effect on employment, backed by court rulings against AI-driven layoffs, a proposal for AI-unemployment insurance and a push on vocational retraining. Kyle Chan of Brookings frames the aim as augmenting rather than replacing workers. The corporate side shows the scale: Richard Liu of JD.com said its 700,000 couriers would be replaced by robots sooner or later, and the firm has signed contracts with about 120 schools to retrain them for work such as robot repair and maintenance.
What you need to know: Reskilling is being treated here as deliberate state policy, in contrast to the company-led model elsewhere, which is a reminder that policymakers have agency over the direction of the technology. The JD.com commitment is a concrete benchmark for what credible, large-scale training provision looks like.
Source: The New York Times, 30 June 2026
US universities show both a mandate and a backlash
June 2026 | SUNY Board of Trustees; NYT Magazine (Linda Kinstler)
The State University of New York adopted a systemwide AI policy across its 64 campuses, embedding AI literacy into general education for all incoming undergraduates from autumn 2026, with bias evaluation and data-privacy requirements, a significant precedent for mandatory AI literacy at scale in public higher education. The other side of adoption showed at California State University, which renewed its OpenAI deal on a smaller $13 million, three-year agreement after a faculty petition of nearly 4,000 signatures against the original $16.9 million, 500,000-licence contract, struck during a $2.3 billion deficit. Roughly half the licences were ever activated. The professors who coped best required two versions of each assignment, one without AI and one with, plus a written account of use.
What you need to know: Together these show institutional AI adoption running ahead of any settled view of its purpose. The transferable detail is pedagogical: making AI use visible through paired assignments and a written account of use is a practical alternative to either banning or ignoring it.
Source: SUNY Board of Trustees and NYT Magazine, June 2026
Further Reading: Find out more from these resources
Resources:
-
Watch videos from other talks about AI and Education in our webinar library here
-
Watch the AI Readiness webinar series for educators and educational businesses
-
Study our AI readiness Online Course and Primer on Generative AI here
-
Read our byte-sized summary, listen to audiobook chapters, and buy the AI for School Teachers book here
-
Read research about AI in education here
-
Watch Rose Luckin demystify AI using baking on Rose's AI here
About The Skinny
Welcome to "The Skinny on AI for Education" newsletter, your go-to source for the latest insights, trends, and developments at the intersection of artificial intelligence (AI) and education. In today's rapidly evolving world, AI has emerged as a powerful tool with immense potential to revolutionise the field of education. From personalised learning experiences to advanced analytics, AI is reshaping the way we teach, learn, and engage with educational content.
In this newsletter, we aim to bring you a concise and informative overview of the applications, benefits, and challenges of AI in education. Whether you're an educator, administrator, student, or simply curious about the future of education, this newsletter will serve as your trusted companion, decoding the complexities of AI and its impact on learning environments.
Our team of experts will delve into a wide range of topics, including adaptive learning algorithms, virtual tutors, smart classrooms, AI-driven assessment tools, and more. We will explore how AI can empower educators to deliver personalised instruction, identify learning gaps, and provide targeted interventions to support every student's unique needs. Furthermore, we'll discuss the ethical considerations and potential pitfalls associated with integrating AI into educational systems, ensuring that we approach this transformative technology responsibly. We will strive to provide you with actionable insights that can be applied in real-world scenarios, empowering you to navigate the AI landscape with confidence and make informed decisions for the betterment of education.
As AI continues to evolve and reshape our world, it is crucial to stay informed and engaged. By subscribing to "The Skinny on AI for Education," you will become part of a vibrant community of educators, researchers, and enthusiasts dedicated to exploring the potential of AI and driving positive change in education.
