In the autumn of 2023, about a thousand students at a high school in Turkey took four 90-minute maths sessions. One group practised with GPT-4 as it comes. A second group practised with a version that gave hints and refused to hand over full answers. A third group had no AI at all. Then everyone sat an exam without any help.
With plain GPT-4, practice scores rose 48% against the control group, and exam scores fell 17%. With the guardrails, practice scores rose 127% and the exam came out the same as the control. The same model improved the students' work and damaged their learning, depending on whether it did the work for them.
We build a learning platform, so I have an interest in this subject. I have tried to lean on trials rather than vendors, and to say where the evidence runs out, which in Arabic is almost immediately.
Why this is urgent
AI use among students went from common to universal in two years. In the UK, the share of undergraduates using it at all went from 66% in 2024 to 92% in 2025 and 95% in 2026, according to HEPI's annual survey. In 2026, 94% said they use it to help with assessed work.
The guidance has not kept up. The Digital Education Council's 2026 survey of 27,284 students and 18,114 faculty in 35 countries found 88% of students using AI in their learning, and only 29% who think their instructors are equipped to guide them. Gallup found in 2026 that six in ten US teachers use AI and just 18% have received any formal guidance. In RAND's 2025 survey, more than 80% of US students said nobody had ever explicitly taught them how to use it.
The workplace is moving at the same time. Employers in the World Economic Forum's survey expect 39% of workers' core skills to change by 2030, and 40% in Saudi Arabia. In the WEF's picture of a workforce of 100 people, 59 need training by 2030, and 11 of them are unlikely to get it. Stanford's Digital Economy Lab found that employment of 22 to 25 year-olds in the most AI-exposed jobs is now 19% below where it would be had it kept pace with less-exposed peers, mainly through fewer hires. The authors say that is descriptive, not causal, and the New York Fed found junior and senior postings in exposed jobs moving in parallel. The debate is open. If the Stanford reading is right, though, the junior jobs where people used to learn a trade by doing it are thinning out, and more of that learning will have to happen deliberately, in training.
What the trials say
The good results are real, and they share a design.
Study | Setting | What the AI did | Result |
|---|---|---|---|
Kestin et al., Scientific Reports 2025 | 194 Harvard physics students, crossover trial | A tutor built on research-based teaching, used at home | 0.73 to 1.3 SD more learning than an active-learning class, in a median 49 minutes against about 60 |
De Simone et al., World Bank 2025 | Nine public schools in Benin City, Nigeria | GPT-4 in twelve after-school sessions, in pairs, with a teacher guiding | 0.31 SD overall, 0.24 SD in English, about $48 a pupil |
Contractor and Reyes, preprint 2026 | 211 Middlebury undergraduates | Access to AI during study | 0.27 SD on test scores, still there a week later |
Harrison et al., preprint 2026 | 644 GCSE science students in England | An AI revision tutor against self-directed revision | 0.33 SD |
Wang et al., 2025 | 900 tutors and 1,800 K-12 students in the US | AI suggestions for the human tutor | Students 4 percentage points more likely to master a topic, 9 with lower-rated tutors |
Anthropic, 2026 | 52 mostly junior developers learning a new library | A coding assistant | 50% on a later quiz against 67% for those who coded by hand |
Bastani et al., PNAS 2025 | About 1,000 Turkish high-school students | GPT-4 with and without tutoring guardrails | Plain version: exam 17% worse. Guardrailed version: no penalty |
The effect sizes come from different designs and are not directly comparable. The pattern across them holds. The positive results come from tools that make the learner do the thinking: a tutor that asks before it tells, a teacher in the room, pairs of students talking through the problem. The Nigeria team put its result in years of schooling. The English gain equals about 1.5 years of business-as-usual schooling there, and the total gain about two. The negative results come from tools that produce the answer. In the Anthropic study, the AI users who still scored well were the ones who asked follow-up questions, asked for explanations alongside the code, or asked only conceptual questions. The Middlebury team found the same split and called it augmentation against automation.
One number is missing from that table on purpose. A 2025 meta-analysis of ChatGPT and learning reported a very large average effect of 0.867 and was retracted this spring over discrepancies in its analysis. A 2026 replacement covering 35 studies puts the average at 0.67. Averages across tools this different hide the variable that matters, which is what the tool was allowed to do.
Feeling that you learned is a different thing
The oldest finding in this area predates chatbots. In Roediger and Karpicke's 2006 experiment, students either re-read a passage four times or read it once and then tested themselves three times.
After five minutes, re-reading won. After a week, the students who tested themselves recalled 61% against 40%, having read the text about a quarter as many times. The re-readers were more confident they would remember. Robert and Elizabeth Bjork's name for this is desirable difficulty: conditions that make performance improve quickly "often fail to support long-term retention and transfer," which leaves learners and teachers "vulnerable to mis-assessing whether learning has or has not occurred."
AI widens that gap between how learning feels and what it produces. In METR's 2025 trial, experienced open-source developers took 19% longer to finish real tasks with AI tools, and afterwards estimated that the tools had made them 20% faster.
The same shape shows up elsewhere, in smaller or weaker studies that point the same way. In an MIT Media Lab preprint, 15 of 18 participants who wrote essays with ChatGPT could not correctly quote a sentence from the essay they had just written, against 2 of 18 in the other groups. In a before-and-after study of 19 experienced endoscopists in Poland, the rate at which they found adenomas without AI fell from 28.4% to 22.4% after AI was introduced into their clinics. In a Microsoft and Carnegie Mellon survey of 319 knowledge workers, higher confidence in the AI went with less critical thinking.
None of these is conclusive alone. The MIT study is a small preprint, the endoscopy study is observational, and the survey is self-reported. Together with the trials, they describe a skill that stays in the person only if the person keeps using it.
How it can be done
The learning science is settled enough to design around. Dunlosky and colleagues' 2013 review rated ten study techniques and gave only two a high rating: practice testing and spacing practice out over time. Five got a low rating, including summarising and re-reading. A general-purpose chatbot makes summarising effortless, so it is very good at producing the low-value activity on demand. The design work is to point it at the high-value ones.
Make the learner answer first. Ask for the attempt before the explanation: a question, a worked problem, a short recall from memory. The model is useful for writing those questions, for checking the answer and for explaining the mistake afterwards. In the Anthropic study, asking for explanations was one of the habits of the AI users who still scored well.
Put guardrails in the tool, not in the student's willpower. Bastani's guardrailed tutor gave hints and withheld full answers, and it removed the exam penalty. OpenAI, Google and Anthropic all ship study or learning modes now. When ChatGPT's study mode launched, OpenAI's Leah Belsky said there is no way for parents or administrators to lock a student into it, so the guardrail is a button the student can switch off at eleven at night before a deadline. A course that relies on a guardrail has to supply one that stays on.
Keep a teacher in the loop. The Nigeria sessions had a teacher guiding pairs of students. The largest gain in the Tutor CoPilot trial went to students of the weakest tutors, because the AI supported the tutor rather than replacing them. The OECD's 2026 Digital Education Outlook asks for educational AI "designed with teachers, enabling them to monitor students' interactions."
Test later, without the tool. Bastani's result only appeared when the AI was taken away. Any course where every assessment allows AI measures the student and the model together. The University of Sydney splits assessment into two lanes: secured, in-person assessment of learning, and open assessment for learning where AI is allowed. Princeton faculty voted this year to proctor every exam, ending an unproctored honour-code tradition that dated from 1893.
Do not police with detectors. In Liang and colleagues' 2023 study, seven AI detectors misclassified on average 61% of TOEFL essays written by non-native English speakers as machine-written. A 2026 preprint found honest light AI editing flagged 38% to 80% of the time, while text run through "humanizer" tools was detected less than 4% of the time. For students in the Kingdom writing in their second language, that first number is a direct risk. Vanderbilt turned its detector off in 2023 over false positives.
In company training
The same rules apply at work, with one addition. The earlier piece on this journal about expertise as the bottleneck argued that AI rewards the expertise a person already has. The trials above say how that expertise gets built: by producing, being wrong, and being corrected.
A prompt-writing workshop does not do that. What does is the unglamorous version: domain training that assesses whether the person can defend the work, with the AI allowed in practice and absent when they are assessed. Microsoft's 2026 Work Trend Index attributes 67% of the difference in AI impact to organisational factors and 32% to individual ones. When managers model AI use, reported critical thinking rises 22 points. In US companies, training spend rose to $102.8 billion in 2024-25 while training hours per learner fell from 47 to 40, according to Training Magazine's annual report. If AI takes over the easy hours, the remaining ones need to be the hard ones.
In Saudi Arabia
The Kingdom has moved faster than almost anyone on access:
An AI curriculum runs "across all stages of general education" from the 2025-26 school year, reaching more than six million students.
The Cabinet declared 2026 the Year of Artificial Intelligence on 10 March.
SDAIA's SAMAI programme passed one million people trained in November 2025, ahead of its three-year target, and SDAIA reported 1.56 million beneficiaries across its programmes in September.
A national data and AI curriculum is mandatory for every undergraduate, with 14 universities signed up in January.
Microsoft pledged in February to help a further three million people in the Kingdom gain AI skills by 2030, including more than 500,000 educators.
The demand is there too. In the WEF survey, Saudi employers expect 45% of work tasks to be done mainly by technology by 2030, and 38% plan to drop degree requirements, against 19% globally.
Two things are missing. First, none of the published numbers measure learning. They count people enrolled, trained or certified, and the Turkish trial shows how far a score earned with the assistant can sit from what a student can do without it. Second, I could not find a single randomised trial of AI tutoring in Arabic. Every result in this piece comes from English, Turkish or Nigerian classrooms, working with models trained on a web where Arabic is about 0.65% of the text, less than Czech. I also could not find the curriculum's weekly hours or grade-by-grade structure, so from outside it is hard to say how much of it is practice and how much is exposure.
What needs to be known
If you are | Do | Stop |
|---|---|---|
A student | Try the problem before you ask. Use the model to quiz you and to explain your mistakes | Asking it for the finished answer to anything you will be tested on |
A teacher or lecturer | Keep one assessment that is in person and without AI. Use AI for question banks and planning | Trusting a detector's verdict on a student's work |
A university | Separate assessed learning from open practice, as Sydney does | Treating licence roll-outs as an AI strategy |
An L&D lead | Assess people defending their work, without the tool | Reporting completions and hours as if they were capability |
A ministry or regulator | Fund trials that measure learning in Arabic, with an unaided test | Counting enrolments as the outcome |
What I think
The evidence is consistent enough to act on. AI that makes the learner do the work helps, sometimes by a lot, and AI that does the work for them harms learning while making it feel better. The trials so far land on one side of that line or the other, and the line is a design choice made by whoever builds the tool or the course.
The gap I care about most is Arabic. The Kingdom is putting AI in front of six million schoolchildren and every undergraduate, and the evidence base for how an Arabic-speaking student learns with an AI tutor is empty. The trials that would fill it are cheap by the standards of this year's announcements. Nigeria's cost about $48 a pupil. Somebody should run one here, with an exam at the end that the AI does not sit.
For a course you run this term
Keep at least one assessment in person and without AI
Make every AI-assisted exercise start with the learner's own attempt
Use AI to write retrieval questions, and space them over weeks
Choose tools where the guardrail cannot be switched off by the learner
Drop AI detectors from any decision about a student
Measure learning a week later, not at the end of the session
One email when we publish. Research, product decisions, and what teams report back.







