If AI Democratizes Knowledge, Why Are Students Falling Behind? A Study of 26,811 Students Shows Why Understanding Cannot Be Outsourced
AI has reduced the cost of accessing knowledge. It has not removed the effort required to build understanding.
Artificial intelligence should, in principle, be a powerful force for educational equality.
In the past, whether a student received personalized guidance often depended on whether the family could afford private tutoring, whether the school had sufficient resources, and whether a patient teacher was available to answer questions. Today, with an internet-connected device, a student can ask AI to explain a mathematical concept, compare historical perspectives, revise an essay, or restate a difficult idea at an appropriate level and in a familiar language.
Not every family can afford an excellent private tutor. A knowledgeable AI tutor, however, can respond at almost any time and serve countless students at close to zero marginal cost.
If knowledge is becoming easier to access, students should logically be able to learn faster, more broadly and more deeply. Yet this leads to a counterintuitive question:
If AI allows every child to consult a highly knowledgeable tutor at an unprecedentedly low cost, why does the research show higher homework scores but lower exam performance, rather than better learning for everyone?
A 30-month study of 26,811 Chinese secondary school students captures precisely this paradox.
If AI has lowered the barrier to knowledge, why has that not automatically translated into better learning?
The problem may be that we are treating two related but fundamentally different things as if they were the same: gaining access to knowledge and forming an understanding of it.
What Did the Study Find?
The study, The Generative AI Learning Penalty: Evidence from Chinese Secondary Education, was written by David Strömberg, Victor Lei and Yanhui Wu and released in 2026 as CEPR Discussion Paper DP21577.
It draws on 30 months of panel data from 26,811 Chinese students in grades 7 to 12. The dataset covers nine subjects and includes homework scores, homework completion time, monthly closed-book examinations, and high-school and university entrance examination results. The researchers used differences in when students began adopting generative AI and applied a staggered difference-in-differences approach to estimate changes before and after adoption.

The results reveal a striking contradiction:
Homework scores increased by an average of 18% after students adopted generative AI.
The time required to complete homework fell by 30%.
Monthly closed-book exam scores declined by 20% within six months.
Scores in two high-stakes entrance examinations declined by 18% and 24%, with the full negative gap emerging after approximately two years.
The estimated learning loss was largest in social-science subjects, followed by STEM and language subjects. The negative gap was also greater among younger students, initially higher-performing students and male students.
In other words, students completed their work more quickly and submitted answers that looked better. But when assessment shifted from the quality of the submitted work to the students’ own performance, the trend moved in the opposite direction.
This Is Not “Brain Decline”
These findings deserve serious attention, but they should not be exaggerated into claims that AI physically damages students’ brains or simply makes children less intelligent.
The study measured homework performance, completion time and examination results. It did not measure brain structure, intelligence, attention or other neurological and cognitive indicators. The findings support a warning about learning performance and skill formation; they do not, by themselves, demonstrate physiological cognitive decline.
It is also important to note that this is a CEPR discussion paper, not a peer-reviewed journal article. Although the authors used longitudinal data and a quasi-experimental method, students were not randomly assigned to AI and non-AI groups. Differences between groups and other events occurring at the same time may therefore not have been completely eliminated.
The research also took place within the context of Chinese secondary education. Its examination system, homework culture and patterns of AI use may not transfer directly to every country, age group or subject. The percentages should be understood as average effects estimated from this particular dataset and model—not as outcomes that every student will inevitably experience after using AI.
The Issue Is How AI Is Used
One of the most important details in the paper is that not every AI user experienced the same degree of decline.
The researchers found that the negative results were concentrated among approximately 80% of AI users. These students displayed behaviour consistent with “homework outsourcing”: they completed assignments unusually quickly while receiving very high homework scores.
By contrast, students who continued to spend roughly as much time on homework as non-users after adopting AI experienced substantially smaller learning losses.
This classification was inferred from completion time and homework scores. The researchers did not directly observe every interaction between every student and an AI system. Nevertheless, the pattern points to an important possibility:
What matters may not simply be whether a student uses AI, but which part of the learning process AI replaces.
If AI is used to explain a concept, generate counterexamples, compare perspectives or identify blind spots, it may deepen a student’s engagement with knowledge. If it is used to generate a submission-ready answer, the time it saves may include the very practice through which understanding would otherwise have developed.
Understanding Cannot Be Outsourced
During a conversation at Sequoia Capital’s AI Ascent 2026, AI researcher Andrej Karpathy highlighted an idea that he felt captured the issue well:
“You can outsource your thinking, but you can’t outsource your understanding.”
Karpathy introduced the line with language suggesting that he was recalling or endorsing a formulation he found compelling. It is therefore more accurate to say that he quoted or emphasized the idea, rather than assert that he was necessarily its original author.
The statement does not tell students to reject AI, nor does it mean that every cognitive task must be performed without assistance. AI can help people search, organize, compare, calculate and communicate. What cannot simply be transferred is the understanding formed through those activities.
An answer can move from an AI system to a student’s screen within seconds. It does not automatically become part of the student’s mental model. A student may feel that an answer “makes sense” without being able to explain:
Which facts support the answer?
What steps connect the evidence to the conclusion?
Which assumptions does the answer rely on?
Are there other reasonable explanations?
Which parts are facts, which are inferences, and which are value judgements?
Why should the answer be trusted?
In an educational context, Karpathy’s idea can therefore be extended:
Humans can think with AI, but they cannot outsource the formation of understanding—or the responsibility for judging the answer.
A Very Unusual Teacher
AI can be a knowledgeable, responsive teacher capable of adjusting an explanation to a student’s level. But it is also a highly unusual teacher: it may misremember facts, omit essential conditions, accommodate the user’s assumptions, or express a false answer in fluent and confident language.
The US National Institute of Standards and Technology describes this problem as “confabulation”: generative AI can confidently present erroneous or false content. NIST explains that this risk arises from the way generative models produce material based on statistical patterns in their training data. The same mechanism can generate a correct response, a factual error or an internally inconsistent explanation.
Fluency is therefore not the same as reliability, and a confident tone is not evidence.
The same model may also produce very different results depending on how it is used. Output quality can be affected by whether the problem is clearly defined, whether relevant context and constraints are supplied, whether the student requests a finished answer or progressive hints, whether the system is asked to provide sources and uncertainty, and whether the student follows up critically.
This does not mean that a “perfect prompt” can guarantee a correct answer. Better use can improve relevance, transparency and verifiability, but it cannot eliminate the model’s underlying risk of error.
AI Is Not Inevitably Harmful
A field experiment involving nearly 1,000 secondary-school mathematics students offers an important comparison. The results were published in 2025 in the Proceedings of the National Academy of Sciences.
The researchers developed two GPT-4-based tools:
GPT Base: An interface resembling relatively unrestricted use of a general-purpose chatbot.
GPT Tutor: A system with educational guardrails designed to guide students through problems and reduce the direct delivery of final answers.
While students were allowed to use the tools during practice, performance increased by 48% in the GPT Base group and by 127% in the GPT Tutor group. In a later assessment without AI access, however, the GPT Base group scored 17% lower than students who had never used the tool. The negative learning effect was largely mitigated in the GPT Tutor group.
The experiment shows that even when the underlying model is GPT-4, “completing the task for the student” and “guiding the student through learning” can produce very different outcomes.
A separate randomized controlled trial at Harvard found that a carefully designed AI tutor helped university students achieve greater learning gains in less time than in-class active learning. That study was published in the peer-reviewed journal Scientific Reports in 2025, with DOI 10.1038/s41598-025-97652-6.

These findings are not necessarily contradictory. Together, they suggest that the question cannot be reduced to whether AI is good or bad for education:
AI’s effect on learning depends substantially on how the tool is designed and how the student interacts with it.
Understanding Is Also a Prerequisite
“Understanding cannot be outsourced” does not merely mean that students should understand an AI-generated answer after receiving it. The deeper issue is that understanding is not only a desired outcome of using AI; it is also a prerequisite for using AI well.
The better a student understands a problem, the more clearly they can define the objective, conditions and constraints. The better they understand the nature of AI, the less likely they are to mistake fluency for truth. The better they understand the structure of an answer, the more capable they become of detecting unsupported assumptions and gaps in reasoning.
This produces two very different cycles:
| Passive dependence | Active understanding |
|---|---|
| Ask immediately without understanding the problem | Clarify the problem, objective and constraints first |
| Receive a polished, complete answer | Request explanations, evidence and alternative views |
| Treat confidence as proof of correctness | Distinguish facts, inferences and uncertainty |
| Accept or submit the answer directly | Question, compare, verify and revise |
| Become less able to judge the next answer | Deepen understanding and become better at using AI |
True AI literacy is therefore not merely the ability to write prompts. It consists of four forms of understanding:
Understand the problem: What am I actually trying to address? What are the objective, conditions and constraints?
Understand AI: What does it do well? Why might it make mistakes, accommodate my assumptions or reproduce bias?
Understand the answer: What evidence supports the conclusion? What reasoning produced it? What are its limitations?
Understand yourself: Why am I accepting this answer? What goals and value trade-offs am I making?
AI can help us identify ways to achieve a goal. It cannot fully determine which goals are worth pursuing, what costs we should accept, or who should ultimately take responsibility for the result.
What Should Parents Ask?
Parents may be less familiar with AI than their children. They do not need to become experts in large language models before they can help. Rather than focusing only on whether a child used AI, parents can cultivate a more thoughtful way of approaching AI-generated answers.
The following questions can provide a starting point for family conversations:
How do you understand the answer AI gave you?
What information and assumptions does the answer depend on?
Is AI presenting a fact, an inference or an opinion here?
Which parts do you still not understand?
What might the answer have overlooked?
Would the conclusion change if you approached the problem from another perspective?
What makes you believe the answer is reliable?
Have you found independent information that supports or contradicts it?
How did the AI response change your original understanding?
Do you agree with its conclusion? Why?
The question “Do you know how AI constructed this answer?” also requires some refinement. A student usually cannot infer the model’s actual internal computational process from its output, and an AI-generated retrospective explanation may not faithfully describe how the answer was produced.
A more precise question is:
Can you reconstruct the reasons, evidence, inferences and assumptions that would need to be true for this answer to hold?
The aim is not to require children to explain every parameter inside a model. It is to help them understand the knowledge structure behind the answer.
Equal Access to Knowledge Is Not Equal Access to Understanding
We can now return to the question raised at the beginning: if AI has reduced the cost of accessing knowledge, why might students’ learning performance decline?
The answer is not that AI provides too much knowledge. It is that acquiring information and forming understanding have never been the same thing.
AI reduces the cost of searching for information, obtaining explanations and generating answers. It does not eliminate the cognitive engagement required to understand a problem. When students use AI to question concepts, compare perspectives, inspect reasoning and discover blind spots, it can be an excellent tutor. When students hand over the problem and receive a finished response, the system may also remove part of the intellectual activity through which understanding would have formed.
The study therefore does not show that AI and educational equality are incompatible. It points to a more difficult conclusion:
AI can democratize access to knowledge. Whether it can democratize understanding still depends on education.
AI may give more children access to the same body of knowledge, but whether they benefit depends on whether they understand the problem, know that AI can be wrong, and have the ability to question, evaluate and revise its output.
If schools and families do not teach these skills, AI may create a new kind of educational divide:
In the past, inequality may have depended on who had access to a good teacher. In the future, it may depend on who knows how to understand, question and make good use of the same AI teacher.
The task of education is not to keep children inside an old world without AI. Nor is it to leave them unprepared in a new world where AI-generated content is accepted without question. Children should learn about AI early, understand both its capabilities and limitations, and learn how to explore knowledge with it.
The idea emphasized by Karpathy—that understanding cannot be outsourced—is not a demand that we refuse AI assistance. It is a reminder that understanding is what prevents us from becoming subordinate to AI-generated answers. It is also what allows AI to become more than a machine that completes tasks for us: a teacher that helps us ask better questions and make better judgements.
When every child can consult an AI teacher whose knowledge is broader than that of any single person, education may no longer be primarily about producing students who remember more answers. It may be about developing people capable of understanding the teacher, questioning the teacher, correcting the teacher and, ultimately, going beyond the teacher.
Research Sources
Strömberg, D., Lei, V., & Wu, Y. (2026). The Generative AI Learning Penalty: Evidence from Chinese Secondary Education. CEPR Discussion Paper No. 21577. RePEc / CEPR paper record
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. DOI: 10.1073/pnas.2422633122
Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. DOI: 10.1038/s41598-025-97652-6
National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST publication
Sequoia Capital. (2026). Andrej Karpathy: From Vibe Coding to Agentic Engineering. Interview video


