In January 2025, the performance of the best artificial intelligence models on a test designed to resist their capabilities was still very limited: GPT-4o scored approximately 5% on ARC-AGI-1, while o3, announced in late December 2024, reached 75.7% at standard budget and 87.5% in high compute mode on the same benchmark—which was already largely circumvented. On ARC-AGI-2, launched in March 2025, the best models scored approximately 3%; thirteen months later, this rate reached 83%. This is not an anecdote about the speed of technological progress. It is a structural problem for schools.

Educational institutions operate on cycles of five to ten years: curriculum overhauls, teacher training, tool adoption. AI models change generations every six to nine months. The gap between these two rhythms is not a lag in upgrades. It is an incompatibility of regime. And what it reveals goes beyond the question of tools: it forces a reformulation of what schools are supposed to transmit, and to whom they transmit it.

The Essentials

  • On ARC-AGI-2, the best models scored approximately 3% at launch in March 2025; thirteen months later, this rate reached 83%. This data is published by ARC Prize (François Chollet, Mike Knoop et al.) and monitored by Epoch AI.
  • Models change generations every six to nine months, whereas institutional pedagogical cycles last five to ten years.
  • AI amplifies existing skills: students and workers who already know how to structure a problem benefit more from it than those who do not yet know how.
  • The stakes in the next two years: define what portion of human cognition AI cannot substitute, and organize schools around this frontier—which is constantly shifting.

In Just Months, AI’s Glass Ceiling Evaporated

The benchmark in question was not an ordinary test. It had been specifically designed to map what models could not do: problems of complex reasoning, multi-step inference, deep contextual understanding. On ARC-AGI-1, GPT-4o still scored ~5% in early 2024 before o3 massively circumvented it in late December 2024. On ARC-AGI-2, launched in March 2025, the best models started again at ~3%—the benchmark was once again fulfilling its exact function as a mirror of limitations.

Thirteen months after the launch of ARC-AGI-2, that mirror reflects something else: 83% success rate. This is not the same technology progressing gradually: these are multiple generations of models succeeding one another, each surpassing the previous one in capacities that, in the previous generation, seemed structurally inaccessible. The researchers who design these benchmarks have said so themselves: they struggle to build tests that resist for more than twelve months.

This rhythm changes the nature of the problem. When a technology progresses slowly, institutions can adapt through cycles. When it progresses through discontinuous leaps every six to nine months, institutional adaptation cycles become structurally insufficient. The question is no longer “how do we integrate AI into schools?” but “what kind of school is still possible when AI redefines the frontier of cognitive capacities faster than curricula reform?”

Olivier Babeau, a liberal economist, professor of management sciences and essayist on digital transformations, poses this question in its sharpest terms: AI renders traditional curricula and the classical relationship to effort obsolete. He addresses the economics of attention in his books, without being an academic specialist in the strict sense. If the tool can write an essay, solve a math problem, generate an outline for a presentation, what is the point of learning to do it yourself? The provocation is intentional. But it points to a real pedagogical malaise.

AI Amplifies What You Already Have

The school system’s instinctive response is often the same: technology helps those who already know, so schools should first transmit fundamentals before authorizing tools. This reasoning is correct on one point and dangerous on another.

It is correct because data confirms the amplification effect. Workers and students who already master the structure of a problem, who know how to formulate a precise request, evaluate an answer, identify a reasoning error—those people derive substantial benefit from generative AI. Those who do not master these competencies receive answers they cannot critique, and find themselves more dependent rather than strengthened.

Daron Acemoglu and Simon Johnson have documented this mechanism at a different scale in their work on technology and power: the gains from an innovation do not distribute spontaneously. They concentrate first where complementary skills already exist. Generative AI follows this logic with almost cruel precision. The analysis Acemoglu and Johnson make of it suggests that this concentration is not inevitable, but it does not correct itself.

The institutional reasoning becomes dangerous when it draws the opposite conclusion: since AI amplifies existing competencies, schools should first ignore AI and concentrate on fundamentals as before. This position confuses pedagogical sequence with isolating the student from reality. A student who leaves high school without ever having used these tools is not better equipped to use them with discernment. He is simply behind.

When Schools Adapt Their Exercises, AI Adapts Its Responses

The fundamental problem is not tool adoption. It is the speed at which reference points collapse.

Teachers who began rethinking their assessments in 2023, in response to ChatGPT’s emergence, designed formats meant to resist AI: personal questions, contextualized exercises, oral work. Two years later, some of these formats are already circumvented by multimodal models capable of processing speech, simulating voice, generating contextualized responses from a profile built on a few previous exchanges.

This is not a race lost from the start. It is a race whose nature is poorly understood. The objective cannot be to build exercises impervious to AI. The objective is to build learning that has value even if the student has access to AI. And these two objectives do not overlap.

The emergence of AI Scientist, which publishes, experiments, and undergoes peer review quasi-autonomously, illustrates an analogous shift in academic research: tasks that seemed to define researcher competence are absorbed by the tool, and the question of what remains properly human becomes urgent. Schools pose the same question, but with additional stakes: their students do not yet have the fundamentals that would make this collaboration fruitful rather than parasitic.

The Teacher as Coach of Judgment, Not Memory

The lever that data reveals is not technological. It is pedagogical, and it concerns the teacher’s role.

Experiments that work share a common architecture: the teacher does not transmit content that the student reproduces; he creates conditions in which the student must exercise a judgment that AI cannot exercise on their behalf. Formulating an intention. Evaluating the relevance of an answer in light of a real problem. Identifying what the tool cannot say. Making a decision in an ambiguous context where data is incomplete.

These competencies are not acquired by watching AI work. They are acquired by doing, by making mistakes, by trying again under the pressure of a human interlocutor who knows where the tool’s limit lies. This is the role that some education practitioners and researchers identify as the pivot of pedagogical transformation: not the teacher as living encyclopedia, nor as guardian of rules against the tool, but as trainer of cognitive effort and critical judgment.

Babeau poses a real tension here: if schools do not transform fast enough, this role migrates to other actors. Training platforms, private coaches, professional environments. This is not an argument against public schools. It is an argument for taking seriously the speed at which their model must evolve. Intellectual emancipation is not a liberal luxury: it is the condition for students to become users of AI rather than passengers.

The tension between this liberal vision and the institutionalist approach of an Acemoglu is not resolved on the pedagogical terrain alone. One insists on the transformation of individual practices; the other reminds us that without deliberate public investment in teacher training and digital infrastructure, transformation benefits schools that already have the means. The data vindicate both at different levels: transformation is necessary, and it will be unequally distributed if left to market dynamics alone.

Will the Long Degree Lose Its Value, or Change It?

The ten-year projection opens a question that education economists are beginning to formulate seriously: if AI continues to progress at the rate documented since 2023, what is a long education worth in twenty years?

The argument against long-term training is simple: much of the stock of knowledge that justified five or seven years of study becomes accessible in a few well-formulated requests. Law, accounting, part of diagnostic medicine, professional writing. If the tool does the work, why pay for learning time?

The argument for long-term training is different from what it once was. It no longer rests on the volume of knowledge transmitted, but on two things that long time allows and short time does not. First, the formation of judgment in situations where data is insufficient, contradictory, or instrumentalized. Second, the capacity to work with human interlocutors in contexts of high uncertainty, what economists sometimes call complex relational competencies.

These two dimensions are precisely those that AI saturates least well, and probably not by accident: they depend on lived experience, accumulated in real situations with real stakes. AI can simulate a difficult client, a hostile colleague, an ambiguous file. It cannot replace having been confronted by them and having to endure.

If this analysis is correct, the value of long-term training does not disappear. It shifts. Curricula that select on memory and reproduction lose value. Those that train judgment, argumentation, the ability to hold steady before ambiguity gain it. This shift reshuffles the cards between programs, between disciplines, between pedagogies. It will not happen spontaneously: assessment systems, competitions, diplomas carry considerable inertia.

The generational stakes are here. Students entering sixth grade today will be on the job market in 2035. The AI models that will exist then have not yet been designed. Training for an unknown job market with tools whose capacities change every six months forces a conclusion that disturbs established curricula: the only thing one can transmit with certainty is the capacity to learn, unlearn, and reformulate.

Working Experiments Don’t Come from Ministries

Concrete examples exist, but they are scattered. In Finland, high schools have integrated AI into critical reading exercises: students receive a response generated by a model on a literary text and must evaluate its relevance, identify its blind spots, correct it. The tool is not banned; it becomes the object of learning. In France, isolated teachers build similar sequences without waiting for official instructions, often using resources shared on informal networks.

These experiments have in common an inversion of the relationship to the tool: rather than protecting the student from AI, they ask him to examine it. This inversion is pedagogically demanding. It assumes a teacher capable of distinguishing a good answer from a plausible one, a judgment on content and not only on method. This is precisely what makes teacher training the least spectacular and most decisive lever.

The problem of equity restores all its force here. Schools that have the means to train their teachers in this pedagogy, to invest in experimentation devices, to recruit hybrid profiles between pedagogy and digital competence are, in essence, the same ones that already produced the best results. The amplification effect that AI exerts on individuals reproduces itself at the institutional level. Without a deliberate policy for diffusing these practices in the most fragile schools, pedagogical innovation widens the gap it claims to narrow.

What Remains to Be Decided

The question is not resolved. Data on the long-term impact of AI on school learning do not yet exist—the tools are too recent, the cohorts too short. What is observed is a series of converging signals: AI progresses faster than institutions adapt, it amplifies preexisting cognitive inequalities, and the correction lever is human rather than technical.

The open question is that of the tipping point. At what moment does a school that has not integrated AI into its pedagogical practices become a school that prepares its students for a world that no longer exists? This point is not fixed. It shifts with each generation of models. And this may be the most uncomfortable data point of the problem: for the first time in a long while, schools must adapt at a pace they do not control.


Sources

  1. Olivier Babeau, “L’école face au déluge de l’intelligence artificielle”, Le Journal du Dimanche — Institut Sapiens: https://www.lejdd.fr/Societe/olivier-babeau-lecole-face-au-deluge-de-lintelligence-artificielle-171174
  2. Daron Acemoglu and Simon Johnson, Power and Progress, Basic Books, 2023 — analysis of mechanisms for capturing technological gains
  3. Data on AI benchmarks: Scale AI / Epoch AI — monitoring of model performance on complex reasoning tasks (no URL — data accessible via public tracking of MMLU, GPQA and ARC-AGI benchmarks)
  4. Work on the economics of competencies and the value of long-term training: OECD, Education at a Glance 2024 (no direct URL — annual report available on OECD website)
  5. ARC Prize Technical Reports (2024 & 2025)
  6. ARC Prize - o3 Breakthrough (arcprize.org)
  7. Power and Progress - Basic Books 2023
  8. Olivier Babeau - Wikipedia
  9. AI Scientist - Nature (2026)
  10. Epoch AI Benchmarks Hub
  11. ARC-AGI-2 Paper (Chollet et al., May 2025)