Back to feed

Two-year study finds banning AI from classrooms leaves students worse off

2 min
Two-year study finds banning AI from classrooms leaves students worse off

This digest was compiled by AI from multiple sources — links to the originals are below.

A two-year study at Vrije Universiteit Amsterdam found students banned from using AI in a law course performed worst among three groups. The no-AI group finished last in both 2024 and 2025, while students given structured AI training initially outperformed peers but lost that edge by the second year.

Key Facts

  • The experiment involved 66 students in 2024 and 164 in 2025, randomly split into three groups.
  • The no-AI group finished last in both years, with many subgroups running out of ideas after 10–15 minutes.
  • The trained group scored well above the other two in 2024, especially on the take-home exam, but the gap nearly closed in 2025.
  • Researcher Thibault Schrepel said his assumption that unguided AI would do more harm than good was wrong.

Study Design

Thibault Schrepel of Vrije Universiteit Amsterdam randomly split students in his “Law of AI” course into three groups. All groups worked in teams of four or five and had 20 minutes to improve a provision of the EU AI Act. The first group could not use ChatGPT, the second received AI-generated revision suggestions without guidance, and the third received hands-on training in legal prompt engineering and checking AI suggestions. Grading covered substance, clarity, proportionality, and innovation, and all students took the same multiple-choice and take-home exams.

Group Performance

The no-AI group mostly made minor wording tweaks and many subgroups ran dry after 10–15 minutes, an effect Schrepel calls “idea exhaustion.” The second group accepted AI suggestions largely without question, and every subgroup kept at least one misleading or legally extraneous term from ChatGPT’s output. Only the trained third group engaged in genuine back-and-forth with the AI, testing different phrasings and digging into substantive questions. The trained group scored well above the other two in 2024, especially on the take-home exam, but by 2025 all three groups performed at roughly the same level.

Researcher's Conclusions

Schrepel attributes the narrowing gap to growing chatbot familiarity, as many students already use these tools in daily life. He adds that ethical use and legal responsibility still need to be taught. One finding held constant across both years: the no-AI group finished last, and Schrepel concludes that banning AI produces worse average outcomes than allowing it. Schrepel writes that the second group’s performance surprised him and that his own assumption turned out wrong.

1 source

Time · lag behind first