Children still outlearn AI on language despite vast data gap
This digest was compiled by AI from multiple sources — links to the originals are below.

Children learn language with vastly less data than large language models, a gap researchers call the data efficiency gap. Stanford cognitive scientist Michael C. Frank notes that LLMs require scraping the entire sum of human knowledge to recreate a milestone children reach in a year. Understanding this gap could lead to more efficient AI models and insights into developing minds.
Key Facts
- An LLM can process over 100,000 times more words than a person encounters while mastering their native language.
- Meta's Llama 3.1, released two years ago, was pretrained on 15 trillion tokens.
- A preteen in a linguistically rich home may have heard around 100 million words, rising to 300 million by age 20 with literacy.
- Georgetown University cognitive scientist Ethan Gotlieb Wilcox says frontier models could be pretraining on 10 times more data than Llama 3.1.
The Data Efficiency Gap
The difference in learning efficiency between children and large language models is called the data efficiency gap. Stanford University cognitive scientist Michael C. Frank contrasts the progress of LLMs with the fact that they require scraping the entire sum of human knowledge to recreate a milestone children reach in a year. Georgetown University cognitive scientist Ethan Gotlieb Wilcox notes that Claude has seen the amount of language an entire city will experience in one generation.
Scale of Model Training
Meta's open-weight LLM Llama 3.1, released two years ago, was pretrained on 15 trillion tokens. Frontier models could be pretraining on 10 times more data, according to Ethan Gotlieb Wilcox. The well of easily available training data could run dry as early as the 2030s.
Human Learning Benchmarks
A preteen raised in a linguistically rich home may have heard around 100 million words. Adding literacy can boost that word count to about 300 million words by age 20. If all the words used to train a modern LLM were printed on paper, the stack would reach past the International Space Station.
1 source
Children still outlearn AI on language despite vast data gap






