Loading prices...
All news
Flat vector illustration of a single glowing cyan point expanding into a dense soundwave-like cluster of particles on a dark background, symbolizing the vast data efficiency gap between how children learn language and how LLMs are trained

Kids learn language on a fraction of the data LLMs need, and nobody knows why

Toddlers typically start producing grammatically correct sentences after hearing somewhere between 10 million and 30 million words. Large language models need trillions of tokens to reach fluent, flexible language use, a gap researchers call the data efficiency gap, MIT Technology Review reported.

Meta's Llama 3.1 pretrained on 15 trillion tokens two years ago, and frontier models today could be training on 10 times more, according to Georgetown cognitive scientist Ethan Gotlieb Wilcox. A preteen raised in a linguistically rich home hears roughly 100 million words, a figure that climbs to about 300 million once reading enters the picture by age 20. Wilcox put the scale gap plainly: a modern LLM has effectively seen as much language as an entire city generates in one generation.

The progress recently has been amazing. But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.

Michael C. Frank, cognitive scientist, Stanford University

Closing the gap matters beyond curiosity. The easily available pool of internet text could run dry as early as the 2030s, and researchers see a more data-efficient training approach as one way frontier labs might keep scaling models once that well tightens. We covered a related piece of that same pressure in OpenAI's $30 billion Georgia data center build, where compute scarcity is reshaping how labs plan capacity alongside the data question.

Researchers Alex Warstadt, Leshem Choshen, and colleagues launched an annual competition called BabyLM in 2022 to test the idea directly. Entrants train models on a corpus of 100 million words, drawn from storybooks, dialogue, subtitles, and transcripts of speech directed at children, then score them on the same grammar benchmarks psycholinguists use with human subjects. The 2024 winner, a model called GPT-BERT, trained on about 100 million words and still beat Meta's Llama 2 70B, trained on roughly 15,000 times more data, on one of the competition's benchmarks.

  • Toddlers reach fluent grammar after roughly 10 to 30 million words of exposure
  • Meta's Llama 3.1 pretrained on 15 trillion tokens in 2024
  • BabyLM's 2024 winner, GPT-BERT, trained on about 100 million words
  • GPT-BERT beat Llama 2 70B, trained on roughly 15,000 times more data, on one benchmark
  • Readily available internet training data could run dry as early as the 2030s

A separate research track tries to close the gap by studying what children see and hear as they take in the world, beyond written text. Stanford's Michael Frank ran a project called SAYCam that recorded two hours a week of three babies' lives between six months and two and a half years old using headcams. Princeton neuroscientist Uri Hasson went further, recording 12 hours a day of 17 children's first 1,000 days for a project described in a 2026 preprint. Neither data set has produced a model that learns language the way a real toddler does.

Berkeley developmental psychologist Alison Gopnik points to a likely reason: children actively explore and experiment with their environment rather than passively absorbing whatever video or text is placed in front of them. Meta researchers have taken an interest, contributing to BabyLM's multimodal track and recently announcing a benchmark for training models on baby headcam footage, but no lab has cracked how to replicate a toddler's curiosity inside a training pipeline.

This piece is informational, not a recommendation to buy, sell, or hold any asset.

Published: 12:35 · 24.08.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!