World Models and How Children Learn
As LLMs have gotten progressively better — winning IMO gold medal, solving long standing problems in several disciplines, coding agents becoming ubiquitous etc. — there has been a push towards Physical AI. To this end, developing World Models has been a rage nowadays. Recently, Jitendra Malik and Yann LeCun shared a brief historical perspective and some recent work on World Models.
It’s fascinating that the early work can be traced back to the early 1940s and how multi-disciplinary — cross-cutting, for example, psychology, cognitive science and neuroscience, model Predictive Control (MPC) — it has been.
It has been interesting to see the learning evolution — Observation → Imitation → Explore-and-Exploit (/reinforcement) — with my 5 yr and 1.5 yr old. I have come across a few instances of the following:
- Carey and Bartlett’s “Fast Mapping” — a mechanism whereby a child can learn a new word or concept from a single, incidental exposure and retain that knowledge long-term.
- Karmiloff-Smith’s “Representational Redescription” — reorganization of current information leading to a child developing explicit knowledge about a complex topic they were never formally taught.
The above, coupled with self-reflection (which has shown benefits to tasks such as coding), could drive learning acceleration and consequently have material implications for science — for example, discovery on the positive side and security on the negative side.