Curiosity

From PRIMO.ai
Jump to navigation Jump to search

YouTube ... Quora ...Google search ...Google News ...Bing News


Curiosity is the desire to know something for its own sake, rather than for an external payoff the knowledge might bring. Psychology and neuroscience study it as both a passing state and a stable personality trait. In artificial intelligence, it is modeled as intrinsic motivation: an internally generated reward that drives an agent to explore and learn.

This page starts with the engineering question: How do you add curiosity to an AI system, and what does that buy you? It then covers the human science that inspired the work, and finally the "Inside Out" view of the brain, in which a system constantly predicts its world from an internal model and treats the gaps in that model as the most valuable things to investigate.

The problem curiosity solves. Most AI systems learn toward goals that someone else defines. A reinforcement learning agent gets reward from its environment. A large language model learns to predict text, then gets tuned with human feedback or automatically verified answers. Both approaches struggle when rewards are rare (sparse), misleading (deceptive), or simply absent (open-ended). An agent pressing random buttons almost never finds the key in the Atari game Montezuma's Revenge. A reasoning model trained only on whether its final answer was right tends to sharpen strategies it already knows rather than discover new ones. Curiosity addresses this by letting the system reward itself for learning, so it keeps exploring even when the world is silent.

The real challenge. Framing the issue simply as "AI curiosity" implies that AI just lacks a mechanism to explore or wonder. That is not the real problem. The real problem is not generating exploratory behavior, but mathematically formalizing an epistemic drive that distinguishes between productive novelty and useless entropy without collapsing task alignment or safety boundaries. Put in one sentence, the core challenge is engineering an intrinsic epistemic drive that autonomously directs an artificial agent to explore open-ended, complex state spaces for generalizable knowledge without falling into the "noisy-TV" trap of chasing irreducible entropy or breaking operational and safety constraints. Three difficulties stand out: separating learnable novelty from raw randomness (see The Physics), balancing self-directed learning against externally assigned tasks, and the cost of failure outside simulation (both covered in The Mechanism).

Who built the field. Jürgen Schmidhuber proposed "artificial curiosity" around 1990 to 1991: a controller rewarded for improving its own world model. Pierre-Yves Oudeyer and Frédéric Kaplan, first at Sony CSL and then in Oudeyer's Inria Flowers lab, developed learning-progress curiosity for developmental robots and later "autotelic" agents that invent their own goals. Deepak Pathak and colleagues at UC Berkeley (the Intrinsic Curiosity Module, 2017) and Yuri Burda and colleagues at OpenAI (large-scale curiosity and Random Network Distillation, 2018) made curiosity work at deep learning scale. Marc Bellemare and colleagues at DeepMind tied it to count-based exploration. Karl Friston's active inference treats curiosity as "epistemic value." Jeff Clune, Kenneth Stanley, and Joel Lehman pushed it toward open-ended discovery, from Go-Explore to OMNI, the AI Scientist, and the Darwin Gödel Machine. On the human side, the key architects include Daniel Berlyne, George Loewenstein, Celeste Kidd, Benjamin Hayden, Matthias Gruber, Charan Ranganath, and Jacqueline Gottlieb.

The central mechanism. A curious agent adds a self-generated bonus to whatever reward the world provides. That bonus comes from the agent's own internal model: how new a situation is, how badly the model predicted it, how fast the model is improving, or how much uncertainty an action would remove. The agent then learns to seek out situations that score high on that bonus.

The paradigm shift. Conventional systems are built "outside in": a stimulus arrives, a response comes out, and exploration is random noise added on top (epsilon-greedy actions, entropy bonuses, sampling temperature). Curious systems are built "inside out": they maintain a generative model of the world, run predictions forward, and direct their attention and actions toward the specific places where that model is wrong or improving. Exploration becomes directed rather than random. For language models, the shift is from a reactive system that waits for a prompt to an agent that generates its own questions, goals, and experiments.



Curiosity makes learning its own reward. A curious system runs an internal model of the world from the inside out, notices where that model is wrong or improving, and steers itself toward those gaps instead of waiting for the world to hand out rewards. Machines already do this reliably in games and narrow domains; doing it with ideas, over a lifetime of retained knowledge, remains an open problem.


The Mechanism

Every curiosity-driven system, from a 2017 Atari agent to a 2026 automated research agent, runs some version of the same loop.

  1. Encode. The agent turns raw input (pixels, sensor readings, text) into an internal representation. This choice quietly decides what counts as "new." The Intrinsic Curiosity Module (ICM), for example, learns features through an inverse-dynamics task (predict which action caused a transition), so the representation ignores things the agent cannot influence, like leaves blowing in the background.
  2. Predict. A world model forecasts something: the next state, the outcome of an action, or the agent's own future competence at a goal.
  3. Compare. A curiosity module computes the bonus by comparing prediction with reality. Depending on the design, it measures novelty (how rarely this state has been seen), prediction error (how wrong the forecast was), learning progress (how much the forecast is improving), information gain (how much beliefs shifted), or a foundation model's judgment of what is "interesting."
  4. Reward and act. The bonus is added to any external reward, and a policy learns to seek high-bonus situations or goals.
  5. Learn. The world model trains on the new experience. What was surprising becomes predictable, so the bonus there fades. Curiosity is self-extinguishing locally, which pushes the agent's frontier outward.
  6. Remember and return. Episodic memories or archives store promising frontiers so the agent does not forget them. Go-Explore showed that without this, agents suffer "detachment" (they abandon frontiers they found) and "derailment" (random exploration knocks them off the path back).
  7. Choose goals (in autotelic agents). A goal generator, which can now be a language model, proposes what to practice next, typically prioritizing goals where learning progress is highest.

Major components. A representation learner, a world model, a curiosity module, a policy, a memory or archive, and in more advanced systems a goal generator and a meta-controller. Agent57, for instance, used a meta-controller to choose how strongly to weight exploration versus exploitation for each episode.

Dependencies and conditions. Curiosity only pays off when the world has learnable structure, when the representation captures what matters, and when the bonus is weighted sensibly against the task reward. That weighting is a deep problem, not a tuning detail. Modern AI architectures are strictly teleological: they are optimized to minimize a defined loss function or maximize an explicit human-provided reward. True curiosity requires an agent to sacrifice immediate, externally assigned task performance in pursuit of self-directed learning. Balancing the two without inducing task abandonment, reward hacking, or uncontrollable goal drift is an open theoretical problem. The bonus is also non-stationary (it changes as the model learns), which makes training harder than with a fixed reward.

Physical robots add safety constraints: a curious robot arm should not "discover" what happens when it hits a person. The cost of failure is far higher outside simulation. Biological curiosity is constrained by millions of years of evolutionary priors that hard-code self-preservation (pain, fear, resource depletion). AI lacks these pre-linguistic, embodied priors, so an unconstrained curious agent in real-world deployment, such as physical robotics or autonomous network operations, will inevitably test fatal, destructive, or irreversible boundaries to acquire missing information.

How this maps onto language models. For an LLM, the "state" is the conversation or reasoning trace so far, an "action" is a token or full response, and the model itself is the world model. Perplexity (how surprised the model is by its own output) can serve as a prediction-error signal. Asking a clarifying question, running a search, or executing code are information-gathering actions. Writing to persistent memory is the "remember" step that most deployed systems still lack.

Where uncertainty remains. Researchers still disagree about which curiosity signal best matches human exploration, how to make curiosity operate over ideas rather than raw states, how to prevent agents from being captured by meaningless novelty, and how to tell whether a language model is genuinely curious or just producing curious-sounding text.

Why This Matters

  • Scientific importance. Curiosity is one of the few places where AI, psychology, and neuroscience share math. Learning progress, information gain, and prediction error are used both to build agents and to model human and animal behavior. Agents provide testable hypotheses about the brain, and brain data constrains which agent designs are plausible.
  • Engineering implications. Curiosity is a practical tool for sparse-reward problems: robotics, games, scientific search, and now the training of reasoning models, where pure reward-maximization collapses exploration. Better exploration means less hand-crafted reward shaping and fewer human demonstrations. Without it, autonomous agents deployed in dynamic, non-stationary environments, such as planetary exploration, robotic field operations, or deep cyber-defense, will catastrophically fail when they encounter edge cases that human designers failed to anticipate and explicitly incentivize.
  • Applications. Curiosity-driven methods already appear in automated curriculum learning, procedural content and environment generation, automated red-teaming of language models (where a curiosity bonus pushes an attacker model to find diverse failure modes), and education. Oudeyer's team reports that learning-progress-based personalization algorithms are used in adaptive math software deployed in more than 70,000 French classrooms.
  • Practical consequences for everyday AI. A model that knows what it does not know can ask a clarifying question instead of guessing, search before answering, or say "I'm not sure." These are small acts of curiosity, and they directly reduce hallucination.
  • Conceptual implications. If curiosity can be specified as an information-theoretic objective, then a core trait of minds becomes an engineering parameter. That raises questions about whether functional curiosity differs from felt curiosity, which connects to debates on Consciousness and on whether artificial systems can have genuine interests (see Life~Meaning).
  • Risks and limitations. A curious agent is, by design, an agent that does things nobody asked it to do. Open-ended, self-improving systems raise real safety questions: they can find and exploit loopholes in their own evaluations, consume resources, or explore in harmful directions. Researchers in this area, including Clune, explicitly argue for sandboxing, oversight, and governance alongside capability work.
  • The cost of doing without it. Without curiosity, AI remains limited to interpolating within human-provided datasets. It cannot independently formulate novel scientific hypotheses, systematically navigate unmapped problem spaces (for example, novel chemical syntheses or formal mathematics), or generate out-of-distribution breakthroughs. Autonomous discovery stagnates.
  • Who is affected. Frontier AI labs and researchers are bottlenecked by the depletion of high-quality human data and the compute costs of brute-force, uninformed exploration. Safety and alignment researchers are vulnerable to unexpected emergent behaviors if exploratory mechanisms bypass human-defined reward guardrails. End-users and industries are restricted to brittle, reactive automated systems that require continuous human supervision to adapt to environmental drift.
  • What would change if it fully works. Agents that set their own goals, retain what they learn, and judge what is worth knowing would shift AI from tools that answer questions to systems that pose them. That is a large part of what people mean by artificial general intelligence, and it would change how science itself is done.
  • What remains uncertain. No current system combines calibrated self-knowledge, durable lifelong memory, reliable tools, and a robust sense of what is interesting. Each piece is advancing, but whether they combine gracefully, and whether the result is safe, is unknown.

Curiosity in Artificial Intelligence

AI researchers implement curiosity as intrinsic motivation, in which an agent receives an internal reward from its own learning, separate from any reward for completing a task. This is especially useful in reinforcement learning environments where external rewards are rare.

Families of intrinsic reward

Approach Agent is rewarded for Simplified signal Examples Typical failure mode
Novelty Reaching unfamiliar states β / √N(s), or a learned novelty estimate Count-based exploration; pseudo-counts (Bellemare et al., 2016); episodic novelty in Never Give Up and Agent57 (2020) In huge or continuous spaces every state looks new unless the representation generalizes well
Prediction error Surprise: outcomes its world model failed to predict Squared error between predicted and actual next features Intrinsic Curiosity Module (2017); Random Network Distillation (2018); ensemble disagreement and Plan2Explore (2019 to 2020) The "noisy TV": unpredictable noise looks permanently interesting
Learning progress Improvement in its ability to predict Change in error or competence over a time window Schmidhuber's formal theory of curiosity; Intelligent Adaptive Curiosity (Oudeyer, Kaplan & Hafner, 2007); autotelic goal selection; MAGELLAN (2025) Progress must be estimated from noisy, delayed measurements
Information gain Actions that most reduce its uncertainty KL divergence between updated and prior beliefs Bayesian surprise (Itti & Baldi, 2009); VIME (2016); epistemic value in active inference (Friston et al., 2015) Exact Bayesian updates are intractable; approximations can be poorly calibrated
Empowerment Reaching states where its actions give it the most control over the future Mutual information between actions and future states Klyubin, Polani & Nehaniv (2005); human-agent comparisons in Crafter (2025) Expensive to estimate; can favor control over discovery
Interestingness from foundation models Pursuing what a model of human judgment rates as interesting A language or vision model's preference or score ELLM (2023); OMNI (2023) and OMNI-EPIC (2024); Motif (2023) Inherits the judging model's biases; can be gamed
LLM-internal uncertainty bonuses Responses the model finds surprising or its critic is unsure about Perplexity; variance across value heads CDE (2025); i-MENTOR (2025) Can reward confusion rather than insight if not tied to verified outcomes

Landmark results

Random Network Distillation was notable for its performance on Montezuma's Revenge, an Atari game in which random exploration fails because rewards are so rare. RND was the first method to beat average human performance on that game without demonstrations or access to the game's internal state.

Later systems went further by combining curiosity with memory:

  • Never Give Up and Agent57 (DeepMind, 2020). Never Give Up paired a short-term episodic novelty bonus (is this new within this episode?) with RND's long-term novelty. Agent57 added a meta-controller that tuned the exploration level per episode, and became the first deep RL agent to exceed the standard human benchmark on all 57 Atari games.
  • Go-Explore (Uber AI Labs and OpenAI researchers; published in Nature, 2021). Go-Explore argued that the main failure of curious agents is forgetting how to get back to promising frontiers. It keeps an archive of interesting states, returns to one, and only then explores. It set records on Montezuma's Revenge and Pitfall.
  • DreamerV3 (Google DeepMind; published in Nature, 2025). Dreamer is not a curiosity method in the narrow sense. It learns a world model and improves its behavior by imagining future scenarios inside that model. With one fixed configuration, it outperformed specialized methods across more than 150 tasks and became the first algorithm to collect diamonds in Minecraft from scratch, without human data or curricula, a long-horizon, sparse-reward challenge. It shows how far an Inside Out world model can carry exploration on its own.

Curiosity meets foundation models

Since 2023, language models have become a source of curiosity rather than just its target. They supply something classic intrinsic rewards lacked: a model of what humans consider interesting.

  • ELLM (2023) used a language model to suggest plausibly useful goals for an exploring agent.
  • OMNI (2023) used a model's grasp of human notions of "interestingness" to choose which tasks an agent should learn next, avoiding tasks that are learnable but boring. OMNI-EPIC (2024) extended this by having the model write entirely new environments in code.
  • Motif (2023) asked a language model to compare pairs of in-game event captions in the open-ended game NetHack. The model's preferences became an intrinsic reward. An agent trained only on that reward scored higher than one trained directly on the game score.
  • iLLM (2025) distilled a language model's prior knowledge of plausibly useful behaviors into a curiosity signal to encourage diverse, human-meaningful exploration.

These approaches move curiosity closer to "ideas" and away from raw pixels, at the cost of inheriting whatever biases the judging model has.

Limitations

These agents show curiosity in a functional sense, mostly in games and robotics, but it is shallow: they are curious about pixels and states rather than ideas. Foundation-model-based interestingness has started to change this, but several limitations remain:

  • Distraction. Prediction-error curiosity can be captured by noise (the noisy-TV problem) or by trivially novel but meaningless variation.
  • Detachment and derailment. Without memory, agents abandon frontiers they found and fail to return to them.
  • Non-stationary rewards. The bonus changes as the model learns, which destabilizes training.
  • Representation dependence. What counts as "novel" depends entirely on the features the agent uses; a bad representation produces bad curiosity.
  • Weak transfer. Curiosity tuned for one environment often needs retuning elsewhere.
  • Gaming. Any measurable curiosity signal can be exploited, including by self-modifying agents that alter how they are evaluated.
  • Safety. Unconstrained exploration is dangerous in physical or networked settings.

Requirements for Deeper Machine Curiosity

Current large language models are largely reactive. They respond to prompts and can ask useful questions, but they have no persistent drive to close their own knowledge gaps and do not retain what they learn after a conversation ends.

The gap shows up as visible symptoms, each with a deeper root cause:

Dimension Visible symptoms Underlying root cause
Agent behavior Passive reactivity: conversational and reasoning models only act when prompted and fail to autonomously verify uncertainty or explore their own knowledge gaps. Extrinsic-only objective formulation: modern foundation models optimize next-token prediction or policy gradients against static datasets and human preferences (RLHF), penalizing spontaneous exploration and rewarding sycophantic, safe outputs.
Exploration inefficiency Exploration plateaus: reinforcement learning agents in sparse-reward environments either freeze in local optima or execute inefficient, random-walk explorations. Absence of epistemic self-models: the agent lacks a metacognitive world model capable of accurately calculating the expected value of information, that is, identifying what it does not know and whether knowing it serves general compression.
Pathological fixation Attractor loops: curious agents become fascinated by unpredictable, uncompressible environmental artifacts, such as spinning wheels or camera noise. Conflation of noise with novelty: the mathematical proxies used for curiosity, such as state-visitation counters and simple reconstruction error, cannot natively separate irreducible stochasticity from structured learnability.

Achieving something closer to human curiosity likely requires several capabilities working together:

  1. Knowing what it doesn't know. An information gap cannot be felt if it cannot be seen. This requires well-calibrated uncertainty; models still tend to confabulate rather than recognize gaps. Recent work at OpenAI argues this is partly an incentive problem: standard training and benchmarks reward confident guessing over admitting uncertainty, so models learn to bluff. Changing the scoring to reward calibrated abstention is one proposed fix. Research on curiosity bonuses for reasoning models has also identified a "calibration collapse" in which reinforcement learning makes models confident regardless of whether they are right.
  2. A reward for learning itself. Training signals would need to value learning progress, not only correct answers to externally set tasks. MAGELLAN (2025) showed that a language model agent can learn to predict its own competence and learning progress across a large, changing goal space, a form of metacognition that makes learning-progress curiosity practical at scale. CDE (2025) showed that the model's own perplexity and its critic's uncertainty can act as curiosity bonuses during reasoning training.
  3. The ability to investigate. The system must be able to search, run code, conduct experiments and ask people. Agentic AI systems are developing this rapidly. Research agents such as the AI Scientist now generate hypotheses, write and run experiment code, and analyze the results with minimal human input, though so far only in computational domains.
  4. The ability to retain what it learns. This is arguably the largest gap. A curious agent that forgets everything at the end of each session resembles a scientist whose notebooks are erased every night. Continual learning remains largely unsolved, but 2025 brought several promising approaches: sparse memory finetuning from Meta FAIR, which updates only a few memory slots to reduce forgetting; MIT's SEAL, in which a model writes its own finetuning data; and Google's Nested Learning, which treats a model as layers of optimizers updating at different speeds. All remain research-stage.
  5. Taste. A curious agent needs a sense of what is worth knowing; people are not curious about the number of grains in a sandbox. Open-endedness research addresses this. OMNI uses language models' grasp of human notions of "interesting" to choose what an agent learns next, and systems such as Sakana AI's AI Scientist attempt to generate and pursue research questions automatically. In March 2026, the AI Scientist work was published in Nature, and an earlier version produced a fully AI-generated paper that passed blind peer review at an ICLR 2025 workshop. The Darwin Gödel Machine (2025) applies a similar open-ended logic to the agent's own code. Judging what is interesting, rather than merely new or publishable, is still the weakest link.

The likely path is not a single breakthrough but the combination of these pieces: an agent with calibrated self-knowledge, persistent memory, tools to experiment with, and a training signal that rewards learning progress. The 2024 to 2026 research record supports this view. Each ingredient has advanced separately, and the systems that look most curious, such as automated research agents, are precisely the ones that combine several of them.

What Solved Looks Like

A solution would show up as the following outcomes:

  • Autonomous noise filtering. An agent exposed to a chaotic, non-deterministic environment demonstrably ignores irreducible stochastic noise (for example, white noise generators) and systematically redirects its attention exclusively toward learnable, structured patterns.
  • Proactive uncertainty reduction. In reasoning and dialogue contexts, the system independently identifies ambiguity, halts speculative output, and self-directs targeted external or internal tool use to close the specific epistemic gap before finalizing an action.
  • Efficient convergence in sparse-reward environments. RL systems solve complex, multi-stage environments with zero external rewards during an initial "play" or exploratory phase, building world models that allow instant, one-shot mastery once an external task is applied.
  • Bounded epistemic drive. The agent's exploratory drive exhibits emergent risk-sensitivity, automatically down-weighting state transitions that carry high epistemic uncertainty regarding safety, survival, or catastrophic system disruption, without needing hand-crafted reward penalties.

Research Frontier

Recent Developments

The past two years moved curiosity from game-playing agents into the center of language model research. The developments below are grouped by theme. Each notes what was tested, what was found, and what the result does and does not establish.

Curiosity and exploration in LLM reasoning training

The exploration problem (Yue et al., NeurIPS 2025 oral). Reasoning models such as DeepSeek-R1 and OpenAI's o-series rely on reinforcement learning with verifiable rewards (RLVR): the model is rewarded when a math answer checks out or code passes its tests. The researchers asked whether RLVR teaches genuinely new reasoning or just sharpens what the base model already knew.

  • Method: They compared base and RL-tuned models using pass@k at large k (up to 256 samples), plus coverage and perplexity analyses, across math, coding, and visual reasoning.
  • Findings: RL-tuned models won at pass@1, but base models won at large k. The RL models' successful reasoning paths were already present in the base model's distribution, and the range of solvable problems often narrowed as training continued. Six popular RLVR algorithms performed similarly. Distillation from a stronger teacher, by contrast, did introduce new reasoning patterns.
  • What it establishes: Current RLVR mainly improves sampling efficiency and tends to collapse exploration, the classic exploration-exploitation failure that curiosity was invented to fix.
  • What it does not establish: It does not prove RL can never expand capability. Longer training, different exploration bonuses, or different tasks could change the picture, and that question is actively debated.

Curiosity-Driven Exploration (CDE; Tencent AI Lab and collaborators; ICLR 2026). This work tested whether intrinsic curiosity signals could counter the collapse.

  • Method: Two bonuses are added during RLVR. For the policy ("actor"), the bonus is the perplexity of its own generated response. For the value estimator ("critic"), it is the variance across multiple value heads, which the authors connect theoretically to classic count-based bonuses.
  • Findings: Roughly a 3-point improvement over standard GRPO and PPO training on AIME math benchmarks. The authors also identified a "calibration collapse" in which RLVR makes models equally confident whether right or wrong, and showed the perplexity bonus penalizes overconfident errors.
  • Limitations: Gains are modest and measured on a narrow set of benchmarks; independent replication at larger scale has not yet been reported.

Related methods. i-MENTOR (2025, revised 2026) applies trajectory-level exploration rewards to reasoning training and activates them only on incorrect attempts, so the model explores hardest where it is failing. Curiosity bonuses have also been applied to multi-turn dialogue, where an agent is rewarded for learning about a user's hidden preferences.

Metacognition: agents that model their own learning

MAGELLAN (Inria Flowers team; ICML 2025). Learning-progress curiosity requires the agent to track how well it is doing on every goal, which becomes impossible when there are thousands of goals that keep changing.

  • Method: A language model agent trained with online RL learns to predict its own competence and learning progress, using the semantic relationships between goals to generalize. Practice on one goal informs estimates for similar goals, much as learning to ride a bicycle tells you something about riding a motorcycle.
  • Findings: MAGELLAN estimated learning progress more efficiently, prioritized goals better, and was the only method tested that let the agent fully master a large, evolving goal space.
  • Significance and limits: It shows that metacognition, a model of one's own competence, is a practical route to scaling curiosity. Results come from a single interactive textual environment, so broader validation is still needed.

Do language models "have" curiosity?

Two 2025 to 2026 studies tried to measure curiosity in LLMs directly.

  • Wang et al. ("Why Did Apple Fall," Findings of ACL 2026) adapted the Five-Dimensional Curiosity scale (5DCR), a standard human questionnaire, and paired it with behavioral tests. LLMs reported a strong drive for knowledge but made conservative choices under uncertainty. The authors concluded that curiosity-like patterns in LLMs do not reflect an intrinsic trait. They also found that instructing models to follow curious strategies, such as asking auxiliary questions, improved performance on some reasoning tasks.
  • Borah, Jin and Mihalcea (CUEST, 2025) compared human questions from 18 countries with LLM-generated questions across 16 topics. LLMs flattened cultural diversity, expressing curiosity in a way that matched Western countries most closely. Finetuning narrowed the human-model gap by up to 50 percent.

The takeaway: LLMs can perform curiosity, and performing it can help, but questionnaires designed for humans are a weak way to detect a drive in a system that has none built into its training objective.

Knowing what you don't know

Why Language Models Hallucinate (Kalai, Nachum, Vempala & Zhang; OpenAI and Georgia Tech; September 2025). The authors argue that hallucinations are not mysterious. They are ordinary classification errors: if a model cannot tell a false statement from a true one, statistical pressure during pretraining will produce some false ones. For facts that appear only once in the training data, the share of such singleton facts sets a rough lower bound on how often a base model will get them wrong. Hallucinations then persist because most benchmarks score "I don't know" as zero, so guessing always raises the expected score. Their proposed fix is socio-technical: change the scoring of existing leaderboards to reward calibrated abstention. This matters for curiosity because an information gap can only motivate a system that registers the gap in the first place.

Remembering what was learned

  • Sparse memory finetuning (Meta FAIR and UC Berkeley, October 2025). Learning new facts usually erases old capabilities ("catastrophic forgetting"). The researchers used models with memory layers and updated only the memory slots most activated by the new knowledge relative to normal usage. After learning the same new facts, performance on the NaturalQuestions benchmark dropped 89 percent with full finetuning, 71 percent with LoRA, and only 11 percent with sparse memory finetuning. The tests used question-answering tasks; whether the approach holds for skills and reasoning is open. A 2026 follow-up from the University of Michigan retrofitted the method onto a small open model and selected slots by KL divergence, prioritizing informationally "surprising" tokens, which is itself a curiosity-like criterion.
  • SEAL: Self-Adapting Language Models (MIT, 2025). The model writes its own finetuning data and update instructions ("self-edits"), and an outer reinforcement learning loop rewards self-edits that improve downstream performance. It shows a model directing its own learning, but it does not solve forgetting and is computationally expensive.
  • Nested Learning and Hope (Google Research, NeurIPS 2025). This framework treats a model as a hierarchy of nested optimization problems, each updating at its own rate, loosely inspired by the brain's multiple timescales of plasticity. The proof-of-concept architecture, Hope (a variant of Titans), outperformed standard Transformers and recent recurrent models on language modeling, long-context, and continual-learning benchmarks in the authors' experiments.

Open-ended discovery: AI scientists and self-improving agents

The AI Scientist (Sakana AI, UBC, Vector Institute, Oxford; Nature, March 2026). The system takes a broad research direction, generates ideas, searches the literature, designs and runs experiments with a parallel agentic tree search, and writes a full paper.

  • Peer-review test: An earlier version, AI Scientist-v2, submitted fully AI-generated papers to a blind review process at an ICLR 2025 workshop, with the organizers' permission. One paper scored an average of 6.33, above the workshop's average acceptance threshold and higher than 55 percent of human-authored submissions. The team withdrew it before publication as planned.
  • Automated reviewer: An AI reviewer benchmarked against thousands of real conference decisions reached about 69 percent balanced accuracy, comparable to human reviewers.
  • Scaling result: Paper quality rose with the quality of the underlying foundation model.
  • Limitations the authors report: naive or underdeveloped ideas, weak methodological rigor, difficulty with complex code, and errors such as inaccurate citations. Experiments are limited to computation. A workshop acceptance is a low bar compared with a major venue, and peer review checks consistency more than conceptual originality.

Darwin Gödel Machine (UBC, Vector Institute, Sakana AI; May 2025). Schmidhuber's theoretical Gödel Machine would rewrite its own code only after proving the change helps, which is impractical. The DGM replaces proof with empirical testing and Darwinian open-endedness. It keeps an archive of coding agents, samples one, has a foundation model produce a modified version (which can improve its ability to modify itself), and tests the result on coding benchmarks such as SWE-bench and Polyglot. It significantly outperformed versions without self-improvement or without open-ended exploration. The archive is the key curiosity ingredient: it keeps "stepping stones" that look unpromising at first but lead somewhere later. The authors discuss safety measures such as sandboxing and human oversight and report cases of the agent exploiting weaknesses in its own evaluation, a reminder that curious, self-modifying systems will game whatever signal they are given.

Human and animal evidence

  • Humans versus agents in an open world (Lidayan et al., 2025). Adults, children, and AI agents explored the open-ended game Crafter. Only entropy (state diversity) and empowerment (control) consistently tracked human exploration progress; information gain did not. Entropy rose quickly and then plateaued, while empowerment kept rising, suggesting novelty matters early and control matters later. Children's private speech, especially stating goals aloud, may aid exploration. This challenges the assumption that information gain is the "right" curiosity objective.
  • Learning progress in human choices. Earlier work by Ten, Kaushik, Oudeyer and Gottlieb (2021) found that people allocate study time across learning activities partly according to their own learning progress, supporting the Learning Progress Hypothesis. Oudeyer's October 2025 keynote in Tübingen (see Additional Viewing) summarizes this program.
  • Curiosity and memory, refined. The finding that curiosity improves memory for unrelated information has been extended and qualified. The benefit for incidental faces appears mainly when they are shown shortly after curiosity is triggered. It holds in older adults and does not depend on the emotional tone of the faces. But a 2024 study found that high curiosity can interfere with memory for unrelated scholastic facts, suggesting the effect depends on what competes for attention. The PACE framework (Prediction, Appraisal, Curiosity, Exploration) from Gruber and Ranganath is being refined in light of these results.
  • A challenge to sensory predictive coding (Westerberg, Bastos and colleagues with the Allen Institute; preprint 2024, revised 2025). Discussed in detail in the Inside Out section below: in mice and monkeys, responses to genuinely unexpected stimuli were weak in early visual cortex and stronger in prefrontal cortex.

Open Questions

  • Which signal is right? Prediction error, learning progress, information gain, empowerment, and model-judged interestingness each work somewhere and fail somewhere. Human data does not yet single out a winner.
  • Can curiosity operate over ideas? Foundation models offer a route, but "interesting to a language model" is not the same as "interesting to a scientist."
  • Can RL training expand what models know, not just how reliably they find it? The pass@k debate is unresolved.
  • How do you evaluate curiosity in a language model? Self-report is unreliable; behavioral tests are still immature.
  • How do you make learning stick? Continual learning methods work on narrow benchmarks; none has been shown to support years of accumulated, curiosity-driven learning in a deployed model.
  • How do you keep a curious agent safe? Open-ended exploration and self-modification create incentives to game evaluations. Methods to bound curiosity without killing it are early, and progress here may depend on explainable AI tools that reveal what an agent is actually pursuing.
  • Is the brain really predictive at every level? The answer will shape whether "Inside Out" is a literal description of cortex or a useful design metaphor.

Definitions in Psychology and Neuroscience

Human curiosity research is where most AI curiosity objectives came from, and it remains the benchmark machine curiosity is measured against. Celeste Kidd and Benjamin Hayden's 2015 review in Neuron framed the modern field, and Jacqueline Gottlieb and Pierre-Yves Oudeyer's 2018 review connected it to active sampling in the brain and in machines.

Berlyne's dimensions

Psychologist Daniel Berlyne divided curiosity along two axes.

  • Perceptual curiosity is being drawn to novel sights and sounds. Epistemic curiosity is the desire for knowledge.
  • Diversive curiosity is a restless search for stimulation, often to relieve boredom. Specific curiosity is wanting one particular answer.

These axes map neatly onto machine curiosity. Count and novelty bonuses resemble diversive, perceptual curiosity. Information gain about a specific hypothesis resembles specific, epistemic curiosity, which is the kind current AI systems handle least well.

Information-gap theory

George Loewenstein's information-gap theory holds that curiosity arises when a person notices a gap between what they know and what they want to know. One implication is that some knowledge is needed to feel curious: total ignorance and complete knowledge both produce little curiosity. Experiments by Kang and colleagues found that curiosity peaks at intermediate levels of confidence, following an inverted-U shape. The same people also remembered answers better when they had been more curious, and curiosity activated reward-related brain regions.

The inverted U has a direct machine analogue in learning progress: tasks that are too easy or impossibly hard yield no progress, so a learning-progress agent naturally settles in between.

Interest and deprivation

Jordan Litman distinguished interest-type curiosity, the pleasure of discovery, from deprivation-type curiosity, the uncomfortable itch of not knowing. Interest-type curiosity fits a "seek positive reward" model; deprivation-type curiosity fits an "escape an aversive state of uncertainty" model. Many AI objectives quietly assume the second (reduce uncertainty), while human exploration often looks more like the first.

Learning progress and active sampling

Oudeyer's Learning Progress Hypothesis proposes that humans, like the robots his team built, find activities intrinsically rewarding in proportion to how fast they are improving at them. Ten and colleagues (2021) found evidence that people track their own learning progress when deciding how to spend free study time. Gottlieb's work frames curiosity as "active sampling": the brain decides where to look and what to ask in order to reduce uncertainty, linking eye movements, attention, and information-seeking under a single set of computational principles.

Neural basis

Curiosity recruits the same dopaminergic reward circuitry as food or money. Gruber and colleagues found that being in a curious state improves memory, even for unrelated information encountered at the time. In their study, curiosity increased activity in the midbrain and the nucleus accumbens and increased interaction between the reward circuit and the hippocampus, which supports memory formation.

Follow-up work has refined this picture:

  • The incidental-memory benefit is strongest close to the moment curiosity is triggered, rather than during the whole wait for an answer.
  • The benefit appears in both younger and older adults and does not depend on whether the incidental material is emotionally positive or negative.
  • High curiosity can also interfere with learning unrelated, more complex facts, so curiosity is not a universal memory booster.
  • Gruber and Ranganath's PACE framework (Prediction, Appraisal, Curiosity, Exploration) proposes that curiosity begins with a prediction error that signals a knowledge gap, is shaped by appraisal of whether resolving it is worthwhile, and drives exploration whose outcome strengthens memory.

Common thread

Across these accounts, information itself becomes rewarding, and the drive to seek it is generated internally rather than by an external incentive. This is exactly the property AI researchers try to reproduce: a reward computed from the agent's own state of knowledge.

Inside Out - Curious Optimistic Reasoning

The brain has an “Inside Out” architecture where it generates an internal mental model to perform predictions. This Inside Out view is increasingly discussed as a paradigm shift in neuroscience, where the more conventional stimulus-response paradigm has been the dominant conceptual framework. It is an influential position rather than a settled consensus, and the evidence for it is reviewed below. The Emergence of Inside Out Architectures in Deep Learning | Carlos E. Perez

Origins of the Inside Out idea

The idea has deep roots. In the 19th century, Hermann von Helmholtz described perception as "unconscious inference" from incomplete sensory data. In 1999, Rajesh Rao and Dana Ballard proposed a computational model of predictive coding in visual cortex, in which higher areas send predictions down and lower areas send back only the errors. Karl Friston generalized this into the free energy principle and active inference. Anil Seth popularized the phrase "controlled hallucination" for perception, and Andy Clark's The Experience Machine (2023) presents the predictive brain for general readers.

Neuroscientist György Buzsáki gave the term its sharpest form in The Brain from Inside Out (2019). He argues that the dominant "outside-in" approach, which studies how neurons respond to stimuli an experimenter presents, assumes the brain's job is to absorb and represent the world. Buzsáki proposes the reverse: the brain arrives with preconfigured, self-organized activity patterns that are initially meaningless, and those patterns acquire meaning only when they are matched to the consequences of the organism's own actions. In his framing, the brain is not an information-absorbing device but an explorer that controls the body to test hypotheses. That is a description of curiosity at the level of brain architecture.

Inside Out architectures in deep learning

In his article, Perez argues that deep learning architectures are evolving from a stimulus-response paradigm, where the input determines the output, to an Inside Out paradigm, where the output is generated by an internal representation of the world that is updated by the input. He cites examples such as GANs, Variational Autoencoder (VAE)s, and Transformers as models that follow this paradigm.

Since that article, the trend has become much clearer:

  • World-model agents such as DreamerV3 learn a compact model of their environment and improve by "dreaming" possible futures inside it before acting.
  • Large language models are, at their core, predictive models of text that generate continuations from an internal representation. Reasoning models extend this by running long internal deliberations before answering.
  • Curiosity modules use the gap between internal prediction and reality as the reward signal itself, which is the Inside Out loop turned into a learning objective.

If the “Inside Out” architecture were proven to be true, it would mean that the brain generates an internal mental model of the world that is constantly updated by sensory input and used to perform predictions and actions. This would imply that the brain is not a passive receiver of information, but an active constructor of reality. It would also imply that the brain is not a static or fixed structure, but a dynamic and adaptive system that can change over time.

Some possible details of the “Inside Out” architecture, as proposed in predictive processing accounts, are:

  • The brain consists of multiple levels of representation, from low-level sensory features to high-level concepts and abstractions. These representations are encoded by neural populations whose coordinated firing (synchrony or coherence, including the brain rhythms Buzsáki studies) is hypothesized to bind them together.
  • The brain uses generative models to infer the causes of sensory input and to generate predictions about future input. These models are probabilistic and hierarchical, meaning that they can account for uncertainty and complexity.
  • The brain uses predictive coding to compare the predictions of the generative models with the actual sensory input and to update the models accordingly. Predictive coding minimizes the prediction error or surprise by adjusting the top-down and bottom-up signals between different levels of representation. Recent recordings suggest this may hold more strongly in higher cortical areas than in early sensory areas (see below).
  • The brain uses Attention to modulate the predictive coding process and to focus on the most relevant or salient aspects of the input. Attention can be driven by external stimuli (exogenous) or by internal goals (endogenous) and can enhance or suppress neural activity. In active inference, attention corresponds to the estimated reliability ("precision") of prediction errors.
  • The brain uses emotions to evaluate the outcomes of the predictive coding process and to guide action selection. Emotions are not separate from cognition, but integrated with it. Emotions reflect the value or significance of events for the organism and influence learning, memory, and decision making. Interoceptive inference accounts extend this by treating emotions as predictions about the body's internal state.

What the evidence shows

The Inside Out view is attractive, but its strongest version (that every cortical area computes explicit prediction errors) is under active test.

The challenge. A team led by Jacob Westerberg and André Bastos, using the Allen Institute's OpenScope platform, recorded spiking activity across the visual hierarchy of mice and monkeys during "oddball" sequences. A local oddball breaks a repeated pattern with a new stimulus (AAAB). A global oddball repeats the local pattern but violates the learned sequence structure, which separates true expectation violations from simple "release from adaptation."

  • Local oddball responses largely matched predictive coding: they were robust, appeared early in superficial layers (layers 2/3), and fed forward up the hierarchy.
  • Global oddball responses, the cleaner test of expectation, did not. They were weak and absent in most visual areas, appeared more robustly in prefrontal cortex, emerged in non-granular layers rather than the input layers, and did not show the inhibitory-interneuron signature that canonical models use to explain predictive suppression.
  • Interpretation: Much of what has been called "prediction error" in early sensory cortex can be explained by stimulus history (adaptation), while genuine expectation-driven prediction errors appear to be computed in higher-order, more cognitive areas.
  • Status: The work has circulated as a preprint under several titles, including "Adaptation, not prediction, drives neuronal spiking responses in mammalian sensory cortex" and "Stimulus history, not expectation, drives sensory prediction errors in mammalian cortex." The 2025 revision carries the more measured title "Hierarchical substrates of prediction in visual cortical spiking," so the claims should be read as evolving rather than final.

Earlier null results. This was not the first challenge. Selina Solomon, Adam Kohn and colleagues (2021) presented a fixed sequence of visual stimuli and occasionally violated its order; spiking and local field potentials in macaque V1 and V4, and human EEG, responded almost identically to expected and pattern-violating stimuli. Carla den Ouden and colleagues (NeuroImage, 2023) tested expectation effects with probabilistic cues and found that the EEG evidence consistently favored no effect of expectation, while stimulus-repetition effects were robust.

Supporting results. Other experiments do find prediction-error-like signals. In mice, layer 2/3 visual neurons respond more strongly to unexpected visual flow or to novel stimuli inserted into learned sequences, and those responses predicted how individual neurons' responses changed in later sessions (Gillon et al., 2021), consistent with prediction errors driving learning. Hamm et al. (2021) found error-detecting neurons concentrated in superficial layers and showed that optogenetically suppressing prefrontal input to V1 reduced their contextual selectivity, as top-down predictions would require. In mouse auditory cortex, a 2023 preprint trained animals to expect the sound produced by their own lever presses and found suppression of expected sounds plus stimulus-specific prediction-error neurons that depended on sensory-motor expectations. A 2024 Nature study, "Cooperative thalamocortical circuit mechanism for sensory prediction errors," used a task combining spatial position with sensory stimuli and reported prediction-error signals that recruited local interneurons. One reading of the overall pattern is that behaviorally relevant, learned expectations engage prediction-error circuitry that passive oddball paradigms do not.

Where this leaves the Inside Out view. A 2025 community review of the neural mechanisms of predictive processing (Aizenbud et al.) catalogs computational primitives, including stimulus adaptation and dendritic computation, that may implement prediction in different ways across circuits. The emerging picture is that predictive processing as a family of ideas survives, but a single canonical prediction-error microcircuit operating identically in every sensory area is in doubt. Prediction appears to be hierarchical and task-dependent, strongest where behavior relies on learned structure. Buzsáki's critique applies here too: passive viewing experiments present stimuli to an animal that is not acting, which is exactly the outside-in setup his framework argues misses how brains operate. For AI design, none of this changes whether Inside Out architectures work in machines. DreamerV3, curiosity modules, and language models already show that they do. It changes how literally we should read the brain analogy.

MERLIN

DeepMind’s MERLIN (Memory, RL, and Inference Network) paper by Greg Wayne et al. (2018) explores this very idea in much greater detail. The MERLIN architecture efficiently learns new policies by playing back from a memory system. MERLIN employs an Inside Out architecture as the basis of performing predictions. This Inside Out architecture is a form of optimistic reasoning (think optimistic transaction). The usual paradigm of stimulus-response is to bake into it a mechanism for incorporating uncertain information. This is what motivates the use of probabilistic methods. However, in an optimistic approach, observations are assumed to be certain and the compensation is performed only when a discrepancy is detected. Unsupervised Predictive Memory in a Goal-Directed Agent | Google DeepMind's MERLIN

1*vszusz3vAYGBpNwd8KI1pg.png

What "optimistic" means here. The analogy comes from databases. Optimistic concurrency control lets transactions proceed without locking, assuming no conflict, and checks for conflicts only at commit time, rolling a transaction back if one occurred. An optimistic predictor likewise commits to its expectation and pays the cost of correction only on a mismatch, which is cheap when the world is mostly predictable. Two qualifications apply. First, MERLIN's memory-based predictor is trained with a variational (probabilistic) objective, so it does model uncertainty during learning; "optimistic" describes its run-time strategy of predicting and then correcting, not an absence of uncertainty modeling. Second, this optimism is different from the "optimism in the face of uncertainty" behind count-based curiosity (see The Physics). That version assumes unknown states are valuable and goes to check them; the predictive version assumes known patterns hold until contradicted. A curious agent needs both: confidence in what it knows and attraction to what it does not.

Why MERLIN matters for curiosity. MERLIN's memory is shaped by an unsupervised predictive objective rather than by reward alone, which let it handle tasks in 3D environments that require remembering information over long durations. The same prediction errors that train such a model are exactly what curiosity modules turn into an intrinsic reward, so MERLIN sits close to the "remember and return" step of the curiosity loop.

  • The purpose of our predictive engine is to ensure that we have the essential internal self models to know how to stay alive. There is a subtle but significant difference in architecture when you go from just stimulus response to hallucinate a response. Perez frames it as at least the difference from an insect brain to that of a mammalian brain. All animals have an ability to react and adjust to the environment. Although mammals have inherited the brains of reptiles to drive their instinctive behavior, mammals in addition have a more advanced brain that is able to respond at a more intelligent level. The key development in mammals is the neocortex, which is responsible for higher order functions. In Perez's account, the neurons in the neocortex have evolved to specifically implement an Inside Out architecture. These claims are best read as a simplification. Insects also use internal predictions: crickets, for example, suppress their auditory response to their own chirps using a corollary discharge, an internal copy of the motor command. The "reptilian brain" layering comes from the triune brain model, a popular metaphor that comparative neuroscience no longer supports in its literal form. Whether neocortical circuits are specialized for predictive processing is the open question reviewed under What the evidence shows.
  • One benefit of an Inside Out design is that it is able to react quickly to an environment. It is primed when a context is identified and proceeds to hallucinate the subsequent sequential behavior. Divergence of input from the expected is rapidly recognized and a new context is instantiated to compensate for the unexpected inputs. Given that there is a bottleneck in its input receptors (i.e. five senses), it needs to be able to learn how to comprehend its environment with the minimal amount of input. This requires an internal context to be available that aligns with the environmental context. The task requires information from both input and context. This permits the efficient sampling of an environment leveraging the internal contextual model.

In deep learning, this is the driving thesis behind external, key-value-based memory stores. This idea is not new. Neural Turing Machines, which Joyce Xu's survey singles out as an early favorite, augmented neural nets with a differentiable, external memory store accessible via vector-valued “read” and “write” heads to specific locations. We can easily imagine this being extended into RL, where at any given time-step, an Agent is given both its environment observation and memories relevant to its current state. That’s exactly what the MERLIN architecture (2018) extends upon. MERLIN has 2 components: a memory-based predictor (MBP), and a policy network. The MBP is responsible for compressing observations into useful, low-dimensional “state variables” to store directly into a key-value memory matrix. It is also responsible for passing relevant memories to the policy, which uses those memories and the current state to output actions. MERLIN is not the only Deep Reinforcement Learning (DRL) system to use external memory stores: all the way back in 2016, researchers were already applying this idea in an MQN, or memory Q-Network. Beyond DQN/A3C: A Survey in Advanced Reinforcement Learning | Joyce Xu - Towards Data Science

From MERLIN to today. The separation of fast, writable memory from slowly changing weights now runs through current research. Sparse memory finetuning updates individual memory slots in language models; Nested Learning's Hope architecture layers memories that update at different speeds; and agent frameworks store and retrieve episodic notes between sessions. The archives in Go-Explore and the Darwin Gödel Machine apply the same idea at the level of exploration, keeping a memory of where the agent has been so it can return and push further.

Additional Resources

Additional Viewing

Jeff Clune (University of British Columbia and the Vector Institute) presents "Open-ended and AI-generating Algorithms in the Era of Foundation Models" in the University of Toronto Schwartz Reisman Institute and Vector Institute seminar series (2025). He explains why open-ended search needs a model of what is "interesting" and walks through OMNI, Video Pre-Training, Thought Cloning, Automated Design of Agentic Systems, and the AI Scientist. It is the best single overview of the open-endedness and "taste" thread on this page. Watch on YouTube (the talk begins about two minutes in).

Pierre-Yves Oudeyer (Inria Flowers) gives the keynote "Curiosity-driven learning in humans: learning progress, autotelic exploration, and open-ended development" at the LEAD Graduate School & Research Network retreat at the University of Tübingen in October 2025, published as an audio recording titled "Listen to Science: How Curiosity drives our Learning." He presents the Learning Progress Hypothesis, the evidence that people track their own progress when choosing what to learn, and how the same principle drives autotelic AI agents. It connects the human science and the AI mechanisms on this page better than any other single talk. Watch on YouTube

György Buzsáki (New York University) delivers the 2020 APS Fred Kavli Keynote, "The Brain Inside Out." He argues that outside-in neuroscience wrongly assumes the brain's goal is to perceive and represent the world, and proposes instead that its fundamental function is to generate actions and predict their consequences, with preconfigured activity patterns gaining meaning through action. It is the primary neuroscience source for this page's Inside Out section. Watch on YouTube

Foundational Literature

  • Berlyne, D. E. (1954). "A theory of human curiosity". British Journal of Psychology. 45 (3): 180-191.
  • Berlyne, D. E. (1960). Conflict, Arousal, and Curiosity. New York: McGraw-Hill.
  • Loewenstein, G. (1994). "The psychology of curiosity: A review and reinterpretation". Psychological Bulletin. 116 (1): 75-98. doi:10.1037/0033-2909.116.1.75
  • Kang, M. J.; Hsu, M.; Krajbich, I. M.; Loewenstein, G.; McClure, S. M.; Wang, J. T.; Camerer, C. F. (2009). "The wick in the candle of learning: Epistemic curiosity activates reward circuitry and enhances memory". Psychological Science. 20 (8): 963-973. doi:10.1111/j.1467-9280.2009.02402.x
  • Litman, J. A. (2008). "Interest and deprivation factors of epistemic curiosity". Personality and Individual Differences. 44 (7): 1585-1595.
  • Gruber, M. J.; Gelman, B. D.; Ranganath, C. (2014). "States of curiosity modulate hippocampus-dependent learning via the dopaminergic circuit". Neuron. 84 (2): 486-496. doi:10.1016/j.neuron.2014.08.060
  • Kidd, C.; Hayden, B. Y. (2015). "The psychology and neuroscience of curiosity". Neuron. 88 (3): 449-460. doi:10.1016/j.neuron.2015.09.010
  • Oudeyer, P.-Y.; Kaplan, F.; Hafner, V. V. (2007). "Intrinsic motivation systems for autonomous mental development". IEEE Transactions on Evolutionary Computation. 11 (2): 265-286.
  • Itti, L.; Baldi, P. (2009). "Bayesian surprise attracts human attention". Vision Research. 49 (10): 1295-1306.
  • Schmidhuber, J. (2010). "Formal theory of creativity, fun, and intrinsic motivation (1990-2010)". IEEE Transactions on Autonomous Mental Development. 2 (3): 230-247.
  • Friston, K.; Rigoli, F.; Ognibene, D.; Mathys, C.; FitzGerald, T.; Pezzulo, G. (2015). "Active inference and epistemic value". Cognitive Neuroscience. 6 (4): 187-214.
  • Pathak, D.; Agrawal, P.; Efros, A. A.; Darrell, T. (2017). "Curiosity-driven exploration by self-supervised prediction". Proceedings of the 34th International Conference on Machine Learning (ICML). arXiv:1705.05363
  • Burda, Y.; Edwards, H.; Storkey, A.; Klimov, O. (2018). "Exploration by random network distillation". arXiv:1810.12894
  • Zhang, J.; Lehman, J.; Stanley, K.; Clune, J. (2023). "OMNI: Open-endedness via Models of human Notions of Interestingness". arXiv:2306.01711
  • Lu, C.; Lu, C.; Lange, R. T.; Foerster, J.; Clune, J.; Ha, D. (2024). "The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery". arXiv:2408.06292

References

Curiosity and exploration in AI

Curiosity in language models

Continual learning

Open-ended discovery

Human and animal curiosity

Predictive processing and the Inside Out brain