Partial Observability

From PRIMO.ai
Jump to navigation Jump to search

YouTube ... Quora ...Google search ...Google News ...Bing News


In the textbook Markov Decision Process (MDP), an agent is handed the complete state of its environment at every step: the full chessboard, the exact position and velocity of the robot, everything that matters for predicting what happens next. Real problems almost never look like that. A robot's camera cannot see behind the box it is about to pick up. A self-driving car cannot see the pedestrian hidden by a parked van, or read the intentions of the driver beside it. A poker bot cannot see its opponents' cards. A customer-service assistant does not know what the customer actually wants until it asks. A coding agent cannot read an entire repository at once, and a web agent sees one page at a time.


Partial observability in one sentence: the agent never sees the true state of the world, only clues about it, so it must remember and infer before it can decide -- and sometimes it must act just to find out more.


These are all problems of partial observability: the agent receives observations that carry only incomplete, noisy, or ambiguous information about the underlying state. Partial observability has three immediate consequences:

  1. The Markov property breaks for observations. The current observation alone no longer tells the agent everything it needs. Two situations that look identical (perceptual aliasing) may require different actions, so a policy that maps the latest observation straight to an action can be arbitrarily bad.
  2. The agent needs an internal state. To act well it must summarize its history of actions and observations, either explicitly as a probability distribution over possible states (a belief state) or implicitly in a learned memory such as a recurrent network, a Transformer context window, a state space model, or a written note in an LLM agent's scratchpad.
  3. Information becomes something worth acting for. Because uncertainty can be reduced by looking, listening, probing, or asking, an optimal agent sometimes chooses actions purely to gather information. This explore-to-know behavior is exactly what naive methods miss.

Partial observability is not a niche corner of reinforcement learning; it is the default condition of robots, autonomous vehicles, strategic games, dialogue systems, and today's agentic LLM systems. The field's main tools are:

  • Formal models -- the Partially Observable Markov Decision Process (POMDP) for one agent; the Decentralized POMDP (Dec-POMDP) for cooperating teams; partially observable stochastic games, I-POMDPs, and extensive-form games for competitive settings.
  • Planning algorithms for when a model or simulator is available -- point-based value iteration, Monte Carlo tree search in belief space (POMCP, DESPOT, POMCPOW), and learned-heuristic planners such as BetaZero.
  • Learning algorithms for when it is not -- recurrent and transformer-based deep RL, learned world models (Dreamer), and privileged-information training (asymmetric actor-critic, teacher-student).
  • Belief tracking for LLM agents -- natural-language belief states, uncertainty-aware memory, external Bayesian belief engines, and clarification/information-seeking policies, an area that has grown quickly through 2025-2026.
The perception-action loop under partial observability
Hidden state st → Observation ot → Belief / memory bt → Policy π(a | bt) → Action at
what is really true (never seen directly) a noisy, partial clue everything the agent has inferred so far decision rule changes the world and what the agent will see next

Key Concepts

State (s)
The complete description of the environment that determines what happens next. In a POMDP it is hidden.
Observation (o)
What the agent actually receives -- a camera frame, a sensor reading, a web page, a user message, a card on the table.
History (ht)
The full sequence of actions and observations so far: a0, o1, a1, o2, ..., ot. Always sufficient for optimal decisions, but it grows without bound.
Belief state (b)
A probability distribution over possible current states given the history. When the model is known, the belief is a sufficient statistic: nothing else from the history is needed to act optimally.
Information state / agent state
Any compact summary of history that is sufficient (or approximately sufficient) for prediction and control -- the hidden state of an LSTM, the latent state of a world model, or an LLM agent's running notes. Approximate information state theory formalizes when such learned summaries are good enough.
Perceptual aliasing
Different underlying states that produce the same observation (two identical-looking corridors). The central reason memoryless policies fail.
Value of information
How much expected reward improves by reducing uncertainty before acting. It is what makes "look before you leap" (or "ask before you act") rational.
Imperfect vs. incomplete information
In game theory, imperfect information means players cannot see some moves or state (hidden cards); incomplete information means they are unsure of others' payoffs or types. The Harsanyi transformation converts the latter into the former (see below).

The Family of Models

Model Agents What each agent sees Objective Typical solution approaches Worst-case complexity (finite horizon)
MDP 1 Full state Maximize own return Dynamic programming, Q Learning, policy gradient P-complete
POMDP 1 Partial, noisy observations Maximize own return Belief-space planning, recurrent/transformer RL, world models PSPACE-complete (Papadimitriou & Tsitsiklis 1987); infinite-horizon optimal planning is undecidable (Madani, Hanks & Condon 1999)
Dec-POMDP n (team) Each agent: own local observations Maximize one shared team reward Centralized training with decentralized execution (CTDE), value decomposition, learned communication NEXP-complete, even for 2 agents (Bernstein et al. 2002)
POSG (partially observable stochastic game) n Local observations Each agent maximizes own reward Equilibrium computation, self-play, opponent modeling At least as hard as Dec-POMDP
I-POMDP (interactive POMDP) n Local observations Own reward, with nested beliefs about other agents' beliefs Particle filtering over agent models, bounded nesting Intractable in general; approximations required
Extensive-form game with imperfect information n (often 2) Information sets (hidden cards, pieces, units) Equilibrium (e.g., Nash) Counterfactual regret minimization, public belief states, belief networks + search Polynomial in game-tree size for two-player zero-sum (sequence form), but real trees are astronomically large

Partially Observable Markov Decision Processes (POMDP)

A POMDP extends an MDP with an observation model. The idea dates to Åström (1965) in control theory and Smallwood & Sondik (1973) in operations research, and was brought to AI by Kaelbling, Littman & Cassandra (1998).

The Belief Update (Bayes Filter)

After taking action a and receiving observation o, the agent updates its belief with Bayes' rule:

b′(s′) = η · O(o | s′, a) · Σs ∈ S T(s′ | s, a) · b(s)

The sum predicts where the world went; the observation term corrects that prediction with evidence; η renormalizes so the probabilities add up to 1. The same predict-correct loop appears as the Kalman filter (Gaussian beliefs), particle filters (sampled beliefs), and the forward algorithm of a Hidden Markov Model -- a POMDP is essentially an HMM that the agent can steer and that pays rewards. See Bayes.

Worked Example: The Tiger Problem

The classic teaching POMDP (Kaelbling et al. 1998): a tiger is behind one of two doors and treasure behind the other. The agent can open-left, open-right, or listen. Listening costs a little (-1) and reports the tiger's side correctly 85% of the time. Opening the treasure door pays +10; opening the tiger door costs -100.

Step Action Observation Belief "tiger is left"
0 -- -- 0.50
1 listen hear tiger left 0.85
2 listen hear tiger left 0.97
3 listen hear tiger right 0.85

With a 50/50 belief, opening a door has expected value 0.5(+10) + 0.5(-100) = -45, so the optimal policy listens -- an action with no direct payoff whose only purpose is to reduce uncertainty -- and opens the door opposite the tiger only once the belief is lopsided enough. Heuristics that assume the state will be known after the next step (such as QMDP, below) put no value on the information itself, which is why they break down on problems where looking is the only way forward.

Why POMDPs Are Hard

  • Curse of dimensionality -- the belief lives in a space with as many dimensions as there are states.
  • Curse of history -- the number of distinct action-observation histories grows exponentially with the planning horizon.
  • Complexity results -- finite-horizon POMDP planning is PSPACE-complete (Papadimitriou & Tsitsiklis 1987); deciding optimal infinite-horizon behavior is undecidable in general (Madani, Hanks & Condon 1999).

As Leslie Kaelbling has argued, these worst-case results are not a reason to give up: humans and animals act well under heavy partial observability, which suggests that the problems the physical world actually poses have exploitable structure (The Last Four Frames Is Not All You Need | L. Kaelbling - MIT CSAIL).

Planning: Solving POMDPs When the Model Is Known

Family Representative methods Core idea Best suited for
Exact value iteration Sondik's algorithm, Witness, Incremental Pruning Compute the full set of alpha-vectors Tiny problems; teaching
Fast heuristics QMDP, most-likely-state Solve the underlying MDP and average over the belief Quick baselines; problems where information gathering does not matter (QMDP assumes all uncertainty vanishes after one step, so it never values gathering information for its own sake)
Point-based offline solvers PBVI (Pineau, Gordon & Thrun 2003), Perseus, HSVI, SARSOP (Kurniawati, Hsu & Lee 2008) Back up the value function only at beliefs the agent can actually reach Discrete problems with thousands to hundreds of thousands of states; SARSOP won the RSS Test of Time Award in 2021
Online Monte Carlo tree search POMCP (Silver & Veness 2010), DESPOT (Ye et al. 2017), AdaOPS Plan from the current belief at decision time using a black-box simulator and particle beliefs Huge discrete problems (POMCP handled Battleship and partially observable PacMan with ~1018 and ~1056 states)
Continuous-space online planners POMCPOW & PFT-DPW (Sunberg & Kochenderfer 2018) Progressive widening over actions and observations; weighted particle beliefs Continuous states, actions, and observations (driving, robotics)
Learned guidance + search BetaZero (Moss et al. 2024), GammaZero, VOIMCP (2026) AlphaZero-style: offline-trained policy and value networks over beliefs replace hand-crafted heuristics in online search; VOIMCP skips observation branching when the value of information is low Long-horizon, high-dimensional problems such as critical-mineral exploration and carbon storage

Surveys: Online Planning Algorithms for POMDPs | S. Ross, J. Pineau, S. Paquet & B. Chaib-draa - JAIR 2008 ... POMDPs and Robotics | H. Kurniawati - Annual Review of Control, Robotics, and Autonomous Systems 2022 ... POMDPs in Robotics: A Survey | M. Lauri, D. Hsu & J. Pajarinen - IEEE Transactions on Robotics 2023

Learning Under Partial Observability

When there is no model or simulator to plan with, the agent has to learn both what to remember and what to do from experience. This is where deep reinforcement learning meets partial observability.

Memory-Based Model-Free RL

Transformers, State Space Models, and the Memory-Architecture Race

Long-horizon memory is the bottleneck in many POMDPs, and every sequence architecture has been tried:

Takeaway: there is no universally best memory. Match the architecture to the kind of memory the task needs (short vs. long, retain vs. revise, sparse vs. dense recall), and evaluate against a fully observable twin to separate memory failures from everything else.

World Models and Learned Belief States

Model-based agents learn a latent dynamics model whose hidden state acts as a learned belief:

Privileged Information: Training With the Answer Key

In simulation, the true state is usually available during training even though the deployed agent will never see it. Two popular ways to exploit this:

The catch -- the imitation gap. A teacher that knows where the hidden obstacle is never needs to look for it, so a student that copies it never learns to look either. Provable Partially Observable RL with Privileged Information | Y. Cai, X. Liu, A. Oikonomou et al. - 2024 formalized when expert distillation fails and gave a belief-weighted asymmetric actor-critic with polynomial sample complexity; Guided Policy Optimization under Partial Observability (2025) co-trains a guider and a learner so the guider stays imitable.

Generalization Is a Partial-Observability Problem

Why Generalization in RL Is Difficult: Epistemic POMDPs and Implicit Partial Observability | D. Ghosh et al. - NeurIPS 2021 showed that an agent trained on a finite set of environments is, at test time, uncertain about which environment it is in. Even when every individual environment is fully observable, the test-time problem is a POMDP, and policies that act as if they knew the environment generalize poorly. This connects partial observability to out-of-distribution generalization and to meta-RL, where methods such as VariBAD (Zintgraf et al. 2020) maintain a belief over the task itself. See also in-context reinforcement learning.

Active Information Gathering

Under partial observability, exploration has two purposes: finding reward and finding out. Approaches include planning in belief space (which values information automatically), information-gain and curiosity bonuses, and Bayes-adaptive methods. Getting this right is still one of the hardest parts of the problem, for deep RL agents and LLM agents alike (see below).

Decentralized Partially Observable Markov Decision Processes (Dec-POMDP)

Models in which agents have incomplete local information and must coordinate their decisions.

A Dec-POMDP is a team of n agents sharing one reward. Each agent i chooses its action from its own local history of observations; the joint action drives the shared hidden state. Formally: ⟨I, S, {Ai}, T, R, {Ωi}, O, h⟩, where I is the set of agents, Ai and Ωi are the actions and observations of agent i, and h is the horizon.

Why this is so much harder than a single-agent POMDP:

Core reference: A Concise Introduction to Decentralized POMDPs | F. Oliehoek & C. Amato - Springer 2016. Recent overview: (A Partial Survey of) Decentralized, Cooperative Multi-Agent Reinforcement Learning | C. Amato - 2024.

Training Paradigms

Paradigm Training Execution Pros Cons
Centralized training & execution (CTCE) One controller sees everything One controller acts for all Simplest coordination Needs perfect, instant communication; joint action space explodes
Centralized training, decentralized execution (CTDE) Uses global state and all agents' data Each agent acts on local observations only Most widely used; exploits simulator information safely Training/deployment information gap (the multi-agent version of the imitation gap)
Decentralized training & execution (DTE) Each agent learns independently Local Fewest assumptions; works when agents meet online Non-stationarity: every teammate is a moving target

Key Algorithms

Deep Recurrent and Distributed Q-Networks

The deep-learning line of work on multi-agent partial observability grew directly out of DRQN:

Games with hidden "types": the Harsanyi transformation. When agents are uncertain about each other's private characteristics -- payoffs, goals, capabilities, or roles -- the setting is a game of incomplete information. John Harsanyi (1967-68) showed how to handle it: add an initial chance move in which "Nature" draws each player's type from a common prior, and let each player observe only its own type. The game becomes one of imperfect information (a Bayesian game) that can be analyzed with standard tools, leading to the Bayes-Nash equilibrium. The same idea underlies type-based reasoning in multi-agent RL: ad hoc teamwork with unknown partners, opponent modeling, and the interactive POMDPs below, where agents keep beliefs over the possible types and models of other agents.

Beyond Teams: POSGs and I-POMDPs

When agents have their own rewards, the setting becomes a partially observable stochastic game (POSG). Interactive POMDPs (Gmytrasiewicz & Doshi - JAIR 2005) give each agent beliefs over both the physical state and the models of other agents -- including their beliefs about it -- truncated at some nesting depth. The AAMAS 2016 paper linked at the top of this page studies model-free learning with PAC guarantees in such settings.

Multi-Agent Benchmarks

Imperfect-Information Games

Games with hidden information have been the field's public scoreboard for partial observability. Unlike chess or Go (AlphaGo Zero), a player cannot simply search forward from "the" position, because it does not know the position; it knows only an information set of positions consistent with what it has seen. See Game Theory and Gaming.

Year System Game What hid the state Key idea
2017 DeepStack / Libratus Heads-up no-limit Texas hold'em Opponent's cards Counterfactual regret minimization (CFR), depth-limited re-solving during play
2019 Pluribus Six-player no-limit hold'em Five opponents' cards Blueprint strategy from self-play plus limited-lookahead search
2019 AlphaStar StarCraft II Fog of war, scouting LSTM/transformer-based agents trained with league self-play; Grandmaster level
2020 ReBeL Poker, Liar's Dice Private hands Public belief states make imperfect-information games searchable like perfect-information ones
2022 DeepNash Stratego Identity of 40 hidden pieces Model-free regularized Nash dynamics; reached the all-time top three on the Gravon platform
2022 Cicero Diplomacy Other players' intentions Language model plus strategic planning; human-level negotiation
2023 Student of Games Chess, Go, poker, Scotland Yard Varies One algorithm for both perfect- and imperfect-information games
2025-26 Ataraxos Stratego (and Barrage Stratego, Hanabi, Dou Dizhu) Hidden pieces and cards Self-play policy-value network + a separate belief network predicting opponents' hidden pieces + test-time search; published in Nature on 30 September 2026 as the first superhuman Stratego AI

The 2026 Ataraxos result is a useful marker of where the field stands. It beat Pim Niemeijer, described as the most decorated Stratego player in history, 15 wins to 1 with 4 draws, and was trained for under about $8,000 -- roughly 1/500th of the compute of DeepMind's DeepNash. The same recipe also set a new state of the art in cooperative Hanabi across two- to five-player variants. The broader lesson: explicitly modeling hidden information with a learned belief network, then searching over sampled hidden states at decision time, now scales to games with astronomically many hidden configurations (more than 1033 possible Stratego setups).

Partial Observability in LLM Agents

Every agentic LLM system is a POMDP agent, whether or not its designers say so:

  • the user's true goal and preferences are hidden and must be inferred from underspecified requests;
  • a web or computer-use agent sees one screen at a time; content below the fold, behind a login, or in another tab is unobserved;
  • a coding agent cannot read a whole codebase at once and must decide what to open, search, or run;
  • tool results are partial, stale, or wrong; the world can change between steps;
  • the context window is finite, so the agent must decide what to remember and what to forget -- see Context and Memory.

Typical Failure Modes

  • Premature commitment -- answering or acting before resolving ambiguity. Benchmarks such as InfoQuest (2025) found that current assistants need many turns to infer hidden user context and are inefficient at asking.
  • Belief inertia -- failing to overwrite outdated beliefs when the world changes. Theory of Space (2026) found foundation-model agents explore redundantly and struggle to revise spatial beliefs after objects are moved, especially in vision-based settings; Seeing Isn't Believing (2026) targets the same failure in embodied agents.
  • Collapsed uncertainty -- storing guesses as facts in memory or summaries, which then reinforce themselves.
  • Information self-locking -- On Information Self-Locking in RL for Active Reasoning of LLM Agents (2026) reports that outcome-based RL can train agents that stop asking informative questions and stop using the evidence they do get.
  • Inconsistent beliefs -- two histories that imply the same posterior can yield different actions from an LLM because their surface text differs.

Design Patterns for Belief Tracking

Pattern How it works Examples
Full history in context Append every action and observation; let the model infer the state implicitly ReAct-style agents (Yao et al. 2022) -- simple, but context grows without bound and errors hide in long transcripts
Natural-language belief bottleneck At each step, update a concise written belief, then act on the belief alone ABBEL (Lidayan et al. 2025) -- near-constant memory and interpretable beliefs; RL with belief grading narrows the gap to full-context agents (BAIR blog, July 2026)
Uncertainty-aware verbalized beliefs Annotate each belief claim with a likelihood phrase ("likely", "almost certainly not"); train the belief model and the policy jointly with RL Agent-BRACE (May 2026)
Probabilistic agent memory Store attribute-level distributions instead of deterministic conclusions, so early guesses do not harden into "facts" BeliefMem (2026)
External Bayesian belief engine A separate inference module maintains a POMDP-grounded belief; the LLM only chooses actions given that belief, which makes it a Markov policy on the belief MDP Belief-State Engine (Sept 2026) -- compared against QMDP and POMCP on the Tiger problem and a red-team attack-graph task
Belief-based world models World models that expose what is known and unknown, not just simulated futures Belief-Based World Models (2026) on ALFWorld and ScienceWorld
Structured / neuro-symbolic beliefs Keep beliefs in a knowledge graph; combine fast reactive and slow deliberative reasoning NeSyFS (2026)
Active information seeking & clarification Decide explicitly when to gather information or ask the user versus act InfoSeeker (2025), ClarifyBench (2025), RegretBench (2026), uncertainty decomposition for clarification (2026)

Benchmarks That Stress Hidden Information

  • ARC-AGI-3 | ARC Prize Foundation - 2026 -- launched 25 March 2026: interactive, turn-based environments with no instructions, where agents must explore, infer the goal, model the dynamics, and plan. At launch every frontier model tested scored below 1% while humans solved all environments. The ARC Prize 2026 competition on it takes final submissions in early November 2026.
  • τ-bench (2024) -- tool-using agents converse with a simulated user who holds information the agent must elicit.
  • ALFWorld, ScienceWorld, TextWorld -- text-based embodied tasks with unobserved object locations.
  • WebArena and OSWorld -- web and desktop agents operating with partial views of complex interfaces.
  • InfoQuest and RegretBench -- multi-turn dialogue with hidden user context or intent.

Practical Guidance for Agent Builders

  1. Keep an explicit state. Maintain a running notes or state file separate from the raw transcript, and update it every step.
  2. Record uncertainty, not just conclusions. Write "probably X (unverified)" rather than "X." Distinguish what was observed from what was inferred.
  3. Separate belief update from action selection. First ask "what do I now believe?", then "what should I do?"
  4. Price information. Ask or explore when a wrong guess is expensive or irreversible; act directly when it is cheap to redo.
  5. Verify before irreversible actions, and re-read the world after acting -- it may have changed.
  6. Test with twins. Evaluate the same task with and without the hidden information to see how much performance is lost to partial observability.

A Safety Angle: Partial Observability of the Overseer

Partial observability affects the humans who train and supervise AI as well. When Your AIs Deceive You: Challenges of Partial Observability in RLHF | L. Lang et al. - NeurIPS 2024 showed that when evaluators see only part of what an agent did, RLHF can reward "deceptive inflation" (making results look better than they are) and "overjustification" (spending effort to look good rather than be good). Better oversight tools are, in part, tools for reducing the overseer's partial observability. See Explainable / Interpretable AI.

Applications

Domain What is hidden Typical approaches
Robotics and manipulation Object poses behind occlusion, contact state, the robot's own position Particle filters and SLAM, belief-space planning (SARSOP, POMCPOW), recurrent policies, teacher-student sim-to-real
Autonomous driving and drones Occluded road users, other drivers' intentions Occlusion-aware and intention-aware POMDP planning, online tree search in continuous spaces
Dialogue and assistants (Conversational AI) User goals, preferences, context Statistical POMDP dialogue managers; today, LLM belief tracking and clarification policies
Web, desktop, and coding agents (Agents/Assistants) Off-screen content, backend state, unread files Exploration, explicit notes and memory, verification steps
Strategy games and negotiation (Negotiation) Opponents' cards, units, plans, preferences CFR, public belief states, belief networks + search, opponent modeling
Healthcare Patient's underlying disease state POMDP treatment and testing policies that weigh diagnostic tests against treatment
Earth science and sustainability Subsurface ore bodies, CO2 plumes, groundwater contamination Large-scale information-gathering POMDPs solved with DESPOT, POMCPOW, BetaZero
Cybersecurity and Defense Attacker presence, location, and capabilities POMDP-based intrusion response and attack-graph planning; partially observable stochastic games
Multi-robot, warehouse, and traffic control Teammates' observations, global state Dec-POMDP formulations, CTDE multi-agent RL, learned communication
Recommendation A user's latent and shifting interests RL and bandits with latent user state

Benchmarks and Environments

Benchmark Year Setting What it tests
Flickering Atari (DRQN paper) 2015 Single-agent, pixels Integrating information across randomly blanked frames
Hanabi Learning Environment 2019 Cooperative card game Theory of mind, implicit communication
Memory Maze 2022 3D navigation Long-term spatial memory
POPGym 2023 15 low-dimensional POMDPs Broad, fast comparison of memory models
SMACv2 2023 StarCraft II multi-agent Decentralized control under stochasticity and partial observability
Craftax 2024 Open-ended survival (JAX) Exploration and memory in a partially observed world
POPGym Arcade 2025 Pixel games with MDP/POMDP twins Controlled studies of memory and observability
MIKASA (Base and Robo) 2025 Memory RL; robotic manipulation Taxonomy-driven memory evaluation
InfoQuest 2025 Multi-turn chat Uncovering hidden user context through questions
Theory of Space 2026 Embodied foundation-model agents Active exploration and spatial belief revision
ARC-AGI-3 2026 Interactive abstract games Exploration, goal inference, and model building with no instructions

Software and Tools

Tool Language What it offers
POMDPs.jl Julia A modeling interface plus a large ecosystem of solvers (SARSOP, POMCP, POMCPOW, DESPOT, QMDP, and more) and belief updaters
pomdp_py Python Python framework for building and solving POMDPs, with POMCP and value iteration
SARSOP and DESPOT C++ Reference implementations of the leading offline point-based solver and online tree-search solver
AI-Toolbox C++ / Python MDP, POMDP, and multi-agent planning and learning algorithms
BetaZero.jl Julia Learned belief-space planning
Stable-Baselines3 Contrib Python / PyTorch RecurrentPPO (LSTM policies) for partially observable environments
CleanRL Python Single-file implementations, including PPO with LSTM
POPGym Python Partially observable environments and memory-model baselines
DreamerV3 Python / JAX World-model RL agent
PettingZoo Python Standard API and environments for multi-agent RL
EPyMARL and JaxMARL Python / JAX Cooperative MARL algorithms (QMIX, MAPPO, and others) and benchmarks
OpenSpiel C++ / Python Imperfect-information games and algorithms such as CFR

Choosing an Approach

If your situation is... Start with...
Small, discrete, model known Exact or point-based solvers (SARSOP)
Large or continuous, simulator available Online belief-space search (POMCP, DESPOT, POMCPOW); add learned guidance (BetaZero) for long horizons
No model, need sample efficiency A learned world model (DreamerV3)
No model, short-to-medium memory Recurrent PPO or R2D2-style recurrent value learning; benchmark on POPGym
Very long-range recall Transformer or SSM memory, tested against recurrent baselines; check whether you need retention or revision
Simulator exposes the true state during training Asymmetric actor-critic or teacher-student, watching for the imitation gap
Several cooperating agents Model as a Dec-POMDP; CTDE methods (MAPPO, QMIX), plus learned communication if allowed
Adversaries hiding information CFR or public-belief-state methods; belief network + test-time search
An LLM agent Explicit, uncertainty-aware belief state; clarification and information-seeking policy; verification before irreversible actions

Open Problems and Frontiers

  • Calibrated beliefs in open-ended worlds. Classical belief tracking assumes a known, finite state space. Text and web environments have neither, and LLM agents still struggle to maintain and revise beliefs over long horizons.
  • Memory that can be revised, not just retained. Recent benchmarks show that remembering is easier than updating; architectures that do both reliably remain open.
  • Knowing what you do not know. Deciding when to explore, ask, or act -- and training agents (with RL or otherwise) that do not lose the habit of information seeking.
  • Decentralized LLM agents. Multi-agent LLM systems inherit Dec-POMDP difficulty: what to share, what teammates know, and how to coordinate without a shared view.
  • Theory for privileged and hindsight information. When does training-time access to the true state provably help, and how can agents be trained to look for what the teacher simply knew?
  • Evaluation hygiene. Controlling for observability (MDP/POMDP twins) so that memory, credit assignment, exploration, and generalization failures are not confused with one another.
  • Oversight under partial observability. Making sure human and automated evaluators can see enough of an agent's behavior to reward what is actually good.

Timeline

Year Milestone
1965 Åström formulates optimal control with incomplete state information
1973 Smallwood & Sondik: POMDP value functions are piecewise-linear and convex
1987 Papadimitriou & Tsitsiklis: finite-horizon POMDPs are PSPACE-complete
1998 Kaelbling, Littman & Cassandra bring POMDPs into mainstream AI
1999 Madani, Hanks & Condon: infinite-horizon POMDP planning is undecidable
2002 Bernstein et al.: Dec-POMDPs are NEXP-complete
2003-2008 Point-based solvers: PBVI, Perseus, HSVI, SARSOP
2010 POMCP brings Monte Carlo tree search to large POMDPs
2015 DRQN: deep recurrent Q-learning for POMDPs
2016-2017 DDRQN, DIAL, Dec-HDRQN; DeepStack and Libratus beat poker professionals
2018 QMIX; asymmetric actor-critic; POMCPOW for continuous POMDPs
2019 R2D2; Pluribus; AlphaStar; the Hanabi challenge
2020 GTrXL; ReBeL and public belief states
2021-2022 Epistemic POMDPs; MAPPO; DeepNash and Cicero; recurrent baselines shown to be strong
2023 POPGym; SMACv2; Student of Games; "When do transformers shine in RL?"
2024 BetaZero; RLHF under partial observability; evidence that transformers represent belief states internally
2025 DreamerV3 published in Nature (April); Belief State Transformer; POPGym Arcade; MIKASA; ABBEL and the first wave of LLM belief-state methods
2026 ARC-AGI-3 launches (March); Agent-BRACE, BeliefMem, Belief-State Engine, and belief-based world models for LLM agents; Ataraxos, the first superhuman Stratego AI, published in Nature (30 September)

Videos and Lectures

Further Reading