Partial Observability
YouTube ... Quora ...Google search ...Google News ...Bing News
- Architectures for AI ... Generative AI Stack ... Enterprise Architecture (EA) ... Enterprise Portfolio Management (EPM) ... Architecture and Interior Design
- Artificial Intelligence (AI) ... Generative AI ... Machine Learning (ML) ... Deep Learning ... Neural Network ... Reinforcement ... Learning Techniques
- Reinforcement Learning (RL) ... Markov Decision Process (MDP) ... Hidden Markov Model ... Monte Carlo ... Q Learning ... Deep Q Network (DQN) ... Actor Critic ... PPO ... Policy ... Policy vs Plan
- Memory ... Recurrent Neural Network (RNN) ... Long Short-Term Memory (LSTM) ... Gated Recurrent Unit (GRU) ... Transformer ... Attention ... State Space Model (SSM) ... Mamba ... World Models
- Agents/Assistants ... Robotic Process Automation ... Personal Companions ... Productivity ... Email ... Negotiation ... LangChain
- Large Language Model (LLM) ... Context ... In-Context Learning (ICL) ... RLHF ... Conversational AI
- Game Theory ... Gaming ... Google DeepMind AlphaGo Zero ... Robotics ... Embodied AI ... Autonomous Vehicles ... Bayes ... Curiosity
- SMART - Multi-Task Deep Neural Networks (MT-DNN) ... Multi-Task Learning (MTL)
- Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability | S. Omidshafiei, J. Pazis, C. Amato, J. How & J. Vian - ICML 2017
- Deep Multiagent Reinforcement Learning for Partially Observable Parameterized Environments | M. Hausknecht & P. Stone - UT Austin
- Reinforcement Learning in Partially Observable Multiagent Settings: Monte Carlo Exploring Policies with PAC Bounds | R. Ceren, P. Doshi & B. Banerjee - AAMAS 2016
- Planning and Acting in Partially Observable Stochastic Domains | L. Kaelbling, M. Littman & A. Cassandra - Artificial Intelligence 1998 ... the classic POMDP reference
- Algorithms for Decision Making | M. Kochenderfer, T. Wheeler & K. Wray - MIT Press ... free textbook; Part III covers state uncertainty and POMDPs
- Belief-space planning and POMDPs | R. Tedrake - MIT Underactuated Robotics
In the textbook Markov Decision Process (MDP), an agent is handed the complete state of its environment at every step: the full chessboard, the exact position and velocity of the robot, everything that matters for predicting what happens next. Real problems almost never look like that. A robot's camera cannot see behind the box it is about to pick up. A self-driving car cannot see the pedestrian hidden by a parked van, or read the intentions of the driver beside it. A poker bot cannot see its opponents' cards. A customer-service assistant does not know what the customer actually wants until it asks. A coding agent cannot read an entire repository at once, and a web agent sees one page at a time.
Partial observability in one sentence: the agent never sees the true state of the world, only clues about it, so it must remember and infer before it can decide -- and sometimes it must act just to find out more.
These are all problems of partial observability: the agent receives observations that carry only incomplete, noisy, or ambiguous information about the underlying state. Partial observability has three immediate consequences:
- The Markov property breaks for observations. The current observation alone no longer tells the agent everything it needs. Two situations that look identical (perceptual aliasing) may require different actions, so a policy that maps the latest observation straight to an action can be arbitrarily bad.
- The agent needs an internal state. To act well it must summarize its history of actions and observations, either explicitly as a probability distribution over possible states (a belief state) or implicitly in a learned memory such as a recurrent network, a Transformer context window, a state space model, or a written note in an LLM agent's scratchpad.
- Information becomes something worth acting for. Because uncertainty can be reduced by looking, listening, probing, or asking, an optimal agent sometimes chooses actions purely to gather information. This explore-to-know behavior is exactly what naive methods miss.
Partial observability is not a niche corner of reinforcement learning; it is the default condition of robots, autonomous vehicles, strategic games, dialogue systems, and today's agentic LLM systems. The field's main tools are:
- Formal models -- the Partially Observable Markov Decision Process (POMDP) for one agent; the Decentralized POMDP (Dec-POMDP) for cooperating teams; partially observable stochastic games, I-POMDPs, and extensive-form games for competitive settings.
- Planning algorithms for when a model or simulator is available -- point-based value iteration, Monte Carlo tree search in belief space (POMCP, DESPOT, POMCPOW), and learned-heuristic planners such as BetaZero.
- Learning algorithms for when it is not -- recurrent and transformer-based deep RL, learned world models (Dreamer), and privileged-information training (asymmetric actor-critic, teacher-student).
- Belief tracking for LLM agents -- natural-language belief states, uncertainty-aware memory, external Bayesian belief engines, and clarification/information-seeking policies, an area that has grown quickly through 2025-2026.
| Hidden state st | → | Observation ot | → | Belief / memory bt | → | Policy π(a | bt) | → | Action at |
|---|---|---|---|---|---|---|---|---|
| what is really true (never seen directly) | a noisy, partial clue | everything the agent has inferred so far | decision rule | changes the world and what the agent will see next |
Key Concepts
- State (s)
- The complete description of the environment that determines what happens next. In a POMDP it is hidden.
- Observation (o)
- What the agent actually receives -- a camera frame, a sensor reading, a web page, a user message, a card on the table.
- History (ht)
- The full sequence of actions and observations so far: a0, o1, a1, o2, ..., ot. Always sufficient for optimal decisions, but it grows without bound.
- Belief state (b)
- A probability distribution over possible current states given the history. When the model is known, the belief is a sufficient statistic: nothing else from the history is needed to act optimally.
- Information state / agent state
- Any compact summary of history that is sufficient (or approximately sufficient) for prediction and control -- the hidden state of an LSTM, the latent state of a world model, or an LLM agent's running notes. Approximate information state theory formalizes when such learned summaries are good enough.
- Perceptual aliasing
- Different underlying states that produce the same observation (two identical-looking corridors). The central reason memoryless policies fail.
- Value of information
- How much expected reward improves by reducing uncertainty before acting. It is what makes "look before you leap" (or "ask before you act") rational.
- Imperfect vs. incomplete information
- In game theory, imperfect information means players cannot see some moves or state (hidden cards); incomplete information means they are unsure of others' payoffs or types. The Harsanyi transformation converts the latter into the former (see below).
The Family of Models
| Model | Agents | What each agent sees | Objective | Typical solution approaches | Worst-case complexity (finite horizon) |
|---|---|---|---|---|---|
| MDP | 1 | Full state | Maximize own return | Dynamic programming, Q Learning, policy gradient | P-complete |
| POMDP | 1 | Partial, noisy observations | Maximize own return | Belief-space planning, recurrent/transformer RL, world models | PSPACE-complete (Papadimitriou & Tsitsiklis 1987); infinite-horizon optimal planning is undecidable (Madani, Hanks & Condon 1999) |
| Dec-POMDP | n (team) | Each agent: own local observations | Maximize one shared team reward | Centralized training with decentralized execution (CTDE), value decomposition, learned communication | NEXP-complete, even for 2 agents (Bernstein et al. 2002) |
| POSG (partially observable stochastic game) | n | Local observations | Each agent maximizes own reward | Equilibrium computation, self-play, opponent modeling | At least as hard as Dec-POMDP |
| I-POMDP (interactive POMDP) | n | Local observations | Own reward, with nested beliefs about other agents' beliefs | Particle filtering over agent models, bounded nesting | Intractable in general; approximations required |
| Extensive-form game with imperfect information | n (often 2) | Information sets (hidden cards, pieces, units) | Equilibrium (e.g., Nash) | Counterfactual regret minimization, public belief states, belief networks + search | Polynomial in game-tree size for two-player zero-sum (sequence form), but real trees are astronomically large |
Partially Observable Markov Decision Processes (POMDP)
A POMDP extends an MDP with an observation model. The idea dates to Åström (1965) in control theory and Smallwood & Sondik (1973) in operations research, and was brought to AI by Kaelbling, Littman & Cassandra (1998).
The Belief Update (Bayes Filter)
After taking action a and receiving observation o, the agent updates its belief with Bayes' rule:
b′(s′) = η · O(o | s′, a) · Σs ∈ S T(s′ | s, a) · b(s)
The sum predicts where the world went; the observation term corrects that prediction with evidence; η renormalizes so the probabilities add up to 1. The same predict-correct loop appears as the Kalman filter (Gaussian beliefs), particle filters (sampled beliefs), and the forward algorithm of a Hidden Markov Model -- a POMDP is essentially an HMM that the agent can steer and that pays rewards. See Bayes.
Worked Example: The Tiger Problem
The classic teaching POMDP (Kaelbling et al. 1998): a tiger is behind one of two doors and treasure behind the other. The agent can open-left, open-right, or listen. Listening costs a little (-1) and reports the tiger's side correctly 85% of the time. Opening the treasure door pays +10; opening the tiger door costs -100.
| Step | Action | Observation | Belief "tiger is left" |
|---|---|---|---|
| 0 | -- | -- | 0.50 |
| 1 | listen | hear tiger left | 0.85 |
| 2 | listen | hear tiger left | 0.97 |
| 3 | listen | hear tiger right | 0.85 |
With a 50/50 belief, opening a door has expected value 0.5(+10) + 0.5(-100) = -45, so the optimal policy listens -- an action with no direct payoff whose only purpose is to reduce uncertainty -- and opens the door opposite the tiger only once the belief is lopsided enough. Heuristics that assume the state will be known after the next step (such as QMDP, below) put no value on the information itself, which is why they break down on problems where looking is the only way forward.
Why POMDPs Are Hard
- Curse of dimensionality -- the belief lives in a space with as many dimensions as there are states.
- Curse of history -- the number of distinct action-observation histories grows exponentially with the planning horizon.
- Complexity results -- finite-horizon POMDP planning is PSPACE-complete (Papadimitriou & Tsitsiklis 1987); deciding optimal infinite-horizon behavior is undecidable in general (Madani, Hanks & Condon 1999).
As Leslie Kaelbling has argued, these worst-case results are not a reason to give up: humans and animals act well under heavy partial observability, which suggests that the problems the physical world actually poses have exploitable structure (The Last Four Frames Is Not All You Need | L. Kaelbling - MIT CSAIL).
Planning: Solving POMDPs When the Model Is Known
| Family | Representative methods | Core idea | Best suited for |
|---|---|---|---|
| Exact value iteration | Sondik's algorithm, Witness, Incremental Pruning | Compute the full set of alpha-vectors | Tiny problems; teaching |
| Fast heuristics | QMDP, most-likely-state | Solve the underlying MDP and average over the belief | Quick baselines; problems where information gathering does not matter (QMDP assumes all uncertainty vanishes after one step, so it never values gathering information for its own sake) |
| Point-based offline solvers | PBVI (Pineau, Gordon & Thrun 2003), Perseus, HSVI, SARSOP (Kurniawati, Hsu & Lee 2008) | Back up the value function only at beliefs the agent can actually reach | Discrete problems with thousands to hundreds of thousands of states; SARSOP won the RSS Test of Time Award in 2021 |
| Online Monte Carlo tree search | POMCP (Silver & Veness 2010), DESPOT (Ye et al. 2017), AdaOPS | Plan from the current belief at decision time using a black-box simulator and particle beliefs | Huge discrete problems (POMCP handled Battleship and partially observable PacMan with ~1018 and ~1056 states) |
| Continuous-space online planners | POMCPOW & PFT-DPW (Sunberg & Kochenderfer 2018) | Progressive widening over actions and observations; weighted particle beliefs | Continuous states, actions, and observations (driving, robotics) |
| Learned guidance + search | BetaZero (Moss et al. 2024), GammaZero, VOIMCP (2026) | AlphaZero-style: offline-trained policy and value networks over beliefs replace hand-crafted heuristics in online search; VOIMCP skips observation branching when the value of information is low | Long-horizon, high-dimensional problems such as critical-mineral exploration and carbon storage |
Surveys: Online Planning Algorithms for POMDPs | S. Ross, J. Pineau, S. Paquet & B. Chaib-draa - JAIR 2008 ... POMDPs and Robotics | H. Kurniawati - Annual Review of Control, Robotics, and Autonomous Systems 2022 ... POMDPs in Robotics: A Survey | M. Lauri, D. Hsu & J. Pajarinen - IEEE Transactions on Robotics 2023
Learning Under Partial Observability
When there is no model or simulator to plan with, the agent has to learn both what to remember and what to do from experience. This is where deep reinforcement learning meets partial observability.
Memory-Based Model-Free RL
- Frame stacking -- the original DQN fed the last four Atari frames to the network. That recovers velocity but nothing older, and fails whenever something important happened more than a few frames ago.
- Recurrent policies -- Deep Recurrent Q-Learning for Partially Observable MDPs (DRQN) | M. Hausknecht & P. Stone - 2015 replaced DQN's first fully connected layer with an LSTM and showed it coped with "flickering" Atari, where whole frames are randomly blanked out. R2D2 | S. Kapturowski et al. - ICLR 2019 made recurrent replay practical at scale by storing recurrent states and using a "burn-in" warm-up. Recurrent PPO (LSTM or GRU) remains the most common practical baseline.
- Recurrent baselines are strong -- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs | T. Ni, B. Eysenbach & R. Salakhutdinov - ICML 2022 found that a carefully tuned recurrent off-policy agent matched or beat many specialized POMDP methods.
Transformers, State Space Models, and the Memory-Architecture Race
Long-horizon memory is the bottleneck in many POMDPs, and every sequence architecture has been tried:
- Stabilizing Transformers for RL (GTrXL) | E. Parisotto et al. - ICML 2020 -- gated Transformer-XL that made attention-based memory trainable with RL.
- When Do Transformers Shine in RL? Decoupling Memory from Credit Assignment | T. Ni et al. - NeurIPS 2023 -- transformers greatly extend how far back an agent can remember, but do not fix long-term credit assignment; the two are different problems and need different tests.
- Structured State Space Models for In-Context RL | C. Lu et al. - NeurIPS 2023 -- SSMs (S4/S5, and later Mamba-style models) offer transformer-like memory with RNN-like inference cost.
- AMAGO | J. Grigsby, L. Fan & Y. Zhu - ICLR 2024 -- scalable off-policy RL with long-context transformers for memory and meta-learning.
- POPGym | S. Morad et al. - ICLR 2023 -- 15 partially observable environments and 13 memory baselines; at publication, the largest head-to-head comparison of RL memory models.
- POPGym Arcade | Z. Wang, Z. He, E. Toledo & S. Morad - 2025 -- pixel-based, GPU-accelerated environments that each come as a fully observable and partially observable "twin," enabling controlled counterfactual studies. Its analysis found that agents with long-term memory can learn brittle policies that generalize poorly, and that recurrent policies can be "poisoned" by old out-of-distribution observations -- a warning for sim-to-real transfer, imitation learning, and offline RL.
- MIKASA: Memory-Intensive Skills Assessment Suite for Agents | E. Cherepanov et al. - ICLR 2026 -- a taxonomy of memory tasks plus 32 memory-intensive robotic manipulation tasks (MIKASA-Robo).
- Memory Retention Is Not Enough to Master Memory Tasks in RL - Reinforcement Learning Journal 2026 -- on tasks where stored information must be revised, not just retained, classic recurrent models held up better than transformer-based agents.
Takeaway: there is no universally best memory. Match the architecture to the kind of memory the task needs (short vs. long, retain vs. revise, sparse vs. dense recall), and evaluate against a fully observable twin to separate memory failures from everything else.
World Models and Learned Belief States
Model-based agents learn a latent dynamics model whose hidden state acts as a learned belief:
- Recurrent state-space models -- the PlaNet/Dreamer family combine a deterministic recurrent path with a stochastic latent, filtering observations into a compact state from which the agent "imagines" futures. Mastering Diverse Control Tasks Through World Models (DreamerV3) | D. Hafner, J. Pasukonis, J. Ba & T. Lillicrap - Nature, April 2025 outperformed specialized methods on more than 150 tasks with a single configuration and was the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. See World Models.
- Deep Variational Reinforcement Learning for POMDPs | M. Igl et al. - ICML 2018 -- trains a particle-filter-style belief inside the agent.
- Predictive state representations (Littman, Sutton & Singh 2001) -- represent "state" purely as predictions about future observations, avoiding hidden variables altogether.
- Approximate Information State | J. Subramanian, A. Sinha, R. Seraj & A. Mahajan - JMLR 2022 -- theory for when a learned history summary is good enough to plan with, with bounds on the loss from approximation.
- The Belief State Transformer | E. Hu et al. - ICLR 2025 -- a next-token predictor trained on both prefixes and suffixes so that it learns a compact belief state, improving goal-conditioned generation and planning.
- Transformers Represent Belief State Geometry in Their Residual Stream | A. Shai et al. - 2024 -- interpretability evidence that sequence models trained only on next-token prediction internally represent Bayesian belief states over hidden data-generating processes. See Explainable / Interpretable AI.
Privileged Information: Training With the Answer Key
In simulation, the true state is usually available during training even though the deployed agent will never see it. Two popular ways to exploit this:
- Asymmetric actor-critic -- the critic sees the full state while the actor sees only observations (Pinto et al. 2018). Naive versions give biased gradients; Baisero & Amato 2022 showed how to make them unbiased by conditioning the critic on both history and state, and Informed Asymmetric Actor-Critic (2025) extends this to critics that see only partial privileged signals.
- Teacher-student distillation -- train an expert with full state, then distill it into an observation-only student. This powered landmark legged-robot results such as Learning Quadrupedal Locomotion over Challenging Terrain | J. Lee et al. - Science Robotics 2020.
The catch -- the imitation gap. A teacher that knows where the hidden obstacle is never needs to look for it, so a student that copies it never learns to look either. Provable Partially Observable RL with Privileged Information | Y. Cai, X. Liu, A. Oikonomou et al. - 2024 formalized when expert distillation fails and gave a belief-weighted asymmetric actor-critic with polynomial sample complexity; Guided Policy Optimization under Partial Observability (2025) co-trains a guider and a learner so the guider stays imitable.
Generalization Is a Partial-Observability Problem
Why Generalization in RL Is Difficult: Epistemic POMDPs and Implicit Partial Observability | D. Ghosh et al. - NeurIPS 2021 showed that an agent trained on a finite set of environments is, at test time, uncertain about which environment it is in. Even when every individual environment is fully observable, the test-time problem is a POMDP, and policies that act as if they knew the environment generalize poorly. This connects partial observability to out-of-distribution generalization and to meta-RL, where methods such as VariBAD (Zintgraf et al. 2020) maintain a belief over the task itself. See also in-context reinforcement learning.
Active Information Gathering
Under partial observability, exploration has two purposes: finding reward and finding out. Approaches include planning in belief space (which values information automatically), information-gain and curiosity bonuses, and Bayes-adaptive methods. Getting this right is still one of the hardest parts of the problem, for deep RL agents and LLM agents alike (see below).
Decentralized Partially Observable Markov Decision Processes (Dec-POMDP)
Models in which agents have incomplete local information and must coordinate their decisions.
A Dec-POMDP is a team of n agents sharing one reward. Each agent i chooses its action from its own local history of observations; the joint action drives the shared hidden state. Formally: ⟨I, S, {Ai}, T, R, {Ωi}, O, h⟩, where I is the set of agents, Ai and Ωi are the actions and observations of agent i, and h is the horizon.
Why this is so much harder than a single-agent POMDP:
- There is no single shared belief. Each agent must reason not only about the world but about what its teammates have seen and will do.
- The Complexity of Decentralized Control of Markov Decision Processes | D. Bernstein, R. Givan, N. Immerman & S. Zilberstein - Mathematics of Operations Research 2002 proved that finite-horizon Dec-POMDPs are NEXP-complete even with two agents -- they provably admit no polynomial-time algorithm, and decentralization, not just uncertainty, is the source of the jump.
Core reference: A Concise Introduction to Decentralized POMDPs | F. Oliehoek & C. Amato - Springer 2016. Recent overview: (A Partial Survey of) Decentralized, Cooperative Multi-Agent Reinforcement Learning | C. Amato - 2024.
Training Paradigms
| Paradigm | Training | Execution | Pros | Cons |
|---|---|---|---|---|
| Centralized training & execution (CTCE) | One controller sees everything | One controller acts for all | Simplest coordination | Needs perfect, instant communication; joint action space explodes |
| Centralized training, decentralized execution (CTDE) | Uses global state and all agents' data | Each agent acts on local observations only | Most widely used; exploits simulator information safely | Training/deployment information gap (the multi-agent version of the imitation gap) |
| Decentralized training & execution (DTE) | Each agent learns independently | Local | Fewest assumptions; works when agents meet online | Non-stationarity: every teammate is a moving target |
Key Algorithms
- Value decomposition -- VDN (Sunehag et al. 2017) sums per-agent values; QMIX (Rashid et al. 2018) mixes them monotonically so each agent's greedy action is consistent with the team's.
- Centralized critics -- MADDPG (Lowe et al. 2017), COMA (Foerster et al. 2018) with counterfactual baselines, and MAPPO (Yu et al. 2022), which showed that well-tuned PPO with a centralized value function is a very strong cooperative baseline.
- Learned communication -- RIAL & DIAL (Foerster et al. 2016), CommNet (Sukhbaatar et al. 2016): agents learn what to tell each other to overcome local observability.
Deep Recurrent and Distributed Q-Networks
The deep-learning line of work on multi-agent partial observability grew directly out of DRQN:
- DDRQN -- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks | J. Foerster, Y. Assael, N. de Freitas & S. Whiteson - 2016 extended DRQN to teams with three changes: feeding each agent its own last action as input, sharing network weights across agents (with an agent ID input), and disabling experience replay, which becomes misleading when teammates are learning simultaneously. The agents solved classic multi-agent riddles that require inventing a communication protocol.
- Dec-HDRQN -- Deep Decentralized Multi-task Multi-Agent RL under Partial Observability | S. Omidshafiei et al. - ICML 2017 combined DRQN with hysteretic Q-learning (learning more slowly from negative surprises, which are often caused by teammates exploring rather than by bad actions), concurrent replay of synchronized trajectories, and policy distillation to merge specialist policies into one multi-task team policy.
- Parameterized action spaces -- Hausknecht and Stone's work (linked at the top of this page) applied deep multi-agent RL to RoboCup soccer ("Half Field Offense"), where actions combine a discrete choice (kick, dash, turn) with continuous parameters.
Games with hidden "types": the Harsanyi transformation. When agents are uncertain about each other's private characteristics -- payoffs, goals, capabilities, or roles -- the setting is a game of incomplete information. John Harsanyi (1967-68) showed how to handle it: add an initial chance move in which "Nature" draws each player's type from a common prior, and let each player observe only its own type. The game becomes one of imperfect information (a Bayesian game) that can be analyzed with standard tools, leading to the Bayes-Nash equilibrium. The same idea underlies type-based reasoning in multi-agent RL: ad hoc teamwork with unknown partners, opponent modeling, and the interactive POMDPs below, where agents keep beliefs over the possible types and models of other agents.
Beyond Teams: POSGs and I-POMDPs
When agents have their own rewards, the setting becomes a partially observable stochastic game (POSG). Interactive POMDPs (Gmytrasiewicz & Doshi - JAIR 2005) give each agent beliefs over both the physical state and the models of other agents -- including their beliefs about it -- truncated at some nesting depth. The AAMAS 2016 paper linked at the top of this page studies model-free learning with PAC guarantees in such settings.
Multi-Agent Benchmarks
- SMACv2 | B. Ellis et al. - NeurIPS 2023 -- revised the StarCraft Multi-Agent Challenge after finding that many original SMAC scenarios could be won by open-loop policies that ignore observations entirely; adds randomized unit types and start positions so agents must actually use what they observe.
- The Hanabi Challenge | N. Bard et al. - Artificial Intelligence 2020 -- a cooperative card game where you see everyone's cards but your own and communication is strictly limited; a benchmark for theory of mind.
- Melting Pot | J. Leibo et al. - ICML 2021 -- test-time generalization to unfamiliar co-players in mixed-motive social scenarios.
- JaxMARL | A. Rutherford et al. - NeurIPS 2024 -- GPU-accelerated multi-agent environments and baselines.
Imperfect-Information Games
Games with hidden information have been the field's public scoreboard for partial observability. Unlike chess or Go (AlphaGo Zero), a player cannot simply search forward from "the" position, because it does not know the position; it knows only an information set of positions consistent with what it has seen. See Game Theory and Gaming.
| Year | System | Game | What hid the state | Key idea |
|---|---|---|---|---|
| 2017 | DeepStack / Libratus | Heads-up no-limit Texas hold'em | Opponent's cards | Counterfactual regret minimization (CFR), depth-limited re-solving during play |
| 2019 | Pluribus | Six-player no-limit hold'em | Five opponents' cards | Blueprint strategy from self-play plus limited-lookahead search |
| 2019 | AlphaStar | StarCraft II | Fog of war, scouting | LSTM/transformer-based agents trained with league self-play; Grandmaster level |
| 2020 | ReBeL | Poker, Liar's Dice | Private hands | Public belief states make imperfect-information games searchable like perfect-information ones |
| 2022 | DeepNash | Stratego | Identity of 40 hidden pieces | Model-free regularized Nash dynamics; reached the all-time top three on the Gravon platform |
| 2022 | Cicero | Diplomacy | Other players' intentions | Language model plus strategic planning; human-level negotiation |
| 2023 | Student of Games | Chess, Go, poker, Scotland Yard | Varies | One algorithm for both perfect- and imperfect-information games |
| 2025-26 | Ataraxos | Stratego (and Barrage Stratego, Hanabi, Dou Dizhu) | Hidden pieces and cards | Self-play policy-value network + a separate belief network predicting opponents' hidden pieces + test-time search; published in Nature on 30 September 2026 as the first superhuman Stratego AI |
The 2026 Ataraxos result is a useful marker of where the field stands. It beat Pim Niemeijer, described as the most decorated Stratego player in history, 15 wins to 1 with 4 draws, and was trained for under about $8,000 -- roughly 1/500th of the compute of DeepMind's DeepNash. The same recipe also set a new state of the art in cooperative Hanabi across two- to five-player variants. The broader lesson: explicitly modeling hidden information with a learned belief network, then searching over sampled hidden states at decision time, now scales to games with astronomically many hidden configurations (more than 1033 possible Stratego setups).
Partial Observability in LLM Agents
Every agentic LLM system is a POMDP agent, whether or not its designers say so:
- the user's true goal and preferences are hidden and must be inferred from underspecified requests;
- a web or computer-use agent sees one screen at a time; content below the fold, behind a login, or in another tab is unobserved;
- a coding agent cannot read a whole codebase at once and must decide what to open, search, or run;
- tool results are partial, stale, or wrong; the world can change between steps;
- the context window is finite, so the agent must decide what to remember and what to forget -- see Context and Memory.
Typical Failure Modes
- Premature commitment -- answering or acting before resolving ambiguity. Benchmarks such as InfoQuest (2025) found that current assistants need many turns to infer hidden user context and are inefficient at asking.
- Belief inertia -- failing to overwrite outdated beliefs when the world changes. Theory of Space (2026) found foundation-model agents explore redundantly and struggle to revise spatial beliefs after objects are moved, especially in vision-based settings; Seeing Isn't Believing (2026) targets the same failure in embodied agents.
- Collapsed uncertainty -- storing guesses as facts in memory or summaries, which then reinforce themselves.
- Information self-locking -- On Information Self-Locking in RL for Active Reasoning of LLM Agents (2026) reports that outcome-based RL can train agents that stop asking informative questions and stop using the evidence they do get.
- Inconsistent beliefs -- two histories that imply the same posterior can yield different actions from an LLM because their surface text differs.
Design Patterns for Belief Tracking
| Pattern | How it works | Examples |
|---|---|---|
| Full history in context | Append every action and observation; let the model infer the state implicitly | ReAct-style agents (Yao et al. 2022) -- simple, but context grows without bound and errors hide in long transcripts |
| Natural-language belief bottleneck | At each step, update a concise written belief, then act on the belief alone | ABBEL (Lidayan et al. 2025) -- near-constant memory and interpretable beliefs; RL with belief grading narrows the gap to full-context agents (BAIR blog, July 2026) |
| Uncertainty-aware verbalized beliefs | Annotate each belief claim with a likelihood phrase ("likely", "almost certainly not"); train the belief model and the policy jointly with RL | Agent-BRACE (May 2026) |
| Probabilistic agent memory | Store attribute-level distributions instead of deterministic conclusions, so early guesses do not harden into "facts" | BeliefMem (2026) |
| External Bayesian belief engine | A separate inference module maintains a POMDP-grounded belief; the LLM only chooses actions given that belief, which makes it a Markov policy on the belief MDP | Belief-State Engine (Sept 2026) -- compared against QMDP and POMCP on the Tiger problem and a red-team attack-graph task |
| Belief-based world models | World models that expose what is known and unknown, not just simulated futures | Belief-Based World Models (2026) on ALFWorld and ScienceWorld |
| Structured / neuro-symbolic beliefs | Keep beliefs in a knowledge graph; combine fast reactive and slow deliberative reasoning | NeSyFS (2026) |
| Active information seeking & clarification | Decide explicitly when to gather information or ask the user versus act | InfoSeeker (2025), ClarifyBench (2025), RegretBench (2026), uncertainty decomposition for clarification (2026) |
Benchmarks That Stress Hidden Information
- ARC-AGI-3 | ARC Prize Foundation - 2026 -- launched 25 March 2026: interactive, turn-based environments with no instructions, where agents must explore, infer the goal, model the dynamics, and plan. At launch every frontier model tested scored below 1% while humans solved all environments. The ARC Prize 2026 competition on it takes final submissions in early November 2026.
- τ-bench (2024) -- tool-using agents converse with a simulated user who holds information the agent must elicit.
- ALFWorld, ScienceWorld, TextWorld -- text-based embodied tasks with unobserved object locations.
- WebArena and OSWorld -- web and desktop agents operating with partial views of complex interfaces.
- InfoQuest and RegretBench -- multi-turn dialogue with hidden user context or intent.
Practical Guidance for Agent Builders
- Keep an explicit state. Maintain a running notes or state file separate from the raw transcript, and update it every step.
- Record uncertainty, not just conclusions. Write "probably X (unverified)" rather than "X." Distinguish what was observed from what was inferred.
- Separate belief update from action selection. First ask "what do I now believe?", then "what should I do?"
- Price information. Ask or explore when a wrong guess is expensive or irreversible; act directly when it is cheap to redo.
- Verify before irreversible actions, and re-read the world after acting -- it may have changed.
- Test with twins. Evaluate the same task with and without the hidden information to see how much performance is lost to partial observability.
A Safety Angle: Partial Observability of the Overseer
Partial observability affects the humans who train and supervise AI as well. When Your AIs Deceive You: Challenges of Partial Observability in RLHF | L. Lang et al. - NeurIPS 2024 showed that when evaluators see only part of what an agent did, RLHF can reward "deceptive inflation" (making results look better than they are) and "overjustification" (spending effort to look good rather than be good). Better oversight tools are, in part, tools for reducing the overseer's partial observability. See Explainable / Interpretable AI.
Applications
| Domain | What is hidden | Typical approaches |
|---|---|---|
| Robotics and manipulation | Object poses behind occlusion, contact state, the robot's own position | Particle filters and SLAM, belief-space planning (SARSOP, POMCPOW), recurrent policies, teacher-student sim-to-real |
| Autonomous driving and drones | Occluded road users, other drivers' intentions | Occlusion-aware and intention-aware POMDP planning, online tree search in continuous spaces |
| Dialogue and assistants (Conversational AI) | User goals, preferences, context | Statistical POMDP dialogue managers; today, LLM belief tracking and clarification policies |
| Web, desktop, and coding agents (Agents/Assistants) | Off-screen content, backend state, unread files | Exploration, explicit notes and memory, verification steps |
| Strategy games and negotiation (Negotiation) | Opponents' cards, units, plans, preferences | CFR, public belief states, belief networks + search, opponent modeling |
| Healthcare | Patient's underlying disease state | POMDP treatment and testing policies that weigh diagnostic tests against treatment |
| Earth science and sustainability | Subsurface ore bodies, CO2 plumes, groundwater contamination | Large-scale information-gathering POMDPs solved with DESPOT, POMCPOW, BetaZero |
| Cybersecurity and Defense | Attacker presence, location, and capabilities | POMDP-based intrusion response and attack-graph planning; partially observable stochastic games |
| Multi-robot, warehouse, and traffic control | Teammates' observations, global state | Dec-POMDP formulations, CTDE multi-agent RL, learned communication |
| Recommendation | A user's latent and shifting interests | RL and bandits with latent user state |
Benchmarks and Environments
| Benchmark | Year | Setting | What it tests |
|---|---|---|---|
| Flickering Atari (DRQN paper) | 2015 | Single-agent, pixels | Integrating information across randomly blanked frames |
| Hanabi Learning Environment | 2019 | Cooperative card game | Theory of mind, implicit communication |
| Memory Maze | 2022 | 3D navigation | Long-term spatial memory |
| POPGym | 2023 | 15 low-dimensional POMDPs | Broad, fast comparison of memory models |
| SMACv2 | 2023 | StarCraft II multi-agent | Decentralized control under stochasticity and partial observability |
| Craftax | 2024 | Open-ended survival (JAX) | Exploration and memory in a partially observed world |
| POPGym Arcade | 2025 | Pixel games with MDP/POMDP twins | Controlled studies of memory and observability |
| MIKASA (Base and Robo) | 2025 | Memory RL; robotic manipulation | Taxonomy-driven memory evaluation |
| InfoQuest | 2025 | Multi-turn chat | Uncovering hidden user context through questions |
| Theory of Space | 2026 | Embodied foundation-model agents | Active exploration and spatial belief revision |
| ARC-AGI-3 | 2026 | Interactive abstract games | Exploration, goal inference, and model building with no instructions |
Software and Tools
| Tool | Language | What it offers |
|---|---|---|
| POMDPs.jl | Julia | A modeling interface plus a large ecosystem of solvers (SARSOP, POMCP, POMCPOW, DESPOT, QMDP, and more) and belief updaters |
| pomdp_py | Python | Python framework for building and solving POMDPs, with POMCP and value iteration |
| SARSOP and DESPOT | C++ | Reference implementations of the leading offline point-based solver and online tree-search solver |
| AI-Toolbox | C++ / Python | MDP, POMDP, and multi-agent planning and learning algorithms |
| BetaZero.jl | Julia | Learned belief-space planning |
| Stable-Baselines3 Contrib | Python / PyTorch | RecurrentPPO (LSTM policies) for partially observable environments |
| CleanRL | Python | Single-file implementations, including PPO with LSTM |
| POPGym | Python | Partially observable environments and memory-model baselines |
| DreamerV3 | Python / JAX | World-model RL agent |
| PettingZoo | Python | Standard API and environments for multi-agent RL |
| EPyMARL and JaxMARL | Python / JAX | Cooperative MARL algorithms (QMIX, MAPPO, and others) and benchmarks |
| OpenSpiel | C++ / Python | Imperfect-information games and algorithms such as CFR |
Choosing an Approach
| If your situation is... | Start with... |
|---|---|
| Small, discrete, model known | Exact or point-based solvers (SARSOP) |
| Large or continuous, simulator available | Online belief-space search (POMCP, DESPOT, POMCPOW); add learned guidance (BetaZero) for long horizons |
| No model, need sample efficiency | A learned world model (DreamerV3) |
| No model, short-to-medium memory | Recurrent PPO or R2D2-style recurrent value learning; benchmark on POPGym |
| Very long-range recall | Transformer or SSM memory, tested against recurrent baselines; check whether you need retention or revision |
| Simulator exposes the true state during training | Asymmetric actor-critic or teacher-student, watching for the imitation gap |
| Several cooperating agents | Model as a Dec-POMDP; CTDE methods (MAPPO, QMIX), plus learned communication if allowed |
| Adversaries hiding information | CFR or public-belief-state methods; belief network + test-time search |
| An LLM agent | Explicit, uncertainty-aware belief state; clarification and information-seeking policy; verification before irreversible actions |
Open Problems and Frontiers
- Calibrated beliefs in open-ended worlds. Classical belief tracking assumes a known, finite state space. Text and web environments have neither, and LLM agents still struggle to maintain and revise beliefs over long horizons.
- Memory that can be revised, not just retained. Recent benchmarks show that remembering is easier than updating; architectures that do both reliably remain open.
- Knowing what you do not know. Deciding when to explore, ask, or act -- and training agents (with RL or otherwise) that do not lose the habit of information seeking.
- Decentralized LLM agents. Multi-agent LLM systems inherit Dec-POMDP difficulty: what to share, what teammates know, and how to coordinate without a shared view.
- Theory for privileged and hindsight information. When does training-time access to the true state provably help, and how can agents be trained to look for what the teacher simply knew?
- Evaluation hygiene. Controlling for observability (MDP/POMDP twins) so that memory, credit assignment, exploration, and generalization failures are not confused with one another.
- Oversight under partial observability. Making sure human and automated evaluators can see enough of an agent's behavior to reward what is actually good.
Timeline
| Year | Milestone |
|---|---|
| 1965 | Åström formulates optimal control with incomplete state information |
| 1973 | Smallwood & Sondik: POMDP value functions are piecewise-linear and convex |
| 1987 | Papadimitriou & Tsitsiklis: finite-horizon POMDPs are PSPACE-complete |
| 1998 | Kaelbling, Littman & Cassandra bring POMDPs into mainstream AI |
| 1999 | Madani, Hanks & Condon: infinite-horizon POMDP planning is undecidable |
| 2002 | Bernstein et al.: Dec-POMDPs are NEXP-complete |
| 2003-2008 | Point-based solvers: PBVI, Perseus, HSVI, SARSOP |
| 2010 | POMCP brings Monte Carlo tree search to large POMDPs |
| 2015 | DRQN: deep recurrent Q-learning for POMDPs |
| 2016-2017 | DDRQN, DIAL, Dec-HDRQN; DeepStack and Libratus beat poker professionals |
| 2018 | QMIX; asymmetric actor-critic; POMCPOW for continuous POMDPs |
| 2019 | R2D2; Pluribus; AlphaStar; the Hanabi challenge |
| 2020 | GTrXL; ReBeL and public belief states |
| 2021-2022 | Epistemic POMDPs; MAPPO; DeepNash and Cicero; recurrent baselines shown to be strong |
| 2023 | POPGym; SMACv2; Student of Games; "When do transformers shine in RL?" |
| 2024 | BetaZero; RLHF under partial observability; evidence that transformers represent belief states internally |
| 2025 | DreamerV3 published in Nature (April); Belief State Transformer; POPGym Arcade; MIKASA; ABBEL and the first wave of LLM belief-state methods |
| 2026 | ARC-AGI-3 launches (March); Agent-BRACE, BeliefMem, Belief-State Engine, and belief-based world models for LLM agents; Ataraxos, the first superhuman Stratego AI, published in Nature (30 September) |
Videos and Lectures
- Partially Observable MDPs -- see the video in the POMDP section above
- Multi-agent partial observability -- see the video in the Deep Recurrent and Distributed Q-Networks section above
- Lecture 15: Partially Observable MDPs | P. Abbeel - UC Berkeley CS287 Advanced Robotics
- POMDPs: Decision Making Under Uncertainty with POMDPs.jl
- Online Algorithms for POMDPs with Continuous State, Action, and Observation Spaces | Z. Sunberg & M. Kochenderfer - ICAPS 2018
- Why Generalization in RL Is Difficult: Epistemic POMDPs | D. Ghosh et al.
- Partially Observable Markov Decision Processes | P. Poupart - DLRL Summer School
Further Reading
- Planning and Acting in Partially Observable Stochastic Domains | L. Kaelbling, M. Littman & A. Cassandra - Artificial Intelligence 1998
- Algorithms for Decision Making | M. Kochenderfer, T. Wheeler & K. Wray - MIT Press 2022 ... free PDF
- Reinforcement Learning: An Introduction (2nd ed.) | R. Sutton & A. Barto - section 17.3, "Observations and State"
- A Concise Introduction to Decentralized POMDPs | F. Oliehoek & C. Amato - Springer 2016
- Online Planning Algorithms for POMDPs | S. Ross, J. Pineau, S. Paquet & B. Chaib-draa - JAIR 2008
- Partially Observable Markov Decision Processes and Robotics | H. Kurniawati - 2022
- Partially Observable Markov Decision Processes in Robotics: A Survey | M. Lauri, D. Hsu & J. Pajarinen - 2023
- (A Partial Survey of) Decentralized, Cooperative Multi-Agent Reinforcement Learning | C. Amato - 2024
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPs | T. Ni, B. Eysenbach & R. Salakhutdinov - ICML 2022
- ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language | A. Lidayan et al. - 2025
- Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability | A. Chattopadhayay & D. Halder - 2026