TMLR 2025/CO-CQL
Leveraging Fully-Observable Solutions for Improved Partially-Observable Offline Reinforcement Learning
Partially observable offline RL must learn a history representation and a control policy from a fixed dataset. CO-CQL uses state information available during training to obtain guidance from a fully observable expert, combining recurrent Conservative Q-Learning with behavior cloning. The deployed policy still uses only observation history.
01 / Motivation
Offline learning is harder with partial observations.
Offline RL is already constrained by the coverage of a fixed dataset. Under partial observability, the agent must also learn a history representation that retains information needed for future decisions. Noisy or incomplete observations therefore complicate both policy learning and value estimation, without new environment interactions to resolve the uncertainty.
Recorded experience may include full environment states alongside the agent’s observations, especially in simulation or when data collection uses additional sensors. CO-CQL studies how this privileged training information can improve partially observable offline learning while preserving a history-based policy at deployment.
02 / Method
Use state information to guide offline learning.
CO-CQL assumes a dataset containing observation-action histories and corresponding environment states, together with a fully observable expert. The recorded states are used to query the expert; its action distribution provides a behavior-cloning target for the history-based policy.
The algorithm combines this auxiliary imitation objective with recurrent Conservative Q-Learning (CQL). The policy and Q-functions condition on history through recurrent networks; CQL penalizes optimistic value estimates for actions poorly supported by the dataset. Full state is used for expert supervision during training, while execution requires only observation history.
03 / Contributions
What this work adds.
- 01
Cross-observability analysis (COOR)
The Cross-Observability Optimality Ratio quantifies how much of the optimal full-state action set is also optimal given partial information. This characterizes when expert guidance is relevant to the partially observable task.
- 02
An empirical approximation (COAR)
The Cross-Observability Approximation Ratio estimates this overlap using well-trained full-state and history-based policies, connecting the analysis to the observed benefit of expert guidance.
- 03
CO-CQL: conservative learning with expert imitation
CO-CQL trains a history-based actor and critic with CQL and an auxiliary behavior-cloning objective from a full-state expert.
04 / Evidence
Four kinds of partial observability.
Experiments cover observation noise, hidden state variables, information gathering and memory, and restricted fields of view. CO-CQL improves partially observable offline learning across these challenges, with the benefit of full-state supervision varying by task.
Hard HeavenHell highlights the limits of expert imitation. The full-state expert reaches the correct exit without visiting the oracle, so its behavior alone cannot teach an agent how to gather the missing information. CO-CQL combines that guidance with value learning to learn both oracle visits and goal-reaching behavior, outperforming BC and standard CQL.
Information gathering during navigation
State-visitation densities over 100 Hard HeavenHell episodes in which the good exit is on the left. Arrows summarize the learned behavior; percentages report occupancy at each location. The comparison separates goal-directed expert behavior, information gathering under CQL, and their combination under CO-CQL.
Each panel compares methods under a different observation restriction. Curves average five independent runs; shaded bands show standard errors.
Scope and limitations
CO-CQL uses a fixed, task-specific weight for expert imitation. Guidance is less suitable when optimal partially observable behavior requires actions the full-state expert can omit, such as information gathering. COOR and COAR analyze this mismatch; they are not used to gate expert supervision at each step.
05 / Citation
Cite this paper.
@article{wijesundara2025leveraging,
title = {Leveraging Fully-Observable Solutions for Improved Partially-Observable Offline Reinforcement Learning},
author = {Chulabhaya Wijesundara and Andrea Baisero and Gregory Castañón and Alan Carlin and Robert Platt and Christopher Amato},
journal = {Transactions on Machine Learning Research},
year = {2025},
url = {https://openreview.net/forum?id=e9p4TDPy6A}
}