← All papers

TMLR 2025/CO-CQL

Leveraging Fully-Observable Solutions for Improved Partially-Observable Offline Reinforcement Learning

Chulabhaya Wijesundara·Andrea Baisero·Gregory Castañón·Alan Carlin·Robert Platt·Christopher Amato

The short version

Partially observable offline RL must learn a history representation and a control policy from a fixed dataset. CO-CQL uses state information available during training to obtain guidance from a fully observable expert, combining recurrent Conservative Q-Learning with behavior cloning. The deployed policy still uses only observation history.

01 / Motivation

Offline learning is harder with partial observations.

Offline RL is already constrained by the coverage of a fixed dataset. Under partial observability, the agent must also learn a history representation that retains information needed for future decisions. Noisy or incomplete observations therefore complicate both policy learning and value estimation, without new environment interactions to resolve the uncertainty.

Recorded experience may include full environment states alongside the agent’s observations, especially in simulation or when data collection uses additional sensors. CO-CQL studies how this privileged training information can improve partially observable offline learning while preserving a history-based policy at deployment.

In HeavenHell, the partially observable agent must visit an information source to identify the correct exit. A fully observable expert can go there directly.
In HeavenHell, the partially observable policy must visit the oracle to identify the correct exit. A full-state expert can skip that step. Pure expert imitation therefore omits behavior needed for information gathering. Original figure (PDF)

02 / Method

Use state information to guide offline learning.

CO-CQL assumes a dataset containing observation-action histories and corresponding environment states, together with a fully observable expert. The recorded states are used to query the expert; its action distribution provides a behavior-cloning target for the history-based policy.

The algorithm combines this auxiliary imitation objective with recurrent Conservative Q-Learning (CQL). The policy and Q-functions condition on history through recurrent networks; CQL penalizes optimistic value estimates for actions poorly supported by the dataset. Full state is used for expert supervision during training, while execution requires only observation history.

CO-CQL learns from an offline dataset and a full-state expert during training. After training, its policy uses only observations and memory to choose actions.
CO-CQL uses asymmetric offline training: full states provide inputs to the expert, while the learned policy and Q-functions condition on observation-action history. At deployment, only the history-based policy is needed.

03 / Contributions

What this work adds.

  1. 01

    Cross-observability analysis (COOR)

    The Cross-Observability Optimality Ratio quantifies how much of the optimal full-state action set is also optimal given partial information. This characterizes when expert guidance is relevant to the partially observable task.

  2. 02

    An empirical approximation (COAR)

    The Cross-Observability Approximation Ratio estimates this overlap using well-trained full-state and history-based policies, connecting the analysis to the observed benefit of expert guidance.

  3. 03

    CO-CQL: conservative learning with expert imitation

    CO-CQL trains a history-based actor and critic with CQL and an auxiliary behavior-cloning objective from a full-state expert.

04 / Evidence

Four kinds of partial observability.

Experiments cover observation noise, hidden state variables, information gathering and memory, and restricted fields of view. CO-CQL improves partially observable offline learning across these challenges, with the benefit of full-state supervision varying by task.

Hard HeavenHell highlights the limits of expert imitation. The full-state expert reaches the correct exit without visiting the oracle, so its behavior alone cannot teach an agent how to gather the missing information. CO-CQL combines that guidance with value learning to learn both oracle visits and goal-reaching behavior, outperforming BC and standard CQL.

Information gathering during navigation

State-visitation densities over 100 Hard HeavenHell episodes in which the good exit is on the left. Arrows summarize the learned behavior; percentages report occupancy at each location. The comparison separates goal-directed expert behavior, information gathering under CQL, and their combination under CO-CQL.

Three Hard HeavenHell visitation maps compare a fully observable expert that skips the oracle, CQL that stays near the oracle, and CO-CQL that visits the oracle and reaches the correct exit.
Left: the fully observable expert bypasses the oracle, illustrating the information-gathering behavior absent from BC supervision. Center: CQL visits the oracle but rarely reaches the correct exit. Right: CO-CQL learns both stages. The oracle is labeled “priest” in the figure. Original figure (PDF)
Scope and limitations

CO-CQL uses a fixed, task-specific weight for expert imitation. Guidance is less suitable when optimal partially observable behavior requires actions the full-state expert can omit, such as information gathering. COOR and COAR analyze this mismatch; they are not used to gate expert supervision at each step.

Read the full paper

05 / Citation

Cite this paper.

BibTeX
Download .bib
@article{wijesundara2025leveraging,
  title = {Leveraging Fully-Observable Solutions for Improved Partially-Observable Offline Reinforcement Learning},
  author = {Chulabhaya Wijesundara and Andrea Baisero and Gregory Castañón and Alan Carlin and Robert Platt and Christopher Amato},
  journal = {Transactions on Machine Learning Research},
  year = {2025},
  url = {https://openreview.net/forum?id=e9p4TDPy6A}
}