Reinforcement learning has long promised machines that can teach themselves complex skills, from locomotion to robotic manipulation, but the standard recipe requires millions of trial-and-error interactions with the environment. In the real world, such online experimentation is often too expensive, too slow, or simply too dangerous. Offline reinforcement learning offers an appealing alternative: train a policy entirely from previously collected datasets, without ever letting the agent act in the live environment. The catch is a stubborn mathematical pathology known as distribution shift, and a new study published in the journal Machine Learning proposes a fresh way to attack it using ideas borrowed from contrastive representation learning.
The study, conducted by Haotian Zhang, Chen Wang, and Jun Moon of the Department of Electrical Engineering at Hanyang University in Seoul, introduces TACCO, short for TD3+BC with Actor-Critic Contrastive Optimization. The method builds on TD3+BC, a minimalist and widely used offline reinforcement learning algorithm, and augments it with supervised contrastive learning branches that explicitly model out-of-distribution actions. The authors define out-of-distribution actions as those that lack local data support in a given state, do not match the current state, or exhibit low reliability. Rather than treating these problematic actions as an afterthought, TACCO makes them a first-class target of representation learning, reshaping the internal embeddings of the agent so that unreliable actions become visibly distinguishable from trustworthy ones.
To understand why this matters, it helps to see how offline reinforcement learning goes wrong. When an agent learns only from a fixed dataset, its value function, the critic that estimates the long-term reward of taking an action in a state, is only trustworthy for state-action pairs that actually appear in the data. During policy improvement, the actor proposes new actions that may drift outside the support of the dataset. The critic, having never seen these pairs, tends to overestimate their value, a phenomenon known as value overestimation. The actor then chases these inflated estimates, proposing ever more exotic actions, and the whole learning process spirals into degradation. This feedback loop between an overconfident critic and an adventurous actor is the central failure mode of offline reinforcement learning.
Existing approaches generally respond in one of two ways. Policy-constraint methods, such as TD3+BC itself, tether the learned policy to the behavior policy recorded in the data, penalizing any action that strays too far from what the dataset contains. Conservative value regularization methods, exemplified by CQL, pessimistically lower the estimated values of out-of-distribution actions so the critic never overhypes them. Both families work, but as the Hanyang team points out, few methods explicitly model the problematic actions from a representation learning perspective. That is the gap TACCO fills: instead of only constraining behavior or deflating values, it teaches the network’s internal representations to separate reliable actions from unreliable ones, providing a structural safeguard that complements the existing toolkit.
TACCO comes in two variants, each targeting a different side of the actor-critic architecture and a different reward regime. The first, TACCO-A, is designed for dense reward environments, where feedback signal is abundant. It attaches a supervised contrastive learning branch to the actor. The actor consists of a shared state-encoding backbone, a standard policy output head, and an additional embedding head that concatenates the hidden state representation with the proposed action and projects the result into a 64-dimensional embedding. Contrastive learning then operates on three kinds of actions: data-driven actions drawn from the offline dataset, random actions, and cross-state policy actions. The training objective pulls embeddings of dataset-supported actions together while pushing embeddings of random and mismatched actions apart, so the actor’s representation space develops a clear geometry in which out-of-distribution actions stand out and can be suppressed.
The second variant, TACCO-B, is tailored to sparse reward environments, where rewards arrive rarely and unreliable value estimates are especially corrosive. Here the contrastive branch moves to the critic. TACCO-B filters training samples on the critic side based on two criteria: high Q-values and low uncertainty. Samples that pass both filters are treated as reliable, while the remainder are flagged as suspect. The method then applies contrastive constraints to both reliable and remaining samples simultaneously, encouraging the critic’s shared state-action encoder, a small network ending in a LayerNorm-normalized 64-dimensional embedding, to learn discriminative representations that reflect this reliability structure. By shaping the critic’s representation space rather than merely clipping its outputs, TACCO-B aims to stabilize value learning even when the reward signal is thin and misleading bootstrapped targets are easy to come by.
The architectural details are deliberately modest, which is part of the method’s appeal. Both variants retain the dual Q-network structure of TD3, with two independent critics that help mitigate function approximation error, and both keep the behavioral cloning term of TD3+BC that anchors the policy to the dataset. The contrastive modules are small multilayer perceptrons, and the authors report that performance is relatively robust to the weighting of the contrastive loss. Sensitivity analyses reveal a more interesting story for the other knobs: in TACCO-A, relaxing the behavioral cloning weight makes the policy more adventurous but destabilizes training, and increasing the contrastive weight steadily compensates for that instability. In TACCO-B, the fraction of high-value samples used for filtering must be tuned carefully, since too few samples starve the model of coverage while too many reintroduce the very low-quality actions the filter was meant to exclude.
The empirical evaluation is broad. The team compared TACCO against six state-of-the-art offline reinforcement learning algorithms, including CQL, TD3+BC, the Decision Transformer, IQL, DTQL, and O-DICE, across 27 tasks spanning the MuJoCo locomotion suite and the Adroit dexterous manipulation suite within the D4RL benchmark. All methods were trained under an identical protocol: one million gradient steps, evaluation every five thousand steps, a batch size of 256, and results averaged over five random seeds with 95 percent confidence intervals, using the standard D4RL score normalization. According to the authors, TACCO achieved superior performance on most of these tasks, which they interpret as validation that mitigating out-of-distribution action interference through contrastive optimization enhances policy stability, in both dense and sparse reward settings.
The broader significance of the work lies in its reframing of a classic problem. Distribution shift has usually been handled as a constraint or a penalty, a blunt instrument applied to actions or values. TACCO instead treats it as a representation problem: if the network’s embeddings make the boundary between supported and unsupported actions explicit, then both the actor and the critic can exploit that structure during training. This connects offline reinforcement learning to the wider wave of contrastive learning that has transformed vision and speech processing, and to a growing line of research that uses contrastive objectives for goal-conditioned control and planning. It also suggests a practical path forward, since the contrastive branches add little architectural complexity on top of a minimalist baseline that many practitioners already deploy.
Challenges remain before such methods reach real-world robotics and industrial control, where offline data is often suboptimal, heterogeneous, and collected under shifting conditions. The authors note that code will be made available on request, and the work was supported by Korean government research programs through KETEP, MOTIE, and IITP. As offline reinforcement learning matures into a serious option for settings where online exploration is prohibitive, techniques like TACCO point toward a hybrid future: conservative enough to trust the data, expressive enough to improve on it, and, crucially, aware of exactly where its own knowledge ends.
Subject of Research: Contrastive actor-critic optimization for mitigating out-of-distribution actions in offline reinforcement learning
Article Title: TACCO: Actor–Critic Contrastive Optimization for OOD Action Mitigation in Offline Reinforcement Learning
Article References: Zhang, H., Wang, C., & Moon, J. (2026). TACCO: Actor–Critic Contrastive Optimization for OOD Action Mitigation in Offline Reinforcement Learning. Machine Learning, 115(10), Article 241. https://doi.org/10.1007/s10994-026-07170-3
Image Credits: AI Generated
DOI: 10.1007/s10994-026-07170-3
Keywords: offline reinforcement learning, contrastive learning, out-of-distribution actions, TD3+BC, actor-critic, D4RL benchmark, value overestimation, distribution shift, representation learning, MuJoCo, Adroit, machine learning
Cite Scienmag News
Denise Maddox. (October 8, 2026). Contrastive Learning Tames Out-of-Distribution Actions in Offline Reinforcement Learning. Scienmag. https://scienmag.com/contrastive-learning-tames-out-of-distribution-actions-in-offline-reinforcement-learning/
Denise Maddox. "Contrastive Learning Tames Out-of-Distribution Actions in Offline Reinforcement Learning." Scienmag, 8 October 2026, https://scienmag.com/contrastive-learning-tames-out-of-distribution-actions-in-offline-reinforcement-learning/. Accessed 8 October 2026.
Denise Maddox. "Contrastive Learning Tames Out-of-Distribution Actions in Offline Reinforcement Learning." Scienmag. October 8, 2026. https://scienmag.com/contrastive-learning-tames-out-of-distribution-actions-in-offline-reinforcement-learning/

