Residual Reward Models:
Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics

Chenyang Cao1  Miguel Rogel-García1  Mohamed Nabail1  Xueqian Wang2 Nicholas Rhinehart1

1University of Toronto  2Tsinghua University 

Abstract

Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback, which limits its real-world applicability. Prior work has proposed learning a reward model from demonstrations and fine-tuning it using preferences. However, when the model is a neural network, transitioning between different loss functions across training phases often leads to unstable optimization and performance degradation. In this paper, we propose a method to effectively leverage prior knowledge with a Residual Reward Model (RRM). An RRM assumes that the true reward of the environment can be split into a sum of two parts: a prior reward and a learned reward. The prior reward is a term available before training, such as an engineering heuristic ``best guess'', a language-generated reward, or a reward function learned from inverse reinforcement learning, and the learned reward is then trained with preferences as a residual offset. Experimental results in Meta-World and DM-Control show that RRMs substantially improve the sample efficiency of common PbRL methods across various prior reward types. Furthermore, we demonstrate the practical efficacy of our method through sim-to-real transfer on a physical Franka Panda robot, accelerating policy learning and achieving high success rates in fewer steps than baselines.

Method Overview

An agent interacts with the reward-free environment and generates trajectories. In order to generate rewards for reinforcement learning, our method assumes access to a "prior" reward that conveys some information about the task, but generally may be different than the true task reward function. These prior rewards form part of a reward function that is trained to be aligned with preference pairs.

An additional version of residual reward model in image-based setting: Residual Reward Model obtains images and proprioceptive states from the environment rather than states. An encoder is used for extracting representations from images and is jointly trained with the RL agent.

Real Tasks Visualization

We evaluate Residual Reward Models on a real Franka arm to finish tasks. Videos are recorded by iPhone15pro.


Pick and Reach
Push

Simulation Tasks Visualization

We show the learned curve of Residual Reward Models and PEBBLE for each simulation task.


State-based Tasks

MetaWorld Button-Press
MetaWorld Sweep-Into
MetaWorld Door-Open
MetaWorld Door-Unlock
DMControl Quadruped-Walk
DMControl Walker-Walk
Manipulator Learning Reach
Manipulator Learning Push
Manipulator Learning Pick-and-Reach

1
Image-based Tasks

MetaWorld Visual Button-Press
MetaWorld Visual Sweep-Into

Main Results

We report the IQM results for our method and baselines on 6 gripper manipulation tasks from Meta-World and 2 locomotion tasks from DM-Control. Our method achieves prominent performance on all tasks, and significantly outperforms the baseline PEBBLE. Among them, the proxy reward represents the negative distance between the task object and the task goal, and the negative distance between the gripper and the task object. In the image-based setting, the proxy reward applies a punishment when the gripper moves outside the predefined region.

Applying RRM to other PbRL baselines and less feedback

Residual Reward Model can be applied to other PbRL baselines directly while requiring less feedback and still demonstrates excellent performance that surpasses baselines.

Citation

If you find this project helpful, please cite us:

@article{cao2025residual, title={Residual reward models for preference-based reinforcement learning}, author={Cao, Chenyang and Rogel-Garc{\'\i}a, Miguel and Nabail, Mohamed and Wang, Xueqian and Rhinehart, Nicholas}, journal={arXiv preprint arXiv:2507.00611}, year={2025} }