BOSS: A Benchmark for Observation Space Shift in Long-Horizon Task

Chaining skills breaks them — and the culprit is a change that shouldn't matter.

Yue Yang 1,†
UNC Chapel Hill
Northeastern University
UNC Chapel Hill
UNC Chapel Hill
UNC Chapel Hill
IEEE RA-L 2025 · ICRA 2026
1The University of North Carolina at Chapel Hill, 2Northeastern University, †Corresponding author: yygx@cs.unc.edu
A robot places a potato into a bowl, then fails to move that bowl into a cabinet because the policy was only trained on an empty bowl.
Observation Space Shift (OSS). Chaining PlaceObject(potato, bowl) → MoveContainer(bowl, cabinet): the potato left in the bowl shifts the observation space of the second skill, which was only trained on an empty bowl — so it fails.

Abstract

Hierarchical robots solve long-horizon tasks by chaining individually trained visuomotor skills. We identify Observation Space Shift (OSS): a preceding skill changes parts of the scene that are irrelevant to the next skill’s logic yet still visible in its camera input, pushing the next policy out of its training distribution. We introduce BOSS, a benchmark built on LIBERO with three progressive challenges, and show that OSS is common and severe across four imitation-learning baselines — and that three intuitive mitigations all fail to resolve it.

up to 67%
average performance drop from a single irrelevant change

Challenge 1

76%
of tasks hit by OSS once three modifications accumulate

Challenge 2

0 / 3
intuitive mitigations sufficient to resolve OSS

augmentation, frozen encoders, 3D policies

1,727
modified task variants generated automatically

Rule-based Automatic Modification Generator

What is Observation Space Shift?

Skill chaining — sequentially executing pre-learned visuomotor skills — is a simple yet powerful way to tackle long-horizon robot tasks. But failures often arise at skill transitions: the terminal state of one skill can fall outside the initial-state distribution the next skill was trained on. Crucially, the culprit is frequently a change that is logically irrelevant to the next skill, yet still appears in its visual observation.

Consider chaining PlaceObject(potato, bowl) with MoveContainer(bowl, cabinet). The second policy was only ever trained on an empty bowl, so the leftover potato creates a visual mismatch that breaks it — even though moving the bowl is still perfectly feasible. We call this Observation Space Shift (OSS).

Formal definition

We model the environment as a POMDP ⟨S,O,A,T,Z,r,γ⟩\langle \mathcal{S}, \mathcal{O}, \mathcal{A}, \mathcal{T}, \mathcal{Z}, r, \gamma \rangle, where the observation space O\mathcal{O} includes third-person camera views and proprioception. Following Task and Motion Planning, a predicate ψi:S→{0,1}\psi_i : \mathcal{S} \to \{0, 1\} is a binary property of the state (e.g. In(potato, bowl)\texttt{In(potato, bowl)}), and an operator op=⟨Pre,Eff,c⟩op = \langle \textit{Pre}, \textit{Eff}, c \rangle specifies a skill’s preconditions, effects, and cost.

A predicate is irrelevant to an operator when ψi∉Pre∪Eff\psi_i \notin \textit{Pre} \cup \textit{Eff} — it does not affect feasibility. OSS is the problem in which changes to such irrelevant predicates in the visual observation space O\mathcal{O}, caused by the effects of preceding skills, degrade the performance of the current visuomotor policy.

Why prior work doesn’t cover it

Prior skill-transition methods don’t address this well. Offline transition policies assume each skill’s initial states are always reachable — here that would mean undoing the potato placement, which is logically invalid. Online fine-tuning assumes policies can be continually retrained, which is impractical when OSS can occur at every transition of a long-horizon task.

The BOSS benchmark

BOSS is built on the LIBERO simulator — a Franka Emika Panda arm across 12 manipulation scenes. We start from 44 single-skill tasks drawn from LIBERO-100, then generate modified counterparts with our Rule-based Automatic Modification Generator (RAMG), which edits PDDL predicates — repositioning or adding objects, toggling fixtures such as opening a drawer, and changing containment relations — while preserving each skill’s feasibility and logical consistency. RAMG can produce up to 1,727 single-modification variants, giving a large, controllable testbed for OSS.

Diagram of the three BOSS challenges, showing a single predicate shift, accumulated predicate shifts, and a full three-skill chain.
The three BOSS challenges, illustrated with the skills OpenDrawer(cabinet, bottom), PlaceObject(potato, bowl), and MoveContainer(bowl, cabinet). Circles mark the modified — irrelevant — predicates that induce OSS.
Challenge 1
Single Predicate Shift

Robustness to a single irrelevant modification introduced by the immediately preceding skill — a potato placed in the bowl the robot must now move.

Challenge 2
Accumulated Predicate Shift

The cumulative effect of two or three modifications from several preceding skills — a potato in the bowl and an open drawer. Harder, and more realistic.

Challenge 3
Skill Chaining

10 real long-horizon tasks, each a chain of three skills, run end to end — measuring directly how OSS degrades full-task success.

Results

We evaluate four widely used imitation-learning baselines: three Behavioral Cloning policies from LIBERO (BC-RESNET-RNN, BC-RESNET-T, BC-VIT-T) and the vision-language-action model OpenVLA.

For C1 and C2 we report the Ratio Performance Delta (RPD) — the relative drop in success rate caused by OSS. For C3 we report the Delta to Upper Bound Ratio (DUBR) — the normalized gap between the chain’s actual success and its OSS-free upper bound. All results average over three seeds.

Four scatter plots, one per baseline, with most points falling below the diagonal.
Challenge 1. Each point is a task pair: success rate unaffected by OSS on the x-axis, on the modified counterpart on the y-axis. Points below the diagonal are hurt by OSS; darker means larger RPD.
Two grouped bar charts showing both the severity and the frequency of OSS rising as modifications accumulate.
Challenge 2. Average positive RPD (top) and OSS occurrence rate (bottom), for one, two and three accumulated modifications.
Grouped bar chart of DUBR per long-horizon task, positive and large for nearly every task and baseline.
Challenge 3. Delta to Upper Bound Ratio across 10 three-skill long-horizon tasks. Some BC-RESNET-RNN bars read 0% because its OSS-free upper bound is already 0% — it fails even without OSS.

Can data augmentation fix it? No.

A natural idea is to expose policies to more visual variety during training. We used RAMG to generate 1,727 modified environments and replayed demonstrations to build a new dataset of ~57,000 demonstrations — nearly 30× the original LIBERO data for these 44 tasks. We then compare Setup A (trained on original data) with Setup B (trained on the augmented data), both evaluated on the same modified C1 tasks.

Table comparing success rates for the four baselines trained on original versus augmented data.
Success rate under Setup A versus Setup B, and their difference A−B.

Other mitigations also fall short

We tested two more intuitive strategies on Challenge 1, and neither resolves OSS.

Frozen robotics-specific vision encoders. Replacing the trainable encoder with frozen pretrained ones (R3M, LIV) still leaves 55% (BC-R3M-T) and 61% (BC-LIV-T) of tasks affected, with average RPD of 29% and 26%.

3D point-cloud policies. Despite better viewpoint generalization, 3D Diffuser Actor still has 73% of tasks affected, with average RPD 37% — even image-free policies remain vulnerable to OSS.

Together with the data-augmentation result, this shows that current baselines lack any mechanism to handle OSS, motivating the need for new, OSS-targeted algorithms.

Getting started

The code, the pre-trained skill policies for all three BC baselines, and the augmented dataset are all public.

git clone https://github.com/Boss-Benchmark/BOSS.git && cd BOSS
conda env create -f environment.yml && conda activate boss
pip install -e .
huggingface-cli download yygx/BOSS-assets --repo-type dataset --local-dir libero/libero/assets
huggingface-cli download yygx/BOSS-checkpoints --local-dir experiments/boss_44/0.0.0
MUJOCO_GL=egl python libero/lifelong/eval_skill_chain.py \
--model_path_folder ./experiments/boss_44/0.0.0/BCTransformerPolicy_seed10000/run_001/ \
--seed 10000 --max_steps 400 --lht 1

BibTeX

@article{yang2025boss,
title={BOSS: Benchmark for observation space shift in long-horizon task},
author={Yang, Yue and Zhao, Linfeng and Ding, Mingyu and Bertasius, Gedas and Szafir, Daniel},
journal={IEEE Robotics and Automation Letters},
volume={10},
number={9},
pages={8882--8889},
year={2025},
publisher={IEEE}
}

Acknowledgements

This work was supported in part by NSF under Award 2222953, in part by the Laboratory for Analytic Sciences via NC State University, and in part by ONR under Award N00014-23-1-2356.