BOSS: A Benchmark for Observation Space Shift in Long-Horizon Task
Chaining skills breaks them — and the culprit is a change that shouldn't matter.
PlaceObject(potato, bowl) → MoveContainer(bowl, cabinet): the potato left in the bowl shifts the observation space of the second skill, which was only trained on an empty bowl — so it fails. Abstract
Hierarchical robots solve long-horizon tasks by chaining individually trained visuomotor skills. We identify Observation Space Shift (OSS): a preceding skill changes parts of the scene that are irrelevant to the next skill’s logic yet still visible in its camera input, pushing the next policy out of its training distribution. We introduce BOSS, a benchmark built on LIBERO with three progressive challenges, and show that OSS is common and severe across four imitation-learning baselines — and that three intuitive mitigations all fail to resolve it.
- up to 67%
- average performance drop from a single irrelevant change
- 76%
- of tasks hit by OSS once three modifications accumulate
- 0 / 3
- intuitive mitigations sufficient to resolve OSS
- 1,727
- modified task variants generated automatically
Challenge 1
Challenge 2
augmentation, frozen encoders, 3D policies
Rule-based Automatic Modification Generator
What is Observation Space Shift?
Skill chaining — sequentially executing pre-learned visuomotor skills — is a simple yet powerful way to tackle long-horizon robot tasks. But failures often arise at skill transitions: the terminal state of one skill can fall outside the initial-state distribution the next skill was trained on. Crucially, the culprit is frequently a change that is logically irrelevant to the next skill, yet still appears in its visual observation.
Consider chaining PlaceObject(potato, bowl) with MoveContainer(bowl, cabinet). The second policy was only ever trained on an empty bowl, so the leftover potato creates a visual mismatch that breaks it — even though moving the bowl is still perfectly feasible. We call this Observation Space Shift (OSS).
Formal definition
We model the environment as a POMDP , where the observation space includes third-person camera views and proprioception. Following Task and Motion Planning, a predicate is a binary property of the state (e.g. ), and an operator specifies a skill’s preconditions, effects, and cost.
A predicate is irrelevant to an operator when — it does not affect feasibility. OSS is the problem in which changes to such irrelevant predicates in the visual observation space , caused by the effects of preceding skills, degrade the performance of the current visuomotor policy.
Why prior work doesn’t cover it
Prior skill-transition methods don’t address this well. Offline transition policies assume each skill’s initial states are always reachable — here that would mean undoing the potato placement, which is logically invalid. Online fine-tuning assumes policies can be continually retrained, which is impractical when OSS can occur at every transition of a long-horizon task.
The BOSS benchmark
BOSS is built on the LIBERO simulator — a Franka Emika Panda arm across 12 manipulation scenes. We start from 44 single-skill tasks drawn from LIBERO-100, then generate modified counterparts with our Rule-based Automatic Modification Generator (RAMG), which edits PDDL predicates — repositioning or adding objects, toggling fixtures such as opening a drawer, and changing containment relations — while preserving each skill’s feasibility and logical consistency. RAMG can produce up to 1,727 single-modification variants, giving a large, controllable testbed for OSS.
OpenDrawer(cabinet, bottom), PlaceObject(potato, bowl), and MoveContainer(bowl, cabinet). Circles mark the modified — irrelevant — predicates that induce OSS. Robustness to a single irrelevant modification introduced by the immediately preceding skill — a potato placed in the bowl the robot must now move.
The cumulative effect of two or three modifications from several preceding skills — a potato in the bowl and an open drawer. Harder, and more realistic.
10 real long-horizon tasks, each a chain of three skills, run end to end — measuring directly how OSS degrades full-task success.
Results
We evaluate four widely used imitation-learning baselines: three Behavioral Cloning policies from LIBERO (BC-RESNET-RNN, BC-RESNET-T, BC-VIT-T) and the vision-language-action model OpenVLA.
For C1 and C2 we report the Ratio Performance Delta (RPD) — the relative drop in success rate caused by OSS. For C3 we report the Delta to Upper Bound Ratio (DUBR) — the normalized gap between the chain’s actual success and its OSS-free upper bound. All results average over three seeds.
Can data augmentation fix it? No.
A natural idea is to expose policies to more visual variety during training. We used RAMG to generate 1,727 modified environments and replayed demonstrations to build a new dataset of ~57,000 demonstrations — nearly 30× the original LIBERO data for these 44 tasks. We then compare Setup A (trained on original data) with Setup B (trained on the augmented data), both evaluated on the same modified C1 tasks.
Other mitigations also fall short
We tested two more intuitive strategies on Challenge 1, and neither resolves OSS.
Frozen robotics-specific vision encoders. Replacing the trainable encoder with frozen pretrained ones (R3M, LIV) still leaves 55% (BC-R3M-T) and 61% (BC-LIV-T) of tasks affected, with average RPD of 29% and 26%.
3D point-cloud policies. Despite better viewpoint generalization, 3D Diffuser Actor still has 73% of tasks affected, with average RPD 37% — even image-free policies remain vulnerable to OSS.
Together with the data-augmentation result, this shows that current baselines lack any mechanism to handle OSS, motivating the need for new, OSS-targeted algorithms.
Getting started
The code, the pre-trained skill policies for all three BC baselines, and the augmented dataset are all public.
git clone https://github.com/Boss-Benchmark/BOSS.git && cd BOSSconda env create -f environment.yml && conda activate bosspip install -e .
huggingface-cli download yygx/BOSS-assets --repo-type dataset --local-dir libero/libero/assetshuggingface-cli download yygx/BOSS-checkpoints --local-dir experiments/boss_44/0.0.0
MUJOCO_GL=egl python libero/lifelong/eval_skill_chain.py \ --model_path_folder ./experiments/boss_44/0.0.0/BCTransformerPolicy_seed10000/run_001/ \ --seed 10000 --max_steps 400 --lht 1BibTeX
@article{yang2025boss, title={BOSS: Benchmark for observation space shift in long-horizon task}, author={Yang, Yue and Zhao, Linfeng and Ding, Mingyu and Bertasius, Gedas and Szafir, Daniel}, journal={IEEE Robotics and Automation Letters}, volume={10}, number={9}, pages={8882--8889}, year={2025}, publisher={IEEE}}Acknowledgements
This work was supported in part by NSF under Award 2222953, in part by the Laboratory for Analytic Sciences via NC State University, and in part by ONR under Award N00014-23-1-2356.