TL;DR When a pretrained perception model fails in the wild, a co-located human performs a short HumanPlay to clean the scene, then labels just the final frame hands-free with eye-gaze + voice. SAM2 converts the gaze points into bounding boxes and propagates masks backwards across the short RGB-D clip — turning one annotation into dense supervision. These failure-driven samples iteratively fine-tune MSMFormer, lifting UOIS accuracy by +54.2 combined score and improving real-robot grasping and pick-and-place success on the SceneReplica benchmark by +3 and +7, respectively.
iTeach overview video thumbnail

Abstract

Robotic perception models often fail in the real world due to clutter, occlusion, and novel objects. Existing approaches rely on offline data collection and retraining — slow, and blind to deployment-time failures. We propose iTeach, a failure-driven interactive teaching framework that adapts robot perception in the wild. A co-located human observes live predictions, triggers a short HumanPlay interaction on a failed object, and records an RGB-D sequence. Our Few-Shot Semi-Supervised (FS3) labeling annotates only the final frame using hands-free eye-gaze and voice; SAM2 propagates the mask across the sequence for dense supervision. Iterative fine-tuning on these samples progressively improves an MSMFormer UOIS model, translating into higher grasping and pick-and-place success on the SceneReplica benchmark and real-robot experiments.

Overview

iTeach overview

A pretrained perception model fails in the wild under clutter, occlusion, and novel objects. A co-located human performs a short HumanPlay interaction, annotates a single frame with eye-gaze + voice, and we propagate the label across the short RGB-D sequence. Failure-driven samples feed an iterative fine-tuning loop; the best checkpoint is redeployed.

System Setup

A robot with RGB-D and onboard GPU · a human with a headset · nothing else.

Deployment setup: a Fetch mobile manipulator carrying an RTX 4090 laptop, and a co-located human wearing a HoloLens 2, in an office corridor

Everything the system needs is in this frame. A Fetch mobile manipulator carrying an RGB-D camera and a Lenovo Legion Pro 7 (RTX 4090) that runs inference, SAM2 propagation, and fine-tuning onboard — and a human wearing a Microsoft HoloLens 2, which draws the live prediction overlay and captures gaze-and-voice annotation. No workstation and no cloud: the laptop is compute, not a screen anyone looks at, so the correction loop never has to walk back to a desk. RGB-D streams to the laptop over wired Ethernet; the HoloLens joins a laptop-hosted Wi-Fi hotspot via the ROS-TCP connector, and the human drives the robot with a PS4 controller while hunting for failures.

FS3 Labeling

01 HumanPlay
HumanPlay interaction
The human rearranges objects to reduce occlusion and produce a clean final frame while a short 5–10 s RGB-D sequence is recorded.
02 Hands-free annotation
FS3 annotation via eye-gaze and voice
Eye-gaze places point prompts on the final frame; a voice command triggers SAM2 to convert them into bounding-box object labels.
03 Label propagation
SAM2 label propagation
SAM2 (video mode) propagates masks backwards from the final annotated frame to all earlier frames, producing dense per-frame supervision.

Results

Perception improves · Manipulation follows · Anyone can teach.

Perception adaptation
UOIS · Combined Score
3.1× +54.6
Perception quality lift
before26.1 → after80.7
UOIS qualitative comparison across iTeach fine-tuning stages
Qualitative UOIS. Left → right: ground truth, pretrained MSMFormer, and iTeach fine-tuning rounds FT1, FT3, FT5. iTeach recovers missed instances and cleans up over-segmentation across tabletop, shelves, sofas, stairs, and floor-level scenes.
iTeach-UOIS on SceneReplica & real robots
  1. 01RGB-D input
  2. 02iTeach-UOIS
    segmentation
  3. 03Contact-GraspNet
    proposals
  4. 04GTO motion
    planning
  5. 05Grasp
    execution

Only stage 02 differs across comparisons — everything else is held fixed.

Grasping success / 100
Prior best (MSMFormer)71
iTeach-UOIS74 +3
Pick & Place success / 100
Prior best (MSMFormer)65
iTeach-UOIS72 +7
Real-world pick-and-place with GTO

Swap in iTeach-UOIS and the same real-robot pipeline starts handling clutter and unseen objects the pretrained baseline fails on — turning perception gains into reliable picks and places in the wild.

Who can teach? · 12-participant user study
12
participants
6 expert6 non-expert
240
annotations
12 people × 20 objects
94.9%
mean box IoU
SD 0.49 · gaze + voice prompt
LOW MED HIGH
21/100
NASA-TLX
low task load

Each participant ran one full teaching interaction unassisted — spot the failure in the headset, perform HumanPlay, then label the final frame with gaze and voice. Expertise made no difference on any measure. Descriptive results, 6 participants per group; whiskers are ±1 SEM.

Completion time SECONDS ↓
Expert354.2
Non-expert348.2
6.1 s apart, well inside ±1 SEM. Full 20-object session; HumanPlay itself is 14.2 s.
Labeling accuracy BOX IOU % ↑
Expert94.95
Non-expert94.78
0.17 points apart — prompting is precise enough that the error bars are barely visible.
Task load NASA-TLX ↓
Expert20.4
Non-expert21.7
1.3 points apart, both sitting in the low-workload band.
Where the workload sits Expert Non-expert
204060 Mental35.0 / 21.0Physical9.2 / 17.5Temporal13.5 / 21.5Performance39.2 / 45.8Effort13.5 / 13.3Frustration12.5 / 11.0
Axes are the six NASA-TLX subscales (0–65 shown); outlines are means, shaded ribbons are ±1 SEM. Totals match, but the composition differs: experts carry the task mentally (35.0 vs 21.0), non-experts physically and temporally (17.5 / 21.5 vs 9.2 / 13.5) — they move more and feel more time pressure, while thinking about it less. Performance is scored as stored, so higher means a less confident self-assessment.
iTeach-HumanPlay Dataset
48
scenes
45 train · 3 test
~13 K
training samples
dense masks via FS3
902
test samples
held-out scenes
5–10 s
per sequence
RGB-D · 640 × 480

Code & Data

BibTeX

@misc{p2026iteachwildinteractiveteaching,
  title         = {iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception},
  author        = {Jishnu Jaykumar P and Cole Salvato and Vinaya Bomnale and Jikai Wang and Ayush Bhardwaj and Jin-Ryong Kim and Yu Xiang},
  year          = {2026},
  eprint        = {2410.09072},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2410.09072}
}

Contact

Send any comments or questions to Jishnu: jishnu.p@utdallas.edu

Acknowledgements

Supported by

DARPA
Perceptually-enabled Task Guidance
HR00112220005
Sony
Research Award Program
NSF
National Science Foundation
Grant No. 2346528

Thanks to Sai Haneesh Allu for assistance during the experiments.

The user study was approved by the UT Dallas Institutional Review Board IRB-21-194 and all participants gave informed consent.