IROS 2026 demo with Robotiq
From robot runs to the next policy.
Action segmentation with synchronized video, tactile readings and robot state.
A run across its recorded streams
Recorded visualization
An episode-level description does not locate a failed step, retry or recovery. Those events need temporal boundaries and evidence from the recording before they can support failure analysis or data curation.
RoboLens turns recorded runs into timestamped action labels and task outcomes linked to video, tactile signals and robot state. These records support human review, subtask evaluation and failure search across runs, with selected intervals available for training and further testing.
At IROS, we collaborated with Robotiq to demonstrate multimodal action segmentation using synchronized video, tactile measurements and robot state recorded from a teleoperated UR arm equipped with a Robotiq gripper.
01From recordings to structured behavior
Label the task, not just the frames.
Start with the operation’s definition of acceptable work: required steps, completion criteria and output quality. A training example needs those steps, boundaries and outcomes. Evaluation also needs to distinguish a successful attempt from a retry, intervention or incomplete task.
The label schema specifies what starts and ends each task phase, how repeated attempts are represented, and what counts as success. Human quality control checks predicted boundaries and outcomes against the video and available signals. Reviewers correct labels, resolve edge cases and flag anything the recording cannot establish.
Open. Pick. Place. Close.
Action annotation · robot teleop
Inspect contact and motion together.
The Robotiq recordings show a different part of the workflow: reading an action alongside tactile and execution data. In this block reorientation, follow the approach, grasp, rotation and release across both cameras and the tactile traces.
| Video | Wrist and scene views at 30 Hz. |
|---|---|
| Tactile | 28 sensing elements per finger at 1 kHz. |
| Robot / gripper | Tool motion, execution state, commands and feedback. |
One run. Connected evidence.
Teleop · reorient the block upright
Tactile response
Raw sensor countsContact does not by itself prove a stable grasp. A retry does not explain why the first attempt failed. Recorded signals help test an interpretation, while ambiguous cases remain flagged for review.
Each interval needs source references and an explicit status: model proposal, reviewed label or unresolved case. That distinction must survive when annotations become training targets or evaluation results.
02Task progress and subtask success
Measure the task and each step.
A failure does not need a recovery event to be detectable. Track how many required subtasks were completed, where progress stopped and whether the final task criterion was met. A run can complete the approach and grasp but fail placement, even if the robot never retries.
Score each subtask against its own completion criterion. An action label locates an attempt; video, robot state and available sensor signals establish its outcome. Keep failed steps separate from steps the robot never reached and outcomes the recording cannot resolve.
Action interval
{
"episode": "rq_009",
"action": "Reorient upright",
"start_s": 9.323,
"end_s": 12.024,
"clock": "native_recording",
"status": "model_proposal"
}- Completion criterion
- Did the block reach and retain the required upright orientation after release?
- Recorded evidence
- Inspect the action and its aftermath across camera views, tactile readings and robot state.
- Human QC
- Check the proposed label, boundaries and outcome. Keep unresolved cases marked for review.
Across evaluation runs, measure each subtask’s success rate: successful attempts divided by scored attempts. Report steps not reached and uncertain outcomes separately. Compare policy versions under the same task conditions and success criteria.
Group results by policy version, robot, site and date. Track step duration alongside success rate: a task may still finish while repeated attempts make it slower. Follow changes back to the corresponding recordings before attributing them to the policy, hardware or operating conditions.
Search failures across runs.
An agent queries RoboLens through MCP to find incomplete attempts and retrieve their evidence. Compare the returned clips and tactile readings with a successful attempt before selecting intervals for review.
Illustrative Claude Code exchange through RoboLens MCP. Query: Find block attempts that failed before reorientation. Pull a successful attempt for comparison. Response: rq_012 stops before rotation: the arm rises while the block stays flat. rq_019 leaves the block upright after release.
rq_012Block stays flat as the arm rises.
rq_019Block remains upright after release.
Summed tactile readings · same scale · raw counts
03From evidence to the next iteration
Use weak subtasks to guide the next iteration.
Select intervals around failed or incomplete steps, including the lead-up and outcome. Include a recovery when one occurs. Review these intervals alongside successful attempts under comparable task conditions.
Those intervals support three separate steps: selecting training data, targeting further collection and constructing evaluation cases.
Deployment run
Select the behavior. Keep its context.
Training examples
Review successful attempts
and failure examples.
Next collection
Target weak subtasks
and missing conditions.
Held-out evaluation
Compare task completion
and subtask success.
Select behavior that matters.
Curate demonstrations, failed attempts and recoveries where available. Review their labels and suitability for the training objective.
Fill a specific gap.
Use low subtask success rates and missing conditions to target new demonstrations, objects or recovery examples.
Test the next version.
Compare overall task completion and subtask success on held-out runs. Check retry counts and interventions alongside them.
For example, retrieve placement attempts that failed after a policy update, inspect them beside successful attempts under the same conditions, and select a reviewed dataset. Keep a separate regression set to test the next version, then track the same measures as new deployment runs arrive.
04Fine-tuned models and human QC
Task-specific VLMs with human QC.
We fine-tune our own VLMs for action annotation and failure search, with the goal of outperforming closed models on these tasks.
Models propose labels and locate events at scale. Human QC prioritizes new tasks, ambiguous predictions and suspected regressions, plus a sample of routine labels to check for systematic errors. Reviewers correct boundaries and outcomes against the source, preserving predictions and corrections separately.
Annotation workflow
From model output to reviewed labels.
- Task definition
- Define subtasks, boundaries, completion criteria and how to handle uncertain cases.
- Model proposals
- Generate timestamped labels linked to the source video and available signals.
- Human QC
- Correct labels and boundaries, check outcomes and record review status before delivery.
- Held-out evaluation
- Compare model versions and closed-model baselines on the same human-reviewed examples, excluded from fine-tuning.
Measure action-label accuracy, boundary quality and failure-search precision and recall. Track human review effort, throughput and cost per usable hour alongside those measures, so a model change is judged on both output quality and processing cost.
