All posts

From robot runs to the next policy.

Action segmentation with synchronized video, tactile readings and robot state.


A run across its recorded streams

Recorded visualization

0.0 / 13.4 s
01 / Synchronized evidenceThe demo recording shows camera views, tactile readings, action labels and motion traces together.

An episode-level description does not locate a failed step, retry or recovery. Those events need temporal boundaries and evidence from the recording before they can support failure analysis or data curation.

RoboLens turns recorded runs into timestamped action labels and task outcomes linked to video, tactile signals and robot state. These records support human review, subtask evaluation and failure search across runs, with selected intervals available for training and further testing.

At IROS, we collaborated with Robotiq to demonstrate multimodal action segmentation using synchronized video, tactile measurements and robot state recorded from a teleoperated UR arm equipped with a Robotiq gripper.

01From recordings to structured behavior

Label the task, not just the frames.

Start with the operation’s definition of acceptable work: required steps, completion criteria and output quality. A training example needs those steps, boundaries and outcomes. Evaluation also needs to distinguish a successful attempt from a retry, intervention or incomplete task.

The label schema specifies what starts and ends each task phase, how repeated attempts are represented, and what counts as success. Human quality control checks predicted boundaries and outcomes against the video and available signals. Reviewers correct labels, resolve edge cases and flag anything the recording cannot establish.

Open. Pick. Place. Close.

Action annotation · robot teleop

0.0 / 25.0 s
02 / Action annotationA separate DROID demonstration makes distinct task phases easy to follow. Labels were proposed by a model and visually corrected by AI. View annotations · Source credits.

Inspect contact and motion together.

The Robotiq recordings show a different part of the workflow: reading an action alongside tactile and execution data. In this block reorientation, follow the approach, grasp, rotation and release across both cameras and the tactile traces.

Example recording: rq_009
VideoWrist and scene views at 30 Hz.
Tactile28 sensing elements per finger at 1 kHz.
Robot / gripperTool motion, execution state, commands and feedback.

One run. Connected evidence.

Teleop · reorient the block upright

15.5 s
Scene
Wrist

Tactile response

Raw sensor counts
Finger 1Finger 2
Full recording traces. The vertical cursor follows the camera playback on the native recording timeline.
Proposed action stagesUnlabeled interval
0.0 / 15.5 s
03 / Multimodal evidence · rq_009The robot grasps the block, turns it upright and releases it. Both cameras and measured tactile readings share native recording time. Stage boundaries are model proposals.

Contact does not by itself prove a stable grasp. A retry does not explain why the first attempt failed. Recorded signals help test an interpretation, while ambiguous cases remain flagged for review.

Each interval needs source references and an explicit status: model proposal, reviewed label or unresolved case. That distinction must survive when annotations become training targets or evaluation results.

02Task progress and subtask success

Measure the task and each step.

A failure does not need a recovery event to be detectable. Track how many required subtasks were completed, where progress stopped and whether the final task criterion was met. A run can complete the approach and grasp but fail placement, even if the robot never retries.

Score each subtask against its own completion criterion. An action label locates an attempt; video, robot state and available sensor signals establish its outcome. Keep failed steps separate from steps the robot never reached and outcomes the recording cannot resolve.

Action interval

{
  "episode": "rq_009",
  "action": "Reorient upright",
  "start_s": 9.323,
  "end_s": 12.024,
  "clock": "native_recording",
  "status": "model_proposal"
}
Completion criterion
Did the block reach and retain the required upright orientation after release?
Recorded evidence
Inspect the action and its aftermath across camera views, tactile readings and robot state.
Human QC
Check the proposed label, boundaries and outcome. Keep unresolved cases marked for review.
04 / From action to outcomeA simplified rq_009 action proposal and the checks needed to assess completion. The action label alone does not establish success.

Across evaluation runs, measure each subtask’s success rate: successful attempts divided by scored attempts. Report steps not reached and uncertain outcomes separately. Compare policy versions under the same task conditions and success criteria.

Group results by policy version, robot, site and date. Track step duration alongside success rate: a task may still finish while repeated attempts make it slower. Follow changes back to the corresponding recordings before attributing them to the policy, hardware or operating conditions.

Search failures across runs.

An agent queries RoboLens through MCP to find incomplete attempts and retrieve their evidence. Compare the returned clips and tactile readings with a successful attempt before selecting intervals for review.

03From evidence to the next iteration

Use weak subtasks to guide the next iteration.

Select intervals around failed or incomplete steps, including the lead-up and outcome. Include a recovery when one occurs. Review these intervals alongside successful attempts under comparable task conditions.

Those intervals support three separate steps: selecting training data, targeting further collection and constructing evaluation cases.

Illustrative workflow

Deployment run

Lead-upTask context
Task stepAttempted action
OutcomeProgress or failure

Select the behavior. Keep its context.

Training examples

Review successful attempts
and failure examples.

Next collection

Target weak subtasks
and missing conditions.

Next policy

Held-out evaluation

Compare task completion
and subtask success.

Use task progress and subtask outcomes to select examples and guide collection. Include recovery trajectories when available.
06 / The iteration loopCurated behavior intervals can support policy training and collection planning. Evaluate the next policy on held-out cases to measure whether it improved.
Training

Select behavior that matters.

Curate demonstrations, failed attempts and recoveries where available. Review their labels and suitability for the training objective.

Collection

Fill a specific gap.

Use low subtask success rates and missing conditions to target new demonstrations, objects or recovery examples.

Evaluation

Test the next version.

Compare overall task completion and subtask success on held-out runs. Check retry counts and interventions alongside them.

For example, retrieve placement attempts that failed after a policy update, inspect them beside successful attempts under the same conditions, and select a reviewed dataset. Keep a separate regression set to test the next version, then track the same measures as new deployment runs arrive.

04Fine-tuned models and human QC

Task-specific VLMs with human QC.

We fine-tune our own VLMs for action annotation and failure search, with the goal of outperforming closed models on these tasks.

Models propose labels and locate events at scale. Human QC prioritizes new tasks, ambiguous predictions and suspected regressions, plus a sample of routine labels to check for systematic errors. Reviewers correct boundaries and outcomes against the source, preserving predictions and corrections separately.

Annotation workflow

From model output to reviewed labels.

Task definition
Define subtasks, boundaries, completion criteria and how to handle uncertain cases.
Model proposals
Generate timestamped labels linked to the source video and available signals.
Human QC
Correct labels and boundaries, check outcomes and record review status before delivery.
Held-out evaluation
Compare model versions and closed-model baselines on the same human-reviewed examples, excluded from fine-tuning.

Measure action-label accuracy, boundary quality and failure-search precision and recall. Track human review effort, throughput and cost per usable hour alongside those measures, so a model change is judged on both output quality and processing cost.

RoboLens and the Robotiq team together at the Robotiq booth at IROS 2026.
With the Robotiq team at IROS 2026.