Accepted at ICML 2026 / Robot learning

VLCoT.Stage-Verifiable Latent Chain-of-Thought
with Local Repair for Long-Horizon
Robotic Manipulation

Verify progress.
Repair only what changed.

Yirong Qiang1,* Jiahe Zhang2,* Ming Zhou3 Yuxiu Pei1 Lianlei Shan4,†

1 University of Electronic Science and Technology of China2 Fudan University3 Shanghai Jiao Tong University4 Tsinghua University

* Co-authors · † Corresponding author

A plan grounded in progress

“Take the cube from the drawer
and place it in the tray.”

Valid prerequisiteRemaining goal
Conceptual illustration · latent stages, not a recorded robot rollout

98.0%

LIBERO average success

Four suites, equally weighted ↗

+3.2pp

Recovery with repair enabled

Compared with repair disabled ↗

19.9%

Lower inference time per episode

vs. full replanning; 0.4 pp lower recovery ↗

01 / The idea

A failed grasp shouldn’t
erase an opened drawer.

Long tasks depend on intermediate progress. When execution deviates, some of that progress is still useful.

VLCoT gives continuous latent stages observable meanings, checks them against execution, and revisits the earliest affected stage. A verified prefix is reused only while its prerequisites still hold.

See how repair finds its starting point

02 / The method

Align. Verify. Repair.

A shared stage semantics connects
reasoning to observed progress.

01

Make stages observable

Align latent segments with object relations and subgoal events from demonstrations. A decoder predicts each stage’s expected endpoint.

Expected outcome
02

Check what actually happened

Use images and proprioception to verify progress. Supported completion advances the stage; insufficient evidence leaves it unconfirmed.

Observation-based evidence
03

Rebuild the affected suffix

After a persistent deviation, revisit still-required prerequisites. Keep the valid prefix and regenerate from the earliest affected stage.

Dependency-aware repair
VLCoT architecture: demonstration-based stage alignment trains latent prediction and observation estimation; the execution loop either confirms completion, keeps insufficient evidence unconfirmed, or repairs a dependency-aware suffix. Enlarge figure
Original method figure Training supplies stage semantics; execution observations supply verification evidence. Vector PDF ↗
j = min({k} ∪ At)

Repair starts at the current stage k, or an earlier stage whose still-required condition has been invalidated. “Verifiable” describes an observation-based interface, not a formal correctness guarantee.

03 / Inside the loop

What should the robot
reconsider?

Change the execution evidence.
Watch the scope of repair change.

Completion supported

Consecutive valid observations support the grasp: update verified history and advance to placement.

A failed grasp

If the drawer is still accessible, retain the opening stage and regenerate grasping and placement.

An earlier prerequisite is lost

If retrieval still needs an open drawer and that condition is contradicted, repair starts at the opening stage.

Insufficient evidence

Keep the stage unconfirmed. Missing evidence does not certify completion or justify a repair trigger.

A normal intermediate motion is not itself a failure. Likewise, releasing an object after successful placement does not invalidate a completed grasp.

04 / Recovery meets computation

How much should
the plan rebuild?

Four measured strategies.
One reported evaluation setting.

Four repair strategies: disabled 22.8% recovery at 11.8 seconds per episode; current-stage-only 24.6% at 14.2 seconds; dependency-aware local repair 26.0% at 14.9 seconds; full replanning 26.4% at 18.6 seconds.
Dashed lines connect measured strategies only; they do not imply intermediate measurements.
All repair measurements & accounting
Measured repair strategies. Recovery: %; Time / episode: s; p95 latency: ms; Vectors / episode: vectors.
MethodRecoveryTime / episodep95 latencyVectors / episode
Repair disabled22.811.81800.0
Full replanning26.418.682062.4
Current-stage-only suffix reconstruction24.614.236023.6
Dependency-aware local repair26.014.943031.2

Same recovery instances, step limits and task-level budget. Mean complete-episode inference time includes all model calls and both successful and failed episodes. Shared visual encoding is counted once. p95 is observation-to-executable-action latency. Regenerated vectors count deviation-triggered repairs only, excluding ordinary rolling extensions. Table 6 · PDF p. 8 ↗

Local repair trades a small decrease in recovery for less computation. Full replanning has the highest measured recovery.

Original plot

05 / Benchmark results

Across tasks.
Across kinds of failure.

Published references provide context.
Controlled ablations test the components.

Full LIBERO results
LIBERO task success. Spatial: %; Object: %; Goal: %; Long: %; Average: %.
MethodSourceSpatialObjectGoalLongAverage
Diffusion PolicyA78.392.568.350.572.4
OctoA78.985.784.651.175.1
DiT PolicyA84.296.385.463.882.4
OpenVLAA84.788.479.253.776.5
π₀-FASTA96.496.888.660.285.5
π₀A96.898.895.885.294.2
OpenVLA-OFTA97.698.497.994.597.1
CoT-VLAB87.591.687.669.083.9
ThinkActB88.391.487.170.984.4
MolmoActB87.095.487.677.286.8
Fast-ThinkActB92.097.290.279.489.7
LaRA-VLAC96.499.898.696.697.9
VLCoTThis work98.098.898.496.898.0

Published references and measurements from this work. Four suites, ten tasks per suite. VLCoT equally weights the four suite scores; published averages retain their original definitions. A: OpenVLA-OFT; B: Fast-ThinkAct; C: LaRA-VLA. Table 1 · PDF p. 7 ↗

Full RoboTwin 2.0 results · 50 tasks
RoboTwin 2.0 · 50 tasks. Clean (Easy): %; Randomized (Hard): %.
MethodClean (Easy)Randomized (Hard)
DP28.00.6
ACT29.71.7
RDT34.513.7
π₀46.416.3
DP355.25.0
VLCoT58.023.0

Per-task success, averaged over the 50-task benchmark. The reported protocol uses 50 clean demonstrations per task and 100 tests per clean/randomized condition. Published baselines retain their original settings. RoboTwin 2.0 (Chen et al., 2025). Table 2 · PDF p. 7 ↗

Full RoboTwin 2.0 results · 10-task subset
RoboTwin 2.0 · 10-task subset. Clean (Easy): %; Randomized (Hard): %.
MethodClean (Easy)Randomized (Hard)
DP43.10.6
ACT45.53.5
π₀52.916.3
RDT56.422.8
ThinkAct62.424.7
Fast-ThinkAct65.726.4
VLCoT68.031.0

The ten-task subset follows the Fast-ThinkAct reference protocol. Its task coverage differs from the 50-task benchmark; these averages must not be combined or directly ranked against that benchmark. Fast-ThinkAct (Huang et al., 2026). Table 3 · PDF p. 7 ↗

Full LIBERO-RECOVER results
LIBERO-RECOVER · recovery levels. L1: %; L2: %; L3: %; L4: %.
MethodL1L2L3L4
π₀46.124.413.54.6
π₀-FAST40.022.916.13.7
GR00T-N1.530.725.315.21.5
OpenVLA-OFT47.726.614.25.6
Wan2-Policy17.816.212.53.9
Cosmos-Predict2-Policy15.514.59.90.0
VLCoT52.032.020.09.0

Each level equally weights Spatial, Object, Goal and LIBERO-100. L1: action retry; L2: action adaptation; L3: object-state recovery; L4: environment-state recovery. These four columns do not average to the overall recovery statistic used in the controlled experiments. LIBERO-RECOVER (Liu et al., 2026). Table 4 · PDF p. 7 ↗

Original LIBERO heatmap
LIBERO success heatmap with 13 methods, four suites and reported averages; common 0–100% color scale. VLCoT has a 98.0% average, LaRA-VLA 97.9%, and OpenVLA-OFT 97.1%.Enlarge figure +
Original figure from the supplied materials. A, B and C identify source publications. Vector PDF ↗

06 / What makes the difference

A useful stage needs
both meaning and feedback.

Separate stage alignment
from the use of repair at execution.

All controlled ablations
Controlled component ablations. LIBERO avg.: %; LIBERO Long: %; Recovery: %; Stage F1: %.
ConfigurationLIBERO avg.LIBERO LongRecoveryStage F1
Base policy: no latent reasoning96.894.220.5N/A
Auxiliary state supervision only97.194.821.3N/A
Without latent stage alignment97.295.022.181.6
Execution-feedback repair disabled97.595.822.892.4
Complete method98.096.826.092.4

Measured means from this work. Controlled implementations use matched initialization, splits, training trajectories, observation access and execution intervals. Recovery first pools instances within each suite, then equally averages the four suites. N/A denotes the absence of a latent-stage verification interface. Table 5 · PDF p. 8 ↗

Additional evidence

Recovery data and visual history help in different ways.

Recovery trajectories change the training data. An initial-frame input changes the observations available to the model. Each augmentation is tested independently; their gains do not establish a combined effect.

Recovery data & initial-frame experiments
Recovery data and observation history. Original: %; + Recovery data: %; + Initial frame: %.
MethodOriginal+ Recovery data+ Initial frame
OpenVLA-OFT20.825.426.8
GR00T-N1.517.821.218.9
VLCoT26.031.028.0

Each augmentation is independently compared with the original configuration. Recovery instances are pooled within each suite before equally averaging the four suites. Recovery-data and initial-frame gains cannot be added to predict a joint result. Published baselines: LIBERO-RECOVER (Liu et al., 2026). Table 7 · PDF p. 8 ↗

07 / Evidence & scope

The numbers, in context.

Read the evaluation appendix
01

Means, without error bars

The training protocol uses three independent seeds for configurations requiring retraining. The tables do not provide standard deviations, confidence intervals, seed identifiers or run-level records. Small differences do not establish statistical significance.

02

Aggregation matters

Per-level recovery equally weights four suites, including LIBERO-100. Overall recovery first pools instances within each suite, then averages suites. Averaging the level columns is a different statistic. Undetected failures and budget exhaustion remain in the denominator.

03

All model calls count

Episode inference time includes encoding, planning, rolling extensions, verification, actions and repair, for successful and failed episodes. Shared encoding is counted once. p95 latency measures a different quantity: observation-to-action delay.

04

Simulation is the boundary

Environment-state recovery remains difficult, and performance drops under domain randomization. Real-robot transfer, an equal-success comparison, and a general success–budget frontier are not established by these experiments.

The site explains the supplied manuscript and reports its measurements. It does not reproduce the robotics experiments. Inspect the data and provenance ↗

08 / A closer look

Read. Inspect. Build on the idea.

The manuscript, reported numbers,
and the research implementation.

Release contents

The separate research implementation includes documented engineering choices and has not run training, inference, simulation, tests or CI for its initial release. Original experiment code, trained weights, raw episode records and robot rollout videos were not supplied. Manuscript LaTeX/source is not included here. The JSON files contain reported numerical summaries, not reproduced experiments. Project page source ↗

About the downloadable manuscript version

The downloadable PDF is the author-supplied ICML 2026 accepted manuscript with named authors, dated September 27, 2026. It contains 13 pages, including references and appendices. The file is preserved unchanged. All 162 numerical table entries agree with this version; source links use its table numbers and PDF pages.

Cite this work

BibTeX ↓

@inproceedings{qiang2026vlcot,
  title = {Stage-Verifiable Latent Chain-of-Thought with Local Repair for Long-Horizon Robotic Manipulation},
  author = {Yirong Qiang and Jiahe Zhang and Ming Zhou and Yuxiu Pei and Lianlei Shan},
  booktitle = {International Conference on Machine Learning},
  year = {2026},
  note = {Accepted; author-supplied manuscript},
  url = {https://apromisedland.github.io/vlcot-paper-page/}
}

Citation uses the author-confirmed acceptance information. Proceedings volume, page numbers and DOI have not been supplied.

Figure

Use + / − to zoom and scroll to pan. Press Escape to close.