Pith. sign in

REVIEW 2 cited by

Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03967 v2 pith:SJ6RYI74 submitted 2024-07-04 cs.CL cs.AIcs.RO

classification cs.CLcs.AIcs.RO
keywords modelsmultimodalframeworkgeneralisationarchitecturalcorrelationsevaluationinput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Evaluating the generalisation capabilities of multimodal models based solely on their performance on out-of-distribution data fails to capture their true robustness. This work introduces a comprehensive evaluation framework that systematically examines the role of instructions and inputs in the generalisation abilities of such models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. The proposed framework uncovers the resilience of multimodal models to extreme instruction perturbations and their vulnerability to observational changes, raising concerns about overfitting to spurious correlations. By employing this evaluation framework on current Transformer-based multimodal models for robotic manipulation tasks, we uncover limitations and suggest future advancements should focus on architectural and training innovations that better integrate multimodal inputs, enhancing a model's generalisation prowess by prioritising sensitivity to input content over incidental correlations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboBERT: An End-to-end Multimodal Robotic Manipulation Model

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A two-stage trained vision-language-action diffusion policy with carefully selected data augmentations reaches mean episode lengths of 4.52 (ABCD to D) and 3.79 (ABC to D) on CALVIN.

  2. SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

    cs.RO 2026-08 conditional novelty 5.0 of 10

    SAFECAST augments hidden-state failure-probe training and conformal calibration with visual and language contrast sets, improving VLA failure detection under distribution shift in several tested settings.

Pith tools