Pith. sign in

Paper Citation Record · LEDGER

Contrastive Representation Regularization for Vision-Language-Action Models

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 7 inbound Pith citation observations for arXiv:2510.01711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01711 v4

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:55:09.062349Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:48:33.666343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T20:16:29.394437Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bce5ec1-0743-4d7c-8548-6566d4ba1f11 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

Contrastive Representation Regularization for Vision-Language-Action Models Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.188067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.188067Z digest=sha256:ba045011d7308aa914bafb4b913ad9fec91a959b1f734ce048ee85c429607a2c

Observation b7286cbf-518e-4ace-96dd-24a59d2fd82b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Contrastive Representation Regularization for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.363486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.363486Z digest=sha256:b42dc22277911599764fda23c389040e4aabc414a81d3587b38cdcdfab797c74

Observation c6bea193-1965-4835-8961-8034b8ffadd5 · outbound

This paper cites π0.5: A vision-language-action model with open-world generalization.

Contrastive Representation Regularization for Vision-Language-Action Models π0.5: A vision-language-action model with open-world generalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.447591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.447591Z digest=sha256:e51e7c7ec5db7d520342e4baa5ff5465f494fc8826328b01fcb387a715e08ab3

Observation a6aa2bef-5345-46b2-bf97-04435c4764aa · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Contrastive Representation Regularization for Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.702040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.702040Z digest=sha256:456c8a95c010fda353cbc2b0b64a458715438ebaa87e9499430a343c8394456a

Observation e592b8ee-39dd-44e2-8d92-2ef177885807 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Contrastive Representation Regularization for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.770243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.770243Z digest=sha256:7b34b418e6357dc82ca349d6dac990e6d3f9aecef424e4e67e9d3e08fe91ad6a

Observation c93a65da-8c90-463e-ae33-75d26468fd64 · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

Contrastive Representation Regularization for Vision-Language-Action Models Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.859469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.859469Z digest=sha256:784fa2a85ad8b3e041087a858b7adac2ccd8509ce3b67ae1da5acbd01d57d175

Observation 37d05fd4-08c9-4391-b1e0-188528c9210e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Contrastive Representation Regularization for Vision-Language-Action Models Representation Learning with Contrastive Predictive Coding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.985254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.985254Z digest=sha256:5e1cff02ea18ebe7698afa70450adfb7b92c4eeb9073a061b11ebf6e31198731

Observation ebc4c4b4-72dd-4335-bef4-e9d2bf973cff · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Contrastive Representation Regularization for Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.056099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.056099Z digest=sha256:ed7a19c34c8fe3b460cfcc0a72d3d9c9d5614166639566a9eee3e548bc1cc387

Observation 721b8a4e-f81f-432b-8093-86e6492b33a4 · outbound

This paper cites RoboBrain 2.0 Technical Report.

Contrastive Representation Regularization for Vision-Language-Action Models RoboBrain 2.0 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.254972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.254972Z digest=sha256:a8792c46b9156c2297e4959ef4f53b0a7dc08f9f80fdfe544891bb30cc672a76

Observation 66231df5-5a33-4d02-bc12-78ed3707f785 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Contrastive Representation Regularization for Vision-Language-Action Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.295168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.295168Z digest=sha256:2dea8cd169898f152d7ac468f6ab0f714b8b42bf41cc80b5d47ef1e405eddb9f

Observation 1f739e7e-dcb6-4c8a-88cf-b8d7f6607c8d · outbound

This paper cites Instructvla: Vision-language-action instruction tuning from understanding to manipulation.arXiv preprint arXiv:2507.17520,.

Contrastive Representation Regularization for Vision-Language-Action Models Instructvla: Vision-language-action instruction tuning from understanding to manipulation.arXiv preprint arXiv:2507.17520,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.305121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.305121Z digest=sha256:5a9b1b73830f8520d6456f4d99bf90de93873399d0fdbc890dc3b03ef492b1b8

Observation 81a5e571-eef7-4a7d-8bfe-cc1b12e59f34 · outbound

This paper cites ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model.

Contrastive Representation Regularization for Vision-Language-Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.357317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.357317Z digest=sha256:38c8b84229391d7b7d6b1e90b6928ac7b8d82fd00fbc92c5f15bde7714730e86

Observation 141c18cf-8d78-4b88-9872-44c05f7541eb · outbound

This paper cites Under review.

Contrastive Representation Regularization for Vision-Language-Action Models Under review

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.409915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.409915Z digest=sha256:7cdc094fb20216a3cae2c39039f81f4fd77bc27426a50538ad8a4b65513f20cb

Observation a2303cfd-1190-4784-bf84-05fd52a4410f · outbound

This paper cites an unresolved cited work.

Contrastive Representation Regularization for Vision-Language-Action Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.523523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.523523Z digest=sha256:c33cdb5115c5088e891b5867099990c02e2a138fded30f42b9df2cb51b118340

Observation a8de92d7-c11c-49e5-9e85-06084d552415 · outbound

This paper cites We omit the use of future tokens (Zheng et al., 2025), as they are beyond the scope of this work.

Contrastive Representation Regularization for Vision-Language-Action Models We omit the use of future tokens (Zheng et al., 2025), as they are beyond the scope of this work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.565336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.565336Z digest=sha256:c4454b2048679214c4b5bc43d03abf64535fd6dce9bd45bf2c98daa3caa1ad09

Observation 4d0239ec-050a-4031-b824-e62640eeb580 · outbound

This paper cites We randomly sample 10 trajectories per task in RoboCasa-Kitchen, totaling 240 trajectories.

Contrastive Representation Regularization for Vision-Language-Action Models We randomly sample 10 trajectories per task in RoboCasa-Kitchen, totaling 240 trajectories

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.790131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.790131Z digest=sha256:452615605130caa3a62ebbf8ca9b5e73965267312fa47eaf11f8b2b4edc3a5e0

Observation 0b036cb1-3a6b-42d5-be7d-10a50ea3bbea · outbound

This paper cites This result indicates the effectiveness of our proposed training framework, together with the augmen- tation strategyview cutoff.

Contrastive Representation Regularization for Vision-Language-Action Models This result indicates the effectiveness of our proposed training framework, together with the augmen- tation strategyview cutoff

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.964633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.964633Z digest=sha256:377c1abcdee3a5362f7f2e07e7abfbbdfe235c8df75dd0a10378f86c967d79ff

Observation c3485f25-d6a6-4468-9a50-ec903283213a · outbound

This paper cites an unresolved cited work.

Contrastive Representation Regularization for Vision-Language-Action Models Unresolved cited work

Reference 27

Resolution
malformed identifier
no resolver link, observed 2026-08-04T12:55:09.062349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:09.062349Z digest=sha256:fea5fa82afc6923a22723371440b90e517c89d504711dcb1ec91261159db3ad5

Observation ffc5700d-2524-4e39-9e9d-3313b1a8dde0 · outbound

This paper cites At inference, we use an action horizonH= 16and execute all actions without re-planning.

Contrastive Representation Regularization for Vision-Language-Action Models At inference, we use an action horizonH= 16and execute all actions without re-planning

Reference 64

Resolution
malformed identifier
no resolver link, observed 2026-08-04T12:55:08.622522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.622522Z digest=sha256:7e958024af370e3bd9bba79993270a16a6a485e43159e8bd6728666950227f00

Observation aacef3bc-9217-40d9-8880-a3d3eea812f0 · outbound

This paper cites A Simple but Tough-to-Beat Data Augmentation Approach for Natural Language Understanding and Generation.

Contrastive Representation Regularization for Vision-Language-Action Models A Simple but Tough-to-Beat Data Augmentation Approach for Natural Language Understanding and Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.140544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.140544Z digest=sha256:f1f017a3e646d927ccd6f49188e0818ddc4c002c99ac9f9f3447cce4661fc197

Observation b7ed0226-df83-4acc-85f7-de7d50c3feaf · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Contrastive Representation Regularization for Vision-Language-Action Models SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.214906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.214906Z digest=sha256:a70e19381e32570dd0deebb22894952933e75548fc42f91b85378b12f36b34d6

Observation bf66936f-26c3-413e-b8b7-75197fa61bb3 · outbound

This paper cites Contrastive Language, Action, and State Pre-training for Robot Learning.

Contrastive Representation Regularization for Vision-Language-Action Models Contrastive Language, Action, and State Pre-training for Robot Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.104076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.104076Z digest=sha256:0effda98646dad9630913d5e17260886b8241f3b580dad68b800f377f6351681

Observation 8684ed19-c732-44f0-b64a-6561ec1fe91b · outbound

This paper cites Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces.

Contrastive Representation Regularization for Vision-Language-Action Models Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.932813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.932813Z digest=sha256:769f7dd8564aebc829992f85f2b93e8d85a8ac37be2e23b074a652a6d56c0267

Observation 0b27409a-ab77-4523-ae8d-7674f4300620 · outbound

This paper cites Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better.

Contrastive Representation Regularization for Vision-Language-Action Models Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.537700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.537700Z digest=sha256:86fcdab66456374f257735038ab7aa946d767035cd47d83971d117f3a4613de6

Observation 4efb3bb9-9d4e-4e97-bfc3-f181a9694406 · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

Contrastive Representation Regularization for Vision-Language-Action Models NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.638645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.638645Z digest=sha256:61c470af9c010e2ffec550754980cbcf047c4f530173b5888cc160ce0cec6f99

Observation 83b67c86-3d92-4be5-8a20-d3e8ea4aec3c · outbound

This paper cites Qwen2.5-VL Technical Report.

Contrastive Representation Regularization for Vision-Language-Action Models Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.264544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.264544Z digest=sha256:ddf497333bd2e286995ab7cb1c832a58d5c494e27fcf9164447e287fb064eb4c

Observation f32c699b-454d-47a1-bcb8-988ae78d1607 · outbound

This paper cites Layer Spatial Object Goal Long Avg.

Contrastive Representation Regularization for Vision-Language-Action Models Layer Spatial Object Goal Long Avg

Reference 2048

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:08.488375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:08.488375Z digest=sha256:9a6b48a5b2cceae59b2e62200d6c5a39aade55532e218c6bc2c0a19359bcf957

Pith citing papers

Observation f656816d-568a-4ff6-9255-f2bae2644b94 · inbound

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models cites this paper.

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models Contrastive Representation Regularization for Vision-Language-Action Models

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T13:44:40.679199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T13:40:59.330788Z digest=sha256:bd28c40b63182eb100df8412b67d9ba7e7e126aa258031f9bbecf61007f0a5de

Observation 284162fa-9d21-4585-b7a7-56e799c22527 · inbound

QuoVLA: Quotient Space for Vision-Language-Action Models cites this paper.

QuoVLA: Quotient Space for Vision-Language-Action Models Contrastive Representation Regularization for Vision-Language-Action Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:34:39.477553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T12:09:12.124995Z digest=sha256:b02b3aa6479a89ab27ab33aff9e81431978294c50b6ad1d6b1fc884429676be7

Observation 9e6b5330-20e8-4f91-ad74-8b56879c6d8c · inbound

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning cites this paper.

Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning Contrastive Representation Regularization for Vision-Language-Action Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.410905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:28:37.970450Z digest=sha256:247d054cdc7e9adbc24c2584b7f95f314cc191412f29e98d794f35de18a98b32

Observation 49449ac5-7ac0-458d-9085-1d1aff007f73 · inbound

FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning cites this paper.

FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning Contrastive Representation Regularization for Vision-Language-Action Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:37:26.694558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:41:51.839543Z digest=sha256:20a0ed79213a40602cf4d5291367529810d2d94de34ff4a6b358540c875137c3

Observation 24f9ca34-e0ed-414f-9595-1f9cc1826962 · inbound

Contrastive Action-Image Pre-training for Visuomotor Control cites this paper.

Contrastive Action-Image Pre-training for Visuomotor Control Contrastive Representation Regularization for Vision-Language-Action Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.886682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:16:33.702802Z digest=sha256:e00f112eb31169fb73e94aa5b0b835416dc6db00efbbf68bf72b16db860f3903

Observation 891ec8fd-1fe0-4561-b352-77ea4a9bb78a · inbound

GeoProp: Grounding Robot State in Vision for Generalist Manipulation cites this paper.

GeoProp: Grounding Robot State in Vision for Generalist Manipulation Contrastive Representation Regularization for Vision-Language-Action Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:16:29.395716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T20:12:59.778086Z digest=sha256:4abc604c3a99d6d9733c7e1b6ab12d0864f0f7e38a3972c92d49973a9b7b4e6d

Observation 8aca349e-d11e-48ff-a233-a54667a6fffb · inbound

Semantic Anchoring for Robotic Action Representations cites this paper.

Semantic Anchoring for Robotic Action Representations Contrastive Representation Regularization for Vision-Language-Action Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T04:48:33.666343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:48:33.666343Z digest=sha256:b15fcb2f4c0c335b4713edb276d2e0ec35fdad05d19a1d969364913c25ec4f32