Pith. sign in

Paper Citation Record · LEDGER

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

As of 4 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 100 inbound Pith citation observations for arXiv:2411.19650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19650 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T07:33:25.188358Z

measured 178 of 178 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 174 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:10:13.556241Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T23:07:48.078477Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact32
  • verified fuzzy41
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bdad2685-5d3e-4b25-8538-2b25b201222c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.544816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:80ba2b0acba8dcf7dd5eb50447fb1fb4d78f0fa9d35a6bf1c947443033ed24ce

Observation a10b3bb1-ed70-422d-9a01-09135e8e5cb1 · outbound

This paper cites GPT-4 Technical Report.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.368933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:cdb80deb86424747f0bcb7ee2a4c167474342408ee4cee8a3360ea659afeae7b

Observation a328a0f5-29de-43d7-8740-83e84761cac0 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.553388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:1f92fb38118eaec493fa4006490200a0a86eafc62b41ecd778d6c82b11884aa1

Observation a9405621-2096-41ca-8f50-d507567ba0ee · outbound

This paper cites Hydra: Hybrid robot actions for imitation learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Hydra: Hybrid robot actions for imitation learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.881899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:07383bcba080b0129170ae1956e7e22eb3075b6f68486a4fd6515483cbfa7651

Observation 7ca60efb-2d6a-426f-a05c-562b94407f37 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.597427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:41feb0940c9627a364a844e94d0ae1ebc0d0e608c68ff8d2a98dd7505a736f19

Observation c8ecfdb4-8137-4178-babc-53f856d25a4e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.417757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:0d57cd03117dce033bde2b5cdab0595b77bde5aa99f26ef97817c9b297fb10f5

Observation 929fd749-adcc-45c6-a506-7b00056469fe · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.432800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:469b8218244c46093324ec88fd33138a96e7a6991f8101e997fd0b70db03dc93

Observation d935b006-cfff-42b3-9c10-4f7454382896 · outbound

This paper cites Language Models are Few-Shot Learners.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Language Models are Few-Shot Learners

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.445351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:29a38a8adfdf513a13a8c7a5a8bbf0b75f42f70773fd7d0e75835966135cc6df

Observation d68ad38f-b0e2-44b4-94c2-a57b93c02d11 · outbound

This paper cites The ycb object and model set: Towards common benchmarks for manipula- tion research.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation The ycb object and model set: Towards common benchmarks for manipula- tion research

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.930779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:41a5355fc80ba52936fb5dc807c4f425212eac02d1e1bf324fbb9b924ff5680b

Observation 939f2fcf-40a1-4df2-8026-e1512e961037 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.455226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:515c1276699741e21e12bc89b1fd60d141033261b204f8995b9d869baf39734b

Observation 7d36be1d-c2e3-4e2b-94d2-20ce17f303d7 · outbound

This paper cites Berkeley UR5 demonstration dataset.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Berkeley UR5 demonstration dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.959625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:1fad6461cf77bfc1fdac74c01eb40527b64091071e7ce41e213580a196026acf

Observation f145f177-8fdd-43b4-83eb-23cbd2598287 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:36:10.299419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:6f757da97a0dd6c47f648738dc5b442795618044d3bc4b9b3f25ff4549e70320

Observation 52d6193b-d590-47fd-8026-a2f839625bfb · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.999357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:109e0c2e5b8baee3bccbe023a1abf14e3447e7fada954f35ba0d09cffb7abfec

Observation bf677c2b-a41d-4af2-b26a-a3ad4a2591ac · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Diffusion policy: Visuomotor policy learning via action dif- fusion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.004349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:125b58c637953696ba84d681366967200926ada15ad98d92d3b2cda5d3465403

Observation d25037d1-e163-4827-a927-102195fe502c · outbound

This paper cites Analysis and observations from the first amazon picking challenge.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Analysis and observations from the first amazon picking challenge

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.015847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:6c66bc901507e75f28f8efc39595641e771d768bcaaa870bb6458b89a96833af

Observation f008091f-c63f-422f-bba6-f6574dd46ef4 · outbound

This paper cites From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation From Play to Policy: Conditional Behavior Generation from Uncurated Robot Data

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:33:25.509355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:e7fe3db5243b6c4c6e01d6dc82436a6382d61682b57ec26e450917f50e5f95e0

Observation 91303da1-9952-483f-9594-c14c5fa0bc11 · outbound

This paper cites an unresolved cited work.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-12T07:33:26.026475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:29297f0193f5e4eb6dcac21e3334b7b82ab1044d3bbbbc0c628b3de126a844c8

Observation b84486d2-215e-4103-a69d-85fbae207ef6 · outbound

This paper cites High Fidelity Neural Audio Compression.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation High Fidelity Neural Audio Compression

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:49:52.306785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:21d55024353e4ece4c8041bcc6bd46eaf8eeb74a167d504d72a783a0d4bb593e

Observation 24be5b98-0f26-4938-8929-e01405908ce7 · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:55:55.512010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:3f10544fe68567a76f041e4f5a39dd46e8c4c668f7d1b0fdacd2bf1f11ddc0ca

Observation 51aa9ab1-8662-4175-af38-c302287a4b89 · outbound

This paper cites The brain basis of language process- ing: from structure to function.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation The brain basis of language process- ing: from structure to function

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.046691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:2afad0c64547aac1d6ac39851c8046012eae23f2640281e590ac1604c23c76b4

Observation 7566d190-7abc-410d-bcba-fb440f316824 · outbound

This paper cites an unresolved cited work.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-12T07:33:26.052152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:bc1399a49da9d1987ff2f82950a1af06fe4ffee875097aba278038f6bef79ea8

Observation c9459b40-ec60-4360-8a12-f0addd15cbc6 · outbound

This paper cites Classifier-Free Diffusion Guidance.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Classifier-Free Diffusion Guidance

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.552088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:d5fc9bf3d87fd4fea42791005c7de963d501106d0a9ae5051bf2006d1d7b17c5

Observation 67c35d05-8266-46cc-9d15-80efad97d79f · outbound

This paper cites Diffusion Transformer Policy.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Diffusion Transformer Policy

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:33:25.563090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:cac457b6560dea64ffcb5cec8ffa37153f5be007eb33819217f8771662db45fe

Observation eb23a595-7923-41c4-96d8-2829cc545675 · outbound

This paper cites Neu- roanatomy, visual cortex.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Neu- roanatomy, visual cortex

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.063041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:00817955cca609b153c2aa4f57a5e3f6803c406f285879f0e9f4c89ff688c952

Observation 73bee261-1e61-44d5-bcca-91b394ea002f · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.072356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:bc24507b874a3ab5907fa82533d564bfeee73f4ee5c9a534c3badd0b13b0d6af

Observation 7efdcb50-c867-4e03-94a2-6a2acbba6ce9 · outbound

This paper cites QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:33:25.574455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:af42b40f25955c2d3d7e226d9341212094569f1375a48b5ba869782c240a492b

Observation 2bb8ec2b-6c06-4786-92cf-32f81a4804e2 · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:33:25.581840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:8e3af8df57e09e5b3f9ec1525bb0216a2557c1cd2f47ddd063f335d057a6844a

Observation da1562b6-609c-42bd-a9c5-36f0fcc0088a · outbound

This paper cites an unresolved cited work.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-12T07:33:26.100360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:8ba1b673cc76cb90e750b011d0dbdd332480fda218c502ebc9779c9e238b0612

Observation 21a8ddec-88b3-4905-af0f-0d39e0f187e7 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.590628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:768e05857ac1ce165dc4c93e3c75dd89c45ffdf18e254cfc358e3157b6fa2f0f

Observation 9c3211c9-ef97-45d2-bef3-fdee0e271d90 · outbound

This paper cites The darpa robotics challenge finals: Results and perspectives.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation The darpa robotics challenge finals: Results and perspectives

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.126449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:f54875752754feb3470b7214af41fd930e252695077b370aff8cf58044b9c09d

Observation 0350bb56-cf30-4593-a6f5-ce0ff128479c · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:9e492455a8bed6f029c62786c8389c8fbfbb1efe87f41c67b22e3875dae73a57

Observation bda642d4-e433-434d-8fc2-2fd4ca3039a5 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:06:19.706850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:3054f2e5e2a80ea8d3e9ffab41294151a8f1826f67f2f72a23b4249e00472a1d

Observation e5b52590-9bb6-4238-80e3-8cadf796d462 · outbound

This paper cites Vision-language foundation models as effective robot imitators.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Vision-language foundation models as effective robot imitators

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.158880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:9a71e3091f3dae4c75e2fab2bfa53126a7620fe6b684846cfaea762ef8074be2

Observation 4276f456-64a7-4f79-864d-4ff7c93e2cdd · outbound

This paper cites Visual instruction tuning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.760411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:249d16c797e3456c7a2012ab454306344dc6f1f93d0b6c0f0b51a48411fbfc4a

Observation a2d5a6fc-f9ee-4d0b-934f-8567f38a26c4 · outbound

This paper cites Robot learning on the job: Human-in-the- loop autonomy and learning during deployment.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Robot learning on the job: Human-in-the- loop autonomy and learning during deployment

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.766337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:7b7a9018a4369d59e2bd30f15a0fea66a0fc4cb688ece01ed899093d02b393f0

Observation 3c17883e-c440-48ed-82a5-cf8c02d54682 · outbound

This paper cites Visual instruction tuning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Visual instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.774218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:a9c01c623b562b34ac3cbcc6b382a5ec0175a7c2aa3323ce02a9adf11c02026b

Observation ba9e594a-06b1-4aec-a0d6-fd4607716c92 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.615474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:3b4a9489c9671a97317779f6a55a3079b894bf2306c94600d2dc7c62145e4c4e

Observation 61928e4d-253a-4c78-9541-ded92475db90 · outbound

This paper cites Multi- stage cable routing through hierarchical imitation learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Multi- stage cable routing through hierarchical imitation learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.790874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:58c396e7df1ebb8dc67228b80e9b046126da6120ee55589a7600d9fe4e1538c4

Observation bad407ba-9c8b-428b-9f5b-b932c445796f · outbound

This paper cites FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:33:25.623067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:f831158392bd25464cba9ded42fa8ea9716180a22ec07ede7cafe719afd2cc58

Observation 84700920-7c13-4375-84bf-c96dc51433f1 · outbound

This paper cites Interactive language: Talking to robots in real time.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Interactive language: Talking to robots in real time

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.799567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:a3df7574ea280ae7082d7608c499a43ee860c75518ea8c6268cdb1b07346a053

Observation fef41e66-8054-4bf3-88ed-952b03dd747b · outbound

This paper cites Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.804119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:904710b679b1c15055871774d13be700cd80373b3d6f927c71226394ef6f8ba5

Observation 1c81b962-d6e3-43db-8602-303796cda353 · outbound

This paper cites Grounding language with visual affordances over unstruc- tured data.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Grounding language with visual affordances over unstruc- tured data

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.813362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:8d9988e1aca5f3425e32000f5aeaa6275ff2a063e9ceca388d3806cc59cb5883

Observation c537171e-d526-47bb-9218-5647c024c325 · outbound

This paper cites Struc- tured world models from human videos.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Struc- tured world models from human videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.825024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:533a04d60b8dc7d6b1b0e51fb8c360fb8a0c3c3e5bdadd21f579d90c6d821c97

Observation 4afa4b1f-f088-4740-8347-e3560d33114c · outbound

This paper cites R3m: A universal visual repre- sentation for robot manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation R3m: A universal visual repre- sentation for robot manipulation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.829540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:22ef1c304dc70e7ee99a61042ac95403936532dcb5e4db6141056a0741f62dec

Observation 64ead727-bb12-401c-8a28-758af0ea720b · outbound

This paper cites Learning and retrieval from prior data for skill- based imitation learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Learning and retrieval from prior data for skill- based imitation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.833425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:676bd709de5a4005a9a41738bd8155d6c79aa0de665f69f5bd565b0e7f80d24b

Observation 9155c0ba-a5e2-4f5b-adaa-7a1a74d07f25 · outbound

This paper cites Improved denoising diffusion probabilistic models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Improved denoising diffusion probabilistic models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.841455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:4092b6b99654c509a30170c0d04eff98974cfc0ecdf6d2a9f10e5b70ab8f2194

Observation b6a80201-e81c-4c84-ae3a-a8b2b45b3f0a · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.629428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:d3ede693e501a1fd92f546e41fcd8f108423ad071788873cc9b1642394c63dce

Observation f8364b6c-e6c2-4abc-9337-3486a98f7904 · outbound

This paper cites an unresolved cited work.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-12T07:33:25.865030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:e0ed8841d7e48c01619002adead000f43fe7daf4506e759a5c957cbbfa2c6ba7

Observation c1108305-3206-42f0-a79a-96c10a057f9b · outbound

This paper cites Imitating Human Behaviour with Diffusion Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Imitating Human Behaviour with Diffusion Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:33:25.636739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:585c20ee192ee1c51341dde829969217b2fe62a0fb50c59a9ed8b7ca6982f275

Observation 7ed1e77b-176f-474f-9ded-f7db409e4ed9 · outbound

This paper cites Scalable diffusion models with transformers.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Scalable diffusion models with transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.905344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:9853735a70f47eedfe67da707b0d34e57ca47f9da6d2b302b57ece917d5ebf58

Observation 6504951c-d04e-4fe5-a51d-074c8032b53a · outbound

This paper cites Shared Control Templates for Assistive Robotics.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Shared Control Templates for Assistive Robotics

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.913335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:07045a46216239085446481b81974fd66bf2fd25fd5b3b05e22eddc5e01f3d11

Observation e6350ad4-357d-4968-87f0-1e63627388fe · outbound

This paper cites Goal-Conditioned Imitation Learning using Score-based Diffusion Policies.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Goal-Conditioned Imitation Learning using Score-based Diffusion Policies

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:33:25.648800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:9d075ecfd61294e924c4104506e455c8046cd41c2795acccfc1fff2b374e67fa

Observation a3555736-4d3f-4625-869c-82cd74ec6237 · outbound

This paper cites Latent plans for task ag- nostic offline reinforcement learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Latent plans for task ag- nostic offline reinforcement learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.939198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:9d3db3c270549fc3b1ce80a2a4f36a4484c51b74dc09f7b8ceb70765d0a3648f

Observation d9feb6a8-1fdd-403d-a0e2-1ea6b04206e3 · outbound

This paper cites Multi- resolution sensing for real-time control with vision-language models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Multi- resolution sensing for real-time control with vision-language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.978418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:3502ba72b101504304f5c952de0f6e2681d0e5844c045fe7378a0df7dae194ec

Observation dad853bb-44d7-453a-b037-1740021f637c · outbound

This paper cites On bringing robots home.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation On bringing robots home

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.022145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:1341edac545f791c9cff9fdb9b04492a3bc7a3636cbf5f13f39ba3676fdca211

Observation c600602f-da84-40da-8297-fed8414fc894 · outbound

This paper cites MU- TEX: Learning unified policies from multimodal task spec- ifications.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation MU- TEX: Learning unified policies from multimodal task spec- ifications

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.033590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:857a29ec3935a46a8ba6c493a57833ac93074765fe6c631352acf4df18212342

Observation a598f94d-4bec-4f05-93b3-70ec06e1ffba · outbound

This paper cites Perceiver- actor: A multi-task transformer for robotic manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Perceiver- actor: A multi-task transformer for robotic manipulation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.038252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:5b09a2f1afc85770764a4c0d951a6d586c692dc14068d15073231f946829efbc

Observation 9c889824-e32a-4ebd-818b-9447e343f0e0 · outbound

This paper cites Denoising Diffusion Implicit Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Denoising Diffusion Implicit Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.660353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:c774c4ddff920803542dfc5e54543fad4b7c5ee7f203310e915155ad2808974e

Observation 136df55e-9bbd-4501-9422-3109b549eb85 · outbound

This paper cites Open-world object manipulation using pre-trained vision-language models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Open-world object manipulation using pre-trained vision-language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.059053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:72c14dac2b43dfb008b5646245c54b65ed424a744ffe18054ab527d4afd2fa4d

Observation a4de158d-d32c-4e43-ac4a-cd2d6f88bc37 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Gemini: A Family of Highly Capable Multimodal Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.674285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:bf2fa3db4fc7fd0559aa70a3625b2b286e62c1b00665ddbf2cca8acc73fbc46b

Observation dc12dba8-a312-40fe-ae79-7c9701e7d606 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.682898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:8eb0b33626fc83db17b818ca79c5d4d2dd2678c378e5ec782a414e829d7f2ec4

Observation 2488e747-09f6-4f1d-8a75-befb7e54c0e3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.688847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:f0ff9431a98e74a186ee11b05382144045b3f1de852018b94ec08bc95ed0a513

Observation a036d4c4-3bcf-4f27-bac1-aae2db5efc80 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.699948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:74a0c2ccc54d97a283872efd9431900c4beac3686cf822182eb861a37ec98242

Observation df40284e-3bf2-40e4-8dc8-320c428a655b · outbound

This paper cites Neural discrete representation learning.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Neural discrete representation learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.146469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:c248949746f44ac97b919c2d9fb9e8eba91c92f3c33cfbefc700a0a993fdcab2

Observation f8065990-ddcf-4541-86c1-3e97d334e754 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Bridgedata v2: A dataset for robot learning at scale

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.782490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:252b019bd9797958952c92ca3e0eee1f5548fac16ba8a6c6fa2f0879bc30c942

Observation f8ff1e47-d93b-41d2-a82e-d5c3acbcd3d0 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.208161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:faa70468b2354933f0732a0b326b27b61a039b1b2267dd493a94580c2d8c80fd

Observation 620e002c-2485-4775-a40d-c18df9282237 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:32:05.974462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:33be068ed1d148fbb33a9cd10c04b809b2be37971fda52fafb2c2f9509dc4768

Observation 5f02ae89-c0be-4da8-919e-6d74475607d5 · outbound

This paper cites Unleashing large-scale video generative pre-training for visual robot manipulation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Unleashing large-scale video generative pre-training for visual robot manipulation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.893351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:83fcbc30ad1e20b002856cff171b2ec3e5cab4fbb325dcf9464e9e618fc958e7

Observation 18e3cea8-af13-42ab-8011-4ed01be40baa · outbound

This paper cites ucsd kitchens Dataset.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation ucsd kitchens Dataset

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.921349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:b150b41b60a3d6e191ef8a707a0d4ae5816a43cef8b26d41cfe4dc14d8c71149

Observation 4c528fd5-a176-4975-9df2-3467b44afdcf · outbound

This paper cites Physi- ology, motor cortical.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Physi- ology, motor cortical

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.055623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:63012ee2a7945ea4f5c46419052a79b758fc80a97de2898a9c2bc39f609f00b6

Observation b73c6851-bd68-4f63-9cf0-aff0883c8baa · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:2f60781b6bbce3ead16f0ae839bfc9f6b3961778357cb130cb5388e807e886ca

Observation 795d8f4c-b9be-48f3-a807-57c729e3ecfa · outbound

This paper cites Soundstream: An end- 11 to-end neural audio codec.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Soundstream: An end- 11 to-end neural audio codec

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.094227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:4691cb6f1eb600a6be1df51fe3ce9400ddfa53b13083bc1eee24e2f7601a0ad7

Observation 89d48fc8-5ef0-4550-a0f9-8b71e5aa984c · outbound

This paper cites Sigmoid loss for language image pre-training.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Sigmoid loss for language image pre-training

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.118291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:7a249c8e332522d6080b3b8529660a524693e6ac8a5614f9db11703cd0cc62bc

Observation 1abb2c93-fe40-4382-8c29-4fa1061f9842 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:33:25.751136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:3671a2ae103ae82d4e0c735ea03e28232668e01e729d5b05d845d40ebd9a23bd

Observation 0cad2161-d4b7-4ae9-8041-726127041dbe · outbound

This paper cites Train of- fline, test online: A real robot learning benchmark.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Train of- fline, test online: A real robot learning benchmark

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.795339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:d9d1b6c73b7bbc12d5cb56ae51ea7de936f933fcdab0c78efe497c5ed5a36cc9

Observation 7d9f5cc2-d716-40c1-9604-9baefe8cbb82 · outbound

This paper cites Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:25.855664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:48ed69edc70d902d80e1e3757dd95a0a6b9a06bf453c1f86f317cefcfa4492c2

Observation 95a1be72-d529-4d67-8bc3-5d5986b4263d · outbound

This paper cites Vi- ola: Imitation learning for vision-based manipulation with object proposal priors.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Vi- ola: Imitation learning for vision-based manipulation with object proposal priors

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.077047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:039d14931bba57112a5bfc9228276425f138d2a44ee88c92ac00dc6b00a8a73d

Observation 4a1b8fc8-4a49-4769-890b-247060848ab6 · outbound

This paper cites pick Coke can.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation pick Coke can

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T07:33:26.136527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:4e5d223857dd69ef315c77fd70d3c7439d52e6379d0ec3be83581cd90ac90128

Pith citing papers

Observation a27b1738-825c-45dd-9294-872b061cb96f · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.510781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:03e855c30354bbc9359606ddc8c53cf9de89a2a892454866341b4265d17b86bd

Observation ec5f6f61-d0e8-4a6b-9991-89f108183ab6 · inbound

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation cites this paper.

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:14:17.890793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T22:14:17.798964Z digest=sha256:41e3228554fc7f29b63b8566619607caed03eb3f2f1bc978471a3acc7c79da7b

Observation c0784892-5359-4eb4-a0a5-9822fc377a53 · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:12:19.718746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:f44a991588af4cc5ab79bdc36e884ffa8a6fc245dc9d02bb1ef2662afb11a534

Observation 6d7b5753-3a88-4cf9-ab63-0a5e52d7584a · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.385552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:44c90f3f611e5d812305ec9ff748c9e4ba698db03319d032afc3fd6c4dd65036

Observation 70a55dd5-a052-4f79-9548-f2001a1cfd4c · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:00:48.836249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:a1c1539b9956052f77452ead3c8ee9ac9cbb3b1e9d2018647ef68ae8ee5d5bad

Observation 9dfaf2b2-2e11-45fb-91d2-6a8693942608 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.957536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:4d583d72d248f6fb75978a2a5dc46e28c260430adb67fb46e12c9a2ae0c8cfd0

Observation f1fb35b8-f6fa-433d-9e14-8971a8802d87 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.331247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:ba4e8b513a54b8773fdc6c09d8b7a096f1cdc45adb0297ee9b745dfc6e2274fb

Observation 02548de6-f486-422f-8613-0afab74b1cbc · inbound

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation cites this paper.

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:40:27.640729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T06:40:27.206337Z digest=sha256:77591eb7d845761ca759d7fcbf5dc8293f0c24be44f44e3081de5ac96392402c

Observation e9881423-aec0-44e2-8a90-2c4b36631448 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 258

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.278329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:17cf86e4aecbcff6a40a8bed84bcc124e9f1e846177ad3d9467282e98b265ebf

Observation 64afa7fd-d989-4831-813c-e537b3a8d5b4 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.468983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:358e136a9529e84402f2a9f29f26f8ece1a51c97a4b94a72bfe7db968ae37b99

Observation d08b26d3-4b3e-40d5-8c62-95b628d49435 · inbound

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation cites this paper.

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:52:57.582714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:52:18.984005Z digest=sha256:a6a931d2e13a2445f8004deb2ef1bb92da144e5264e4aeac41179e7c565ff865

Observation d4222ffe-a0d0-430f-b66f-a968ee9b1f6e · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.682858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:c17be461f28da007f77c712b56df14b41ac9da3896e1775a29e48b7c76ef8ad9

Observation 94de6eee-62c7-4672-954c-2ffe4c28fdd8 · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:52:03.083704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:68f05d50f1e6c9e8f35c200c488da523aab5f5f7974dd94f253533334e4f49f0

Observation 379ce469-9b70-47ae-b86b-f13114d76289 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:28:16.330264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:29ed6da21d12579d9cdb87dde0bf07b374dd8c0aa75653c0d1b81e6c7bc7e1e2

Observation 5499412b-8750-4524-8f5c-eb674b7e6d91 · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.556241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.556241Z digest=sha256:9c93b8916c758bc58236452dde0c1b5c1f60b54cebae23abf67f142b566b568f

Observation d935dbba-20ba-4735-800f-079d712a1010 · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.797170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.797170Z digest=sha256:927539e94c2b3c21e6b3fcb0ad5681cff9d8965b20da1fbebac7974ed73b91ee

Observation 80494ba4-6565-40e0-a562-7a1910025732 · inbound

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue cites this paper.

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T16:17:26.003505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:17:26.003505Z digest=sha256:b5f929f59ffc2b599697b0f4f3a4e44c71a839312c9d314d6d42983e31957b1a

Observation 4bf2ccd7-86ba-40a2-9672-9a0f50fc6633 · inbound

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations cites this paper.

RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:36.912617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:58:36.912617Z digest=sha256:85f4ebd20383d83eb2f69c6531ea1ec7dc071f049a43512825a3e7e3f06b171d

Observation 0ec6e6c9-6307-4f76-b434-01caf8d000ca · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:49.586172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:49.586172Z digest=sha256:6f86ea03372b92569f9da2fbf02203dda8ea56fa2aeecfd9454fd9082d1b584e

Observation e592b8ee-39dd-44e2-8d92-2ef177885807 · inbound

Contrastive Representation Regularization for Vision-Language-Action Models cites this paper.

Contrastive Representation Regularization for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:07.770243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:07.770243Z digest=sha256:7b34b418e6357dc82ca349d6dac990e6d3f9aecef424e4e67e9d3e08fe91ad6a

Observation 89caafa7-a060-4bfc-9977-7146770f40de · inbound

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation cites this paper.

R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:51:09.010975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T08:46:51.278024Z digest=sha256:8d22275911b86cf446e2c1691b041fc492de283faee644057b155d7c998c8229

Observation 23db38a5-6796-4753-a499-10603157dce8 · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:30:18.178835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:d96f6845ee7a1be37595bc1ebf24a654b5a3ccfbc8a6b86beeae09e4c2386c9a

Observation dd7dfac0-5aee-4ef5-85e6-b94a5a7173eb · inbound

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference cites this paper.

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T21:13:56.896555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:13:56.896555Z digest=sha256:c99605d04bc1483791992f01a34113213f3aada9f28de1a8018df10f66e642d4

Observation 25861ae6-a327-4ebe-a1c4-e0075c521b12 · inbound

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding cites this paper.

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:22:04.626857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T20:21:12.375936Z digest=sha256:415b8bb3d2475a1e8531b043b4ec2d0510412cc32d512bb12d842fdb594e9dbd

Observation 8f219e04-3574-4462-b777-6eb5ff28b57e · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:09:09.437606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:bcd62b62a60d26b6f1ce4941a6b228ba565b3e15f460ca10adec918d00fc6372

Observation 7c19f252-ddaa-487b-baa0-9c347290fcd0 · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:29:10.015043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:51ec2d89c2ef642251115c1685e29fedacc3ba7c382a0c5f07ac84b8ebb57a77

Observation 0d09a04f-9252-4e29-b2c9-4f203e45552c · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:27.625412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:27.625412Z digest=sha256:67611dec0fbe7ba0a986d818059066745ecf4f7d3bdf0ce594ca49898ec44652

Observation dfa15989-a43f-40c4-9fd3-1a001a401664 · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:01:20.316615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:10d4a7440eed4401b63d79d312c1689c4429cc1dd471c4fc6ccbc3c060ab5317

Observation 97dccf5f-d256-472f-b89a-48a53538c6ff · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.810263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.810263Z digest=sha256:9fef59191eb97762338d77a523092ce775e57d164231c99bfd006de053523f5e

Observation 357484ab-3dc4-4dbf-87b0-9427706542bd · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:41:10.973732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T18:39:59.449746Z digest=sha256:c2739688d1e536e6403ad8184c8a0f2d886aae4964c70e1e7b9f8d5d1a621dd8

Observation 1d9dc54b-437c-4ba3-bebc-ea0e974b9d6e · inbound

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation cites this paper.

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T13:35:54.140995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:35:54.140995Z digest=sha256:49fe3eca5a457c9a3f224dd11734a72816f27ac0573a47556ab80e0dc19666fa

Observation d38c2a55-30ef-4e5f-a63e-cb167be56822 · inbound

MobileManiBench: Simplifying Model Verification for Mobile Manipulation cites this paper.

MobileManiBench: Simplifying Model Verification for Mobile Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:23:42.560812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:23:42.560812Z digest=sha256:35c567307422cf9b9bcdaf18ea131b822d34990762085125398eaf66ebbaf1dc

Observation 6214642d-950d-46ac-a6d1-7c53307f2485 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.993667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.993667Z digest=sha256:d7fc38f888bf8e46d1f73509b10e97bddbd154bcf807059ae6ed48eba2bed663

Observation 204061ce-907d-43ba-9c62-c9c7c6148957 · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:14:10.889567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:b796c77fde7de7df12f31accab87257e0dec10a19977823818e4819028acae80

Observation a5f896ce-bf66-40c6-96e1-0b15b85664ef · inbound

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning cites this paper.

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:12:11.737818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T03:11:52.645633Z digest=sha256:fafd28559354be9412c73ab2f7699bebec9726adeec29246db07522b5446bb80

Observation 69740d5b-f5af-4cf8-a247-527255dc14e6 · inbound

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation cites this paper.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.201565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:e07d769b7dc9cdd89c1393ebbfe8db91b046c5ba70a4d3c429a1dcf1c332d871

Observation 605b2227-62e3-416a-9137-7f91d5a6a167 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.630635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:bc06b361da691e5dd5566db0a71e744effd82259c4ffc3f0abbb107358f6a78f

Observation 269df38e-9a8f-4dc3-a11a-03bba2072aaa · inbound

RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design cites this paper.

RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T19:42:46.744880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:42:46.744880Z digest=sha256:d6513b293c05330fb2536d657bfa5d319d25c5b6fe6672ec30e59583f9b26380

Observation d2e0f352-eaa6-44a3-8345-87c49b874441 · inbound

RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies cites this paper.

RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-15T15:08:31.217309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T15:08:31.217309Z digest=sha256:5abee7a6fdec8225e00ce049a97ffb95a342b2c7e738c3d0c2bf54ec9ad414a3

Observation 9d71669b-a873-44c3-8a8d-b3e10713ce77 · inbound

Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy cites this paper.

Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-15T12:59:26.784081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:59:26.784081Z digest=sha256:4f1423502f48ca8a2caaf30bd7c01a51247fa3e28a19b9faf32235c72b36e6ea

Observation 2687cd36-760a-4f77-b6a1-36810ebfd7ea · inbound

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models cites this paper.

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:50:37.174987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T12:50:02.670780Z digest=sha256:48eb4b3a7dd40fdb50c88bff31cfa672acaae6c5cba0a50a43c51b208f9d74ad

Observation 621e5b19-016e-47cd-bb3c-41389e7c8d86 · inbound

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models cites this paper.

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:25:30.882438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:24:45.578472Z digest=sha256:97e19c08ccec96cce5f7e2f2b61b3cd4406064f5beee2557b0cea85cd8eaaa23

Observation 6a77913b-c5f6-409c-86a6-2271fc550a74 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T11:50:03.926363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:d915faf5767134b6d2cf7921a4701cae02439c105a429f756aef60a27bb83639

Observation c2f83144-a14d-40da-a357-b1b76f4569e3 · inbound

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models cites this paper.

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.310088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T07:37:34.182285Z digest=sha256:2fab54930b13eb0dea139d10b94f622e0f6439f6f25d7c0f0ecf04d4698747b0

Observation 93d23a39-77c5-4eef-8406-535997a96b5c · inbound

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models cites this paper.

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:58:26.724692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:54:36.125258Z digest=sha256:971b5239303bd5940d07d2b54482eab71431af2e339a25709cb6af9d3cf22f1e

Observation 27da5bd5-abe3-4225-984f-1a0e694a3056 · inbound

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA cites this paper.

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:26.297923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:21:13.658826Z digest=sha256:9f414b4a09a30851dae213738ea15b497bad75f39d14e86c71eaa8a9b8130dae

Observation 4bdba29e-5207-4985-a5d2-61fc395d6b0d · inbound

BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination cites this paper.

BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:52.472677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:56:42.059913Z digest=sha256:9e85126b480ac99775f941a78ca95abeb0195666fdd0c25cadf0fbb33f37c615

Observation 4a664327-97fa-4bc3-bc67-5ee2a948409f · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:25:58.995772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:560eb17ace8c548d7ab459e03e0905891619eb18e05387915393aa2a4bfb40f5

Observation dfe11aac-4bb4-4ccf-921c-573478b4986d · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:64d1adddcb9683b49fde7368b196b58c15931c1adf194b863a1faf3d73e2fd06

Observation 6707b904-7c09-4902-8169-9c4919e93b82 · inbound

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation cites this paper.

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:11:01.569931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:11:29.558490Z digest=sha256:850393476fba3bde177bb8102d6d4840160703adda3fd22edaa868f760ad1125

Observation 07d74926-32b0-45a4-a490-6635cdf4a33b · inbound

DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models cites this paper.

DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:05.092759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:38:07.136529Z digest=sha256:b95af5c364f47f9f8a03fe4f83dfe853dcc3700b315cd3fc44e360ede4132577

Observation 8f585254-030f-4ae2-bd7e-fda12d0e6807 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.844783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:01150a456f86b687552ab60eb6e5d08d9f0b010d23f1de06b6191227ea54d085

Observation c45aa43b-8f0e-4a04-bd90-f9fcf59b1b88 · inbound

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning cites this paper.

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:48.443266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T05:06:38.517652Z digest=sha256:8fe0a9a83d023404f6dfeddc790c1d351802af3aea49282e6e6e05f75d28fc33

Observation f31d60cc-5fa1-4afb-b8bd-b7b680dd5f3c · inbound

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation cites this paper.

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:19.829059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:46:21.784769Z digest=sha256:5dc117574ee482410a3999925b974645371e179de9e13c030487a3af85944e9e

Observation 948a5e35-2362-45ce-a839-a7ea9b88b9b3 · inbound

Mask World Model: Predicting What Matters for Robust Robot Policy Learning cites this paper.

Mask World Model: Predicting What Matters for Robust Robot Policy Learning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:11:06.268353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T02:14:17.676675Z digest=sha256:fbad8ac4ab5ecc9a1d58df0f75aa7afde92526b9fc9dfa1c45542366f4d2e002

Observation 5c08c2c7-fee6-49ca-b403-41a23b5a2ded · inbound

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations cites this paper.

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:51:29.216669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T08:56:32.164424Z digest=sha256:64215e8a54b4d75ff078d8b24883f5fdd6781e81925dfb9ae65fd40b8395f1d3

Observation 07f9ad9f-4ea7-449c-8bfb-f0928ed58e89 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:31:30.005255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:94d07465da66e34faa73395de2e9e022138f692b2454bdf1a0981b34b7ddc01c

Observation 93557888-4f12-4824-8f50-9563626760ef · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:06.318330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:cc14745a311dc7cd267419140884f74ffad7013fbc5119ec7475a71634a5e62c

Observation 9edd534a-2cce-41dc-9b8f-c36e3dac82c1 · inbound

Being-H0.7: A Latent World-Action Model from Egocentric Videos cites this paper.

Being-H0.7: A Latent World-Action Model from Egocentric Videos CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:01:04.765357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T20:48:01.461993Z digest=sha256:2327dad6a4679a4a94753fc694b59c217d4bd1453928750122f5e3be87b850de

Observation 42599b07-c3eb-4c0b-9b31-3d02dfb79048 · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.490784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:27c823ebab96ed9b7a23161656cfd1e5d294853443c63b98d1fb5aa0fa1c64b4

Observation ae6ce2d8-bf45-4619-9bc9-e132bbeba89c · inbound

Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference cites this paper.

Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:11:14.495324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T17:49:40.820795Z digest=sha256:875e5d8bda0da9654d04bf2c06a3cba033c3ae39c432842e4964cc347a3f7254

Observation 0f171601-d7fd-40cc-b4a2-c69f1a79ddda · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:51:30.791952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:1ea83a87cac013a0dee93a2879e80dcc944446fa2a2f418325b31e3e64ceab00

Observation 6724b243-19e6-4503-b68b-bb428aa4cde3 · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:34.793904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:b95e9b6fd8ae593696be75ddb360cd587d07bd5a94d51784da07445e5513dc87

Observation 25cfe5ad-7f94-4a79-8368-92dbbbe4039d · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.311573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:094845085043dc52fecbbf9c2329d7a8815ceff907a214c7a7d18c87131a8ec7

Observation d1b4786e-4d04-43e0-a9c1-6e4a93580a56 · inbound

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation cites this paper.

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:09.151382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:54:48.058400Z digest=sha256:4e074bc44ca6002773f5d9e3692340e89bd1a30a1d527111d7ddf69e416b6d1f

Observation c83e848b-98ea-4814-a5a1-782200fac122 · inbound

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation cites this paper.

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:10.831102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T08:52:07.545191Z digest=sha256:06df2670099b84176dfb70d75be2170eb4783942c1c5897850f7b0c1f421e6b4

Observation 64e7d716-a1ac-4fed-8b1b-dc5b30601863 · inbound

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models cites this paper.

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:57.872697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:24:34.761241Z digest=sha256:1a1534d4c94232d526a432f3d6df01c5a86929c2871b79e754e1342c31fdf3df

Observation 357629e7-9cd4-4b9a-9ae0-b57766986edb · inbound

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models cites this paper.

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:23:51.423710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T23:21:10.571364Z digest=sha256:325aa53bd63474fcb0ab4050877f809141b88b8f81d7c714e9576bac166f6a1d

Observation b1253d25-9c34-4b1e-93c3-7625cbea20d6 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:52.967184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:0b61aff5c5bcf82e9d3b8570ba7dfcfa386da249b1e5bd4c1bc455cb526728fd

Observation 6719f627-0f5b-4ec2-aecd-3486be808042 · inbound

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations cites this paper.

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.446889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:51:28.068087Z digest=sha256:ecd512c2f7b12fd0721caf41dc569245953494ad96a489ec810b399a7f2f967d

Observation 691951da-9af1-4738-bcad-72d9c9b8fdb5 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.566927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T03:39:41.090350Z digest=sha256:13800c5b883267266b1c01f01c0f7643f55c28dc66cb9d67aa63c1b62131ee16

Observation 37ddafba-0a1a-4bb8-92be-315466bcab0b · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:26:29.258603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:53:54.608425Z digest=sha256:f8730d0b547196bb93116a84a7c6d5564c14845f323752b200f4022d38acc25d

Observation 73b32afc-3b72-4fed-bef6-2d8b957e73b0 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:19:49.729723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T06:16:48.180290Z digest=sha256:5b63fb2354b6f77c3d564f065b83759a56c7b23d509fcef9c841de86bf4c7cea

Observation 88ff709f-1491-4697-b1fe-c9f88588626c · inbound

Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval cites this paper.

Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:27.208209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:34:27.255728Z digest=sha256:75f0b5d73ee6bcfca97df1e5740040045297522fcda324d6e70afe3f551a2c8e

Observation 94a70600-bc5e-4ef3-87c4-0944b1245702 · inbound

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs cites this paper.

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:27.269149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:46:01.737920Z digest=sha256:50e83e2502e325975c15379b60ccfeaa4c988df6289c098ada95c4b449e0a837

Observation bc576819-4022-4039-9e16-42084e59754c · inbound

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs cites this paper.

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:52:32.104394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T07:50:13.845621Z digest=sha256:bcf18b3c5c3d9d97a0a17df0631053ae3f319422b38ff747eedccc97be9acf6e

Observation 4259af01-e8ad-43a1-b06e-dbfb31c362e1 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.891072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:14:54.885244Z digest=sha256:94a65f98ec695c08e615af0b0bd077a359bc25b6aabf8fd7c2d600968664d76a

Observation ef9acaa9-a358-4a16-86d5-af225a980fe8 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:17:59.468951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:14:56.501485Z digest=sha256:f6cb23187d2d2d09d40858144adb3deed742163bf76f66547bc355d8a12c4c23

Observation 12b7695b-a460-4cef-aacd-ad59a1cffc09 · inbound

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark cites this paper.

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:26.545012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:50:18.706396Z digest=sha256:55723b4eb992c487aaec46f13481ed65ae91ba123201f9b33d7319425729a6ca

Observation 82a1eb27-0d2e-41b7-a33a-a8f5af80db32 · inbound

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models cites this paper.

PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:26.040878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:37:53.732205Z digest=sha256:41184b906ec8fb3ccbaca9823b044f7f7fb10361afde8b6451929be2381ea809

Observation 0f748f54-4a10-4a34-b3cb-b7e1a97ddc5e · inbound

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models cites this paper.

HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:24.239747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:28:18.730751Z digest=sha256:0ce0fa72e0a43fc6135a4002f9f56513377d31062c0dca232167a83ae0a9b68d

Observation ebcf0aee-58ff-4df0-8f2b-43fb44824646 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:07:00.237644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:4848d2e9725986f47e93831c712ecd4f78c059aba995244350936a03b1978a98

Observation 95b53f66-8db5-4743-9071-0ab1dc07222a · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:52.585972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:52.585972Z digest=sha256:7e14845c2f3e36e869e2c638f8930351ca7256ea4216b61f24c2c7ff26eaf892

Observation ad5bd448-e4c7-4b4c-b9e7-b48bb1a7632c · inbound

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model cites this paper.

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T05:47:21.264665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T05:46:28.582844Z digest=sha256:9fb0c6fa28cf912de132a5962ce0d756f52e1eef56c7cde983c1056c90f40e55

Observation fef89edf-faa5-4a28-ae26-b5dc647fc156 · inbound

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation cites this paper.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.521709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:ef9d5e4d65f2617c59e441feef4791d86724bf9e26f851c255fcef9fc4462068

Observation 10ba2a1d-cdfb-46c9-9ca9-12c50e46d99c · inbound

BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning cites this paper.

BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:52:32.926996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T17:51:48.905620Z digest=sha256:c6e03beceb2bacc8ce078e4073e459ced98c6eae6bed73f0a1a93a7ddab3d243

Observation a2a914b1-053b-489d-b470-07912a04480e · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:47:37.226881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T18:43:08.029165Z digest=sha256:a044e45f44f1dc0bcb396d6eb45101758f3cf39bddde40a174f9bbe63671cc2e

Observation 498dc31b-ca37-4578-a8f7-db801df05129 · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:45:05.712384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T21:42:05.592911Z digest=sha256:98b18be34e3fe359989360d1c21ca1852575cd21cf86d0cc8cf5bf3ff2d287b4

Observation 6fbdd620-dbe4-49b4-b666-a77330246e70 · inbound

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models cites this paper.

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:32:34.400201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T18:31:09.883545Z digest=sha256:95dc36a22d733c4cf90a7be78027ee9ce3449d59f2fc1ea5b3a452df1e0aa4ed

Observation ab9a7db4-067e-4973-8e44-01b03e2fbb69 · inbound

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models cites this paper.

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T14:10:05.961824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:10:05.961824Z digest=sha256:88ff591976ced1b1dcff4b45d4bcea0ad8e47bc6567725276c9678977a8e9cff

Observation d33b3de2-4ea3-4822-851e-42c751257825 · inbound

FrameSkip: Learning from Fewer but More Informative Frames in VLA Training cites this paper.

FrameSkip: Learning from Fewer but More Informative Frames in VLA Training CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:02:32.398581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T18:02:04.383290Z digest=sha256:d766cd3b5c3ba7ee8bff8edc3c86cc106614b5925df0f2539281005b3b1b089d

Observation 15a093e0-c38f-4079-aeec-b38dd886283c · inbound

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation cites this paper.

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:46.767128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:55:56.318157Z digest=sha256:ab31d44d5c9bd1e92d250fd5c723fde34dd6da940657b6e10cbc196994774612

Observation e6bae885-bf40-4605-b46e-d91184b38a36 · inbound

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation cites this paper.

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T18:59:49.358909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:59:49.358909Z digest=sha256:b4358dd9f5930b9c7f5c61abe53dfe4429e63c4f1442e640d22bc0defc5a3568

Observation cab310b9-d243-4fd0-a165-0b0d9a99ac7f · inbound

PhysBrain 1.0 Technical Report cites this paper.

PhysBrain 1.0 Technical Report CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:37:39.952920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T16:34:44.204055Z digest=sha256:c5194ce1b9188d7e39fd9346b35f6319a12b95bdab81ee1bfdf12230114e3f0c

Observation 61ae4a8d-066b-45b2-93e2-eef339ae3071 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:17.157221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:dc16bdacc0b8e9b431df400cbf1a06cb3844204880f434eb45e0ad2da789ea26

Observation 4ad8c2ca-eaff-4a4b-a800-2b715f7c9aa3 · inbound

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics cites this paper.

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:58:11.308054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T09:55:17.865828Z digest=sha256:29b0c61e44a475dd3fba6a64c242f944435cc5eb9065225efcee54cb6b0764e2

Observation fa106154-94f9-4d79-9c29-53b48a0b9986 · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:18:07.230609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:d06c1e289d7b3f5de5f468ba50a290f620dafa6da0b65dafd941413d66e116a9

Observation 6eca272d-c2f6-4805-89ee-bb29627eaa98 · inbound

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction cites this paper.

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T03:39:29.581699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T03:36:23.117616Z digest=sha256:84328fe4c3ca1a87f127f99b80682e6fde6fb265f2de4434406ef722e8c9a7a3

Observation 2084485e-40bc-4efa-bde5-20ce469339f3 · inbound

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models cites this paper.

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:58.485034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T17:54:56.386488Z digest=sha256:1705055d077368d89f2acec90436ccdfeabef089fe08df855c6de892789313ad

Observation 1a03a248-e5b8-43b6-91ce-19f4c831ff50 · inbound

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance cites this paper.

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:44:48.577367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T15:36:13.134340Z digest=sha256:516fc20e6372256d623559d664dc8a427f01711a1e5b62943afa8c339bc1bbce