Pith. sign in

Paper Citation Record · LEDGER

Hume: Introducing System-2 Thinking in Visual-Language-Action Model

As of 10 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 36 inbound Pith citation observations for arXiv:2505.21432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21432 v4

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:55.651566Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:57.174397Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact2
  • verified fuzzy24
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 81007fef-4341-4caf-a768-4a4bb9bd60e8 · outbound

This paper cites Vision-language foundation models as effective robot imitators.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Vision-language foundation models as effective robot imitators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.176151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.176151Z digest=sha256:b67d0c813438bd2dfdad7c6c5920701b3cc79204b4ff098da5f5dae1cb58a8a5

Observation a679ba6e-d086-424b-9449-b56e638d267f · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.316972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.316972Z digest=sha256:21ee941ad064b4b9b73ee2c32356ac80888784e1116120803852eac9ce33b8f6

Observation fbb5c2de-20a4-4c51-9fb5-d85fe417456e · outbound

This paper cites Fastumi: A scalable and hardware-independent universal manipulation interface with dataset.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Fastumi: A scalable and hardware-independent universal manipulation interface with dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.649706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:46.445264Z digest=sha256:50ab7a6889c1308c38de513ad2b427274624e7c59832005c4d34415e5b955279

Observation f7a06d05-3a29-4c1f-aa89-74229fc47b72 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.578291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.578291Z digest=sha256:27e6474446ff5b9535387816600b1395e53ab1e4e5fce6a0100dd87246cbc7b0

Observation 72bb9ad0-c3ad-4caf-92d2-0ae9f341887a · outbound

This paper cites Learning 2d invariant affordance knowledge for 3d affordance grounding.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Learning 2d invariant affordance knowledge for 3d affordance grounding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.457287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:46.673699Z digest=sha256:10719946537e62126103a53a08b727ead71279245d9ff202de953a8e976b57fb

Observation 523ae9b6-ae5a-4258-84ba-9c07faaa7f70 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:46.813552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:46.813552Z digest=sha256:81370a9600d21ca32d98059e8a688073579e61f03675fbfdf315d345e69c8657

Observation b16ca1ec-1d7f-44e0-94fd-da39d0af58cc · outbound

This paper cites Improving domain generalization in self-supervised monocular depth estimation via stabilized adversarial training.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Improving domain generalization in self-supervised monocular depth estimation via stabilized adversarial training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.302774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:46.961211Z digest=sha256:cf70f7fe3e7f97c4751405d04157670f4aad88521d25be8c45a0efe94c90f81f

Observation 4a1439db-2ae7-4131-a03b-830828a63ccd · outbound

This paper cites MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.040696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.040696Z digest=sha256:e338cae4291e305539fbb340632a0ccb478a245109392463eba1ce877f22d75a

Observation e4c00cb3-77d4-4214-b80e-5fa8af747019 · outbound

This paper cites ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A Star.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model ORLA*: Mobile Manipulator-Based Object Rearrangement with Lazy A Star

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.143573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.143573Z digest=sha256:cb3c2884cafdf25fa7f4cddacf7a5878faab218056487d2515056a22008346be

Observation 5f130c18-80c5-4cd6-9397-eb6338f72955 · outbound

This paper cites Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.253710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.253710Z digest=sha256:e9706283e4a499de1214b86e882b372bcd3428c27ff4f1be0f57201471e0b6b2

Observation 94eaf881-2020-449e-ab3a-2b66d3dc8480 · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Gemini Robotics: Bringing AI into the Physical World

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.358067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.358067Z digest=sha256:f1ca504fc1e28661507a28a042b5dd360ad8911dff73b392b9d07cf75b80f2d2

Observation 7b973670-3451-42d6-ada0-178115d937c5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.463400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.463400Z digest=sha256:08cac31e87144f228b40496ad07e6a7d0543d10b2399c6843f7bd8a1257a3755

Observation c5d2e319-e561-438f-b073-5c14a9713627 · outbound

This paper cites Thinking, fast and slow.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Thinking, fast and slow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.579910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.579910Z digest=sha256:38b73cfcaa513213dbcfeb26cb7e3e9c958f367e0b236bb1257b4cbd49b8075d

Observation 7005a5ee-e8a7-4078-b940-90adc3152e70 · outbound

This paper cites COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.705703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.705703Z digest=sha256:9368150b4aa64cdf4e13a02d5117b7b037872ebb0e1ea3431ffbee5da54aa78f

Observation dad5dba8-74ae-4f28-932a-592062ea71e2 · outbound

This paper cites Kinematic- aware prompting for generalizable articulated object manipulation with llms.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Kinematic- aware prompting for generalizable articulated object manipulation with llms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:01.141533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:47.802501Z digest=sha256:ddba46c0ef0d880d91fb5305ec45308b8f3105071506d94e1854d2210f408a60

Observation 8c35e3b9-4873-4117-a821-d10f28fa2b16 · outbound

This paper cites MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:47.950909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:47.950909Z digest=sha256:44ac4e99979aae0dfe725a1180937825d06006e94802b0318b2f3c303cceddad

Observation 6dddf33d-15cc-4694-939f-d600362909b7 · outbound

This paper cites Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.084439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.084439Z digest=sha256:f06cc0f0e088d3ee05845ac4e3e748b1734d1383c795138fa317512572fae76d

Observation 597e2a01-7034-4650-b63b-32cafef1ea61 · outbound

This paper cites Robotic policy learning via human-assisted action preference optimization.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Robotic policy learning via human-assisted action preference optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.230494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.230494Z digest=sha256:c5e4d144c41f1a453763d19535d47daf233f7bf91986d68dbacb3c55bed4d65f

Observation 92e6060e-9b8b-44a7-a2c4-33d3bfb0e455 · outbound

This paper cites Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.991650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:48.361825Z digest=sha256:7a03566a23bbecdc60bef1f54a6ae60672d3e17c756d92d347833764332920fd

Observation 85cd5166-56d9-467d-ae2d-9aa866a98f5e · outbound

This paper cites Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.487515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.487515Z digest=sha256:d2dc79392219f6c5bd1194b022b496f990a0ef7e814e600af308bcc905227c9d

Observation 1aa81f56-cba4-40c7-b47d-aabcf0641517 · outbound

This paper cites Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Spatial-temporal graph diffusion policy with kinematic modeling for bimanual robotic manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.854828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:48.605623Z digest=sha256:f0a0fd53b66f8c93b5887b6797db5b508630d116e016e8427ddb3eec0db2d70e

Observation 1af8d243-7c3a-4325-b55d-d0415085ca79 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.688344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:48.729701Z digest=sha256:d7342bafc620e95cd38233f77e08d35a6ed4018eeb3738089fd38195dc35e4a4

Observation 33ca2d67-8cba-48d1-8c09-e441a74ddea4 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.892757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:48.892757Z digest=sha256:565e56701e4c0385c73e83da527d3010d1eac348547605427c5935d1fee6a010

Observation afbe1c2c-f490-4a80-a099-310c6b6a865f · outbound

This paper cites A dual process vla: Efficient robotic manipulation leveraging vlm.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model A dual process vla: Efficient robotic manipulation leveraging vlm

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.459634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:49.081807Z digest=sha256:735e89e9805a3c4749a236fe22e4d0641b392820d41020f6b3f8b3b7c6137131

Observation 0593f14e-de49-4900-97f9-ced694cb1150 · outbound

This paper cites HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.257537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.257537Z digest=sha256:c22dce61507b842ba9701d96ee20018b09239f729a61f210683d0f76320c6fee

Observation fbc7f4d3-eef8-4e1c-9eb2-c8f2cbfe781d · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.385522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.385522Z digest=sha256:ce6e1f7dca6415ab73830073638832a4085b81567174db98e12c2a806126df24

Observation 5ca8f132-157a-4cc0-a0b7-b44fbf22ab28 · outbound

This paper cites DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.511193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.511193Z digest=sha256:351238bd211c3b77e15c2d8ea3dce55827d7d1509b8ae69caa23ae78e266a66e

Observation 3c1db6a1-da02-49f4-a743-2c29061bc1a0 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.614480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.614480Z digest=sha256:abf0665746b3d28c30fb74c5fdefdf29eefed0359a86a25af77af24183525384

Observation faf03009-fb2c-461a-bd4d-14ef462cf812 · outbound

This paper cites Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.760140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:49.760140Z digest=sha256:6d1cb9d3601406016c5bc9f15caea957753e71849ccff319fe6c16b3491d9f8d

Observation f3b374f1-6e31-40ef-983a-45e2b9a5fa7a · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control, 2025.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Helix: A vision-language-action model for generalist humanoid control, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.206921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:49.915193Z digest=sha256:6d652274dc2ae61aa8a7b21adf190d714966c7f4669ec3243f0ceca8219758ea

Observation 32fb4798-d4e5-4b97-82c7-86ae6160cf74 · outbound

This paper cites Flow Matching for Generative Modeling.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.042676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.042676Z digest=sha256:d9ef1e99302a96a4abd95b64ab93815f4503fd4d5dbfa6194e27fe25284b21ca

Observation e36e8996-aef0-42d3-9da8-5e096f4a255c · outbound

This paper cites Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Beyond Optimal Transport: Model-Aligned Coupling for Flow Matching

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.178954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.178954Z digest=sha256:6c21e167b24094e5cf95f2dfce71744ee8a41a555cddf2e41add20073df3ec8f

Observation 4e2d2e85-f720-40dd-9f4e-71daf55ff675 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.325026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.325026Z digest=sha256:c04abe3f6bd910e74d7ad67643ebeaa01f4f01e6f86f8db28f536fbad77065f2

Observation ed099548-4729-4806-a43b-4047d2b92ee5 · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Evaluating real-world robot manipulation policies in simulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:34:00.079764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:50.473479Z digest=sha256:c8f970137dc78c11d3fcba48bf699cc11c4ccdb19ca39d12fe0471a954d7112b

Observation b6980a7b-23da-4fac-9da9-898329f2d91a · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.601029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.601029Z digest=sha256:f45a3f4987b57c3adf5d247cb0693609319dd40d634b37461f6112e1dfb1ccb5

Observation 4661f28f-b4fd-428c-8cab-dd0ea2467d6c · outbound

This paper cites FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.714630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.714630Z digest=sha256:d753dc20a7efddf672fda7d4d396f3c95a33a866dff6585bcc80d60675c24e7c

Observation 257b1298-91e4-4d6a-b99e-c8e12cacef87 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.849364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:50.849364Z digest=sha256:16a2b5b021fd2db5294eff5184728a5a1a2b57775d159fb67cef2aef6eb0b525

Observation 15194271-763a-4a9e-ae20-58bc0171772c · outbound

This paper cites RoboMP2: A robotic multimodal perception-planning framework with multimodal large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RoboMP2: A robotic multimodal perception-planning framework with multimodal large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.951725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:50.961129Z digest=sha256:a34eaf011b64deb3d5ff758dca740da6a3562d42ce1296387dfd2935d995da2a

Observation 6f71dcb0-3c70-43a6-b824-801f1038533a · outbound

This paper cites Universal Actions for Enhanced Embodied Foundation Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Universal Actions for Enhanced Embodied Foundation Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.178608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.178608Z digest=sha256:c538f417509799b3788aa3c9ba2485345c6170b1160fdd982b45e6011dc8bc06

Observation 98106cf3-bab8-479b-bef4-aedf443dc267 · outbound

This paper cites Learning causality-inspired representation consistency for video anomaly detection.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Learning causality-inspired representation consistency for video anomaly detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.715736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:51.301356Z digest=sha256:bb951342ccbefa10efcd7db401dc478ca9ec5b54fba21b66e42e1d889cf547f0

Observation 7987215a-b994-45c6-abc3-f9d0c0442d99 · outbound

This paper cites Pali-x: On scaling up a multilingual vision and language model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Pali-x: On scaling up a multilingual vision and language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.525438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:51.416783Z digest=sha256:2b40f57f0d0b63d3ff96bc96d11a0ead1e341817bed309c47a9d1fcc351d2a1e

Observation a686961f-ed9d-4b19-866a-bda44f39afa6 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.312277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:51.548815Z digest=sha256:ffda652984b58d68cb2b8757b42a1c3067f19a41ce0a7085244353bb3186c424

Observation f30ddaf8-0366-41b9-ac3c-a4e833fb91c2 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Open x-embodiment: Robotic learning datasets and rt-x models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:59.120220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:51.664459Z digest=sha256:0160feddf094ce3fa3e8ab56da2a56dcfc8764245b39de5cd4cdef720246cb59

Observation bb30c695-ea59-47f2-b80d-691d8ebc69ca · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Reflexion: Language agents with verbal reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.772547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.772547Z digest=sha256:a81bd7fb63d3b81dac9d5a77e45a46763a877da315497d2efb412dfdcc29af40

Observation d0c08af9-ddd5-4a01-aed7-1077f92b290c · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Reasoning with Language Model is Planning with World Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.884717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.884717Z digest=sha256:198567ae06dc5b2301eea0e5bca0ba52b8160cbc49c0250e06733ceeebf0c082

Observation 76800bd8-a57b-47d3-be37-85e5bc17e353 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.985432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:51.985432Z digest=sha256:d010f8224833b29bd630ff5c9da3f8b001ecc5d1b05639c9c9d32cad404938ad

Observation 232a0add-3c70-40f2-85a8-dc5ba1c1ae8a · outbound

This paper cites Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.124767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.124767Z digest=sha256:9ba7d2d86e9d688aa9b22c0c4191dbc282a3918b68c737726353be1e13d7d209

Observation c515e4aa-907b-4748-b0fe-ab70894abfb7 · outbound

This paper cites Mllmguard: A multi- dimensional safety evaluation suite for multimodal large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Mllmguard: A multi- dimensional safety evaluation suite for multimodal large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.909259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:52.282259Z digest=sha256:a33d4bfad0779a294ff88f207560331ece40173f5bf815125d3fa09543b18b72

Observation 51c50869-8ef4-42d7-9746-4530cff6eec3 · outbound

This paper cites SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.387699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.387699Z digest=sha256:0069b323b397739ae55c0ccf2d96ae732606a2ef51380e9babd1fb0f6097b239

Observation b14d58a8-8cfa-42de-b0fc-c8e86a6fd48f · outbound

This paper cites MorphMark: Flexible Adaptive Watermarking for Large Language Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model MorphMark: Flexible Adaptive Watermarking for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.513946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.513946Z digest=sha256:ce25e297fc6f806e98de5312044bf1f346553c8674e4c35c4b026aa7967afbe0

Observation 0ade5a7e-5cda-44de-b5bf-de1ce2217f55 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Tree of thoughts: Deliberate problem solving with large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.729436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:52.622187Z digest=sha256:f9ce71a89f786329eacedff5b7db716d809311d0573f542143d688eb138a6cfe

Observation 52dd79ae-1572-46ca-b6ee-a04dcf031ad7 · outbound

This paper cites AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model AlignBot: Aligning VLM-powered Customized Task Planning with User Reminders Through Fine-Tuning for Household Robots

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.754745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.754745Z digest=sha256:7914e8d708740addd2932b17edbb92681f81d94e6c29f798a55831288ea66ce9

Observation 56722f56-4188-4b25-aede-774f1ca05879 · outbound

This paper cites Sets: Leveraging self-verification and self-correction for improved test-time scaling.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Sets: Leveraging self-verification and self-correction for improved test-time scaling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.896307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.896307Z digest=sha256:2466eb03d9fe69028431abdc9a5aa17e39b34482af3d684edc4061db04cf4569

Observation 23c7015d-acb4-4e75-9d04-2bd052aeccba · outbound

This paper cites Interpretable Contrastive Monte Carlo Tree Search Reasoning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Interpretable Contrastive Monte Carlo Tree Search Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.995028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:52.995028Z digest=sha256:9702dc50ffc475d5a7fc18c56994aadc098a6ddcf575c10745fe5c6117b77a47

Observation 51468219-21ee-4e85-a594-c758e0825966 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.140367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.140367Z digest=sha256:8e28dcbc5842d9752b27fa26a36190c2d23e27e70ff12b223c44484a15308345

Observation b0acd536-3231-4720-905d-cf6adbc16476 · outbound

This paper cites Cascaded Diffusion Models for High Fidelity Image Generation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cascaded Diffusion Models for High Fidelity Image Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.244353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.244353Z digest=sha256:643a1eeb11d33dd84caacd523e8678f54635c733ee4cb3ced367e43fbc8b9285

Observation 55a1e608-4ed0-4fa6-bde0-4ddcb2a9a26a · outbound

This paper cites Revis- iting multi-agent world modeling from a diffusion-inspired perspective.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Revis- iting multi-agent world modeling from a diffusion-inspired perspective

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.327795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:53.327795Z digest=sha256:4656d846a4c1c7f75f102094b3b58c57ce93ff3f0b1fede37a909a89476933bc

Observation f73061fa-010c-4826-9dc2-865bbd20f8fd · outbound

This paper cites f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.228091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:53.420181Z digest=sha256:06d4e2ecf9bc00f7f36447983e00754bc21398a2ad86dfbf64c08dae1b97b94d

Observation d2f7f3f1-10d6-4d08-ab54-bf8ea252f4b1 · outbound

This paper cites Bring Metric Functions into Diffusion Models.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Bring Metric Functions into Diffusion Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.005341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:53.554629Z digest=sha256:47adf88309f6fb8340569615726c1c5c39214f8c80b9d9ba0f4bc8526e0b16ac

Observation 50763b71-7e9d-4562-b70a-c7dee4706c6c · outbound

This paper cites Spectral-cascaded diffusion model for remote sensing image spectral super-resolution.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Spectral-cascaded diffusion model for remote sensing image spectral super-resolution

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.526003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:53.717819Z digest=sha256:eb97f55d6ffbf9ccb62eac3a23efaba60c07735349ebf611c228fc1ac5d22243

Observation f51f89c3-d1cb-4027-84ce-d0d70f442a8f · outbound

This paper cites High-resolution frame interpolation with patch-based cascaded diffusion.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model High-resolution frame interpolation with patch-based cascaded diffusion

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.305583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:53.835285Z digest=sha256:769447cb47e516ca86f92009bd1011d3af3a50c9eb6a65495170871f15e0b2db

Observation c527b747-cc31-4b73-8b80-f6f1e271296e · outbound

This paper cites Cascaded diffusion models for virtual try-on: Improving control and resolution.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cascaded diffusion models for virtual try-on: Improving control and resolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:58.127619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:53.990234Z digest=sha256:98c77eae5fce09e3a7a6999e7052d1b13f055750f8806fd564de91ee8e323985

Observation 2c28b4f4-4442-4192-947e-bf9bac79cfd2 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.126278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.126278Z digest=sha256:f04f7ebaf67f4ea3b3ca769f1406eca6abe83e314876d0f6ee2ec2329511f254

Observation 05da854f-a948-4814-ad54-9ffdfb1cc9e1 · outbound

This paper cites Pre-training for robots: Offline rl enables learning new tasks from a handful of trials.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Pre-training for robots: Offline rl enables learning new tasks from a handful of trials

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.971727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:54.254443Z digest=sha256:be55e70fab877c91ae2ba0d485b435a403fcee23a06be4f8da6519b2e6694569

Observation 287d1e5a-a6e8-4633-a58d-c5f73ef9088f · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.814329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:54.384981Z digest=sha256:fa2dc0bd075fb17a7ec7414bbecc8319a59b9d55032dcb643a912d8583b887dc

Observation f5826893-a8b4-4a18-848d-bfba099c12f1 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.509544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.509544Z digest=sha256:a638e0dcbe5a92ff9b203a57f08373d28343c2587674fd874cf1d61d39b8c01d

Observation b73c3ab7-a074-4ed8-bac6-3ee1f980e446 · outbound

This paper cites Octo: An open-source generalist robot policy.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Octo: An open-source generalist robot policy

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.674777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.674777Z digest=sha256:4ba700f23b5cddfe0e5f090c8abd53079f3bf5c68de82d49aec66bdc007526a7

Observation 949e21aa-1443-418a-a674-d4b10016708d · outbound

This paper cites Scaling proprioceptive-visual learning with het- erogeneous pre-trained transformers.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Scaling proprioceptive-visual learning with het- erogeneous pre-trained transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.645007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:54.796683Z digest=sha256:211d059272fb8a75301d0231fde1f068b3ed3d0f2b63d8a52e310c16e54931b7

Observation fb8bd160-8b34-48c3-9899-1b76a00db3e0 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.889344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:54.889344Z digest=sha256:8778d4c3f1d828c64346f0f6fe4a2e5d1bacf464dba57395b2f67392f04acbf5

Observation 4ee513b6-aba5-475c-9d35-14fb20a3afc8 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.004692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.004692Z digest=sha256:a6c441debe154c824eac2441f2a002253082df3bcb40226b3d5a37719017801e

Observation ad6b665d-616d-44f5-bdca-262e55c9930f · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Diffusion policy: Visuomotor policy learning via action diffusion

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.157444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.157444Z digest=sha256:8e095be447d9756298c6a346b95b1d8d34798bc2315b8cbe7df71684fc2b596f

Observation 7ac1ebce-6d32-4038-8412-5bb69cd72bc6 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.223798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.223798Z digest=sha256:f7ad81ac506bdef712f2868c057343ddad05756574852ef420328211685c5f7b

Observation ae1bf471-dc67-454c-9dff-58f2d586f78c · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.329461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.329461Z digest=sha256:e3107d8f8d7d6815bb380e8f03b23b4fe2661ac9f6ec0ffdae4bf6ab0565267f

Observation 1791c3be-837b-4f62-b6bf-1553d4807b86 · outbound

This paper cites Continuous control with deep reinforcement learning.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Continuous control with deep reinforcement learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.428373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.428373Z digest=sha256:a7e00c2eabab6de1932723591599b81cefa9cc777f62615dd90b78a361b50e03

Observation 9b18831f-3bec-46ab-82ca-e17095e176ce · outbound

This paper cites Addressing function approximation error in actor-critic methods.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Addressing function approximation error in actor-critic methods

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:33:57.417704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:33:55.564589Z digest=sha256:349ba31bfeed8ffe288449a537610990419d9854f562a80a5be1d4c7e31b43fd

Observation 5d8e87df-c426-4323-8554-e9c8d0a2fea1 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Hume: Introducing System-2 Thinking in Visual-Language-Action Model Soft Actor-Critic Algorithms and Applications

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.651566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:33:55.651566Z digest=sha256:ea7fdeda80152389c1990c5a471085ef5a87344717822fcfa6498a3bb328cdba

Pith citing papers

Observation 7f957329-1602-4060-b1fb-d8301b5ef22a · inbound

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving cites this paper.

FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:19:42.874884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T19:19:42.748573Z digest=sha256:eee7ecc917d4cfff44c1144af0046a4ecc142babab04dab11f3b3ee7a7f6c981

Observation 9943f6b5-3f45-44ca-a277-c0385c11385e · inbound

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface cites this paper.

Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.315435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.315435Z digest=sha256:d50520d12ccc79bcf41014f3949863c1a33292b6917579bdcf37e9fa390eeef1

Observation bf1e6f74-a5fa-4853-a89d-9d6deb14871f · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.581802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:5d788507456f3c5cb2c6eb3b5baab14dbe8297f42afefde3e87378c8f803d3b8

Observation e6c42ed0-3a1e-41ec-be9f-2ac5e9d0b556 · inbound

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions cites this paper.

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:42:47.069003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T12:42:46.978148Z digest=sha256:d56101fdb609c6fdf99fbaf68d54c4640894fe1faec59d7c549169ef0a97bf31

Observation 3edc250c-f891-48c5-987d-ad183462ea08 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:39.902748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:c213d9ba25753c67a6c4d10394d38a06f5d574969bb5110a2abcf79ef7e8cf16

Observation c5a28893-c809-42c3-bf40-11c38ea3f232 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:11.780961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:11.780961Z digest=sha256:a02491bd38b3d1998b30d6233d352f5d23a19e2d1bf7bbcc99b81bbf8ddb7808

Observation fa07afbc-06fa-432a-b218-afd5a0eb7310 · inbound

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation cites this paper.

VGAS: Value-Guided Action-Chunk Selection for Few-Shot Vision-Language-Action Adaptation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:00:26.042642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T07:00:01.741166Z digest=sha256:a8354e85b49570921fcf94a17188cfebc9f5614a47c90b3267c0b800761cff0f

Observation 150d728d-7652-4a3c-bc24-46f9e1ba84ac · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.030889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.030889Z digest=sha256:2c4b464c6785b08f593c3899df4451a38fc0116fd9d8abed7739d041d4941563

Observation 73a9281f-48a8-4fa0-a5d9-b79e4083dc45 · inbound

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data cites this paper.

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T12:10:54.741769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:10:54.741769Z digest=sha256:eeff6550b2024d5a6e8a1dc65e5203806a74528b59e779d1281c80f8773670a2

Observation e773579a-6a7d-4e0d-8a2f-5b3e8483c3fb · inbound

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models cites this paper.

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:50:37.167968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T12:50:02.670780Z digest=sha256:3f4a7e9f5d707052bddf7dd21242085e1ed9bc571055e7c16edba85b7267efcd

Observation 8deee7bd-f6c2-4f84-b333-13298a69cef4 · inbound

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving cites this paper.

Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.527563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T11:56:40.836234Z digest=sha256:9dcce89cdb54c3c28c30b2fae248ccc44564df43368815d158b96781557dfa33

Observation 635e0017-9b9f-4d17-b2f7-0e961982ed91 · inbound

Spatial navigation in preclinical Alzheimer's disease: A review cites this paper.

Spatial navigation in preclinical Alzheimer's disease: A review Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-13T19:53:53.219607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:53:53.219607Z digest=sha256:dc09a538d0034be78a184f81db3a0ae7f3112b6eb197e1bcf42e78e39f929126

Observation 5241d33f-930f-4b74-8528-f7da22a0e8d4 · inbound

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models cites this paper.

UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:43:19.018696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T21:40:30.771335Z digest=sha256:cc503c9b58785a66e5386a98a68e50e009f9d85145458865c50994ca803aa05a

Observation 90f67b5f-aacf-4d18-a407-f2dfa1a89af7 · inbound

Deep Image Clustering Based on Curriculum Learning and Density Information cites this paper.

Deep Image Clustering Based on Curriculum Learning and Density Information Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:03:28.527131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:03:05.442554Z digest=sha256:4552059e6be4c2c154d708efce875a4464c71736826ef91815eed6d5fc67ba9b

Observation 28a5c88f-922d-4a9e-afe3-e79a16620ddd · inbound

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models cites this paper.

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:01.503826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T16:57:52.595165Z digest=sha256:6af07deeb6beeb89a92155d57a5bccfb53b624928aed1648af243b5194fe07c3

Observation bce58944-e575-4e19-aeeb-2f96c4ab1f0d · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:04.809662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:14:03.865035Z digest=sha256:dc291daaac615b90ac08233dcceea15aa65e586fb6ff1a5014e462b79fd12e0e

Observation c66e8685-deec-4791-a0d1-cde96d09ed9f · inbound

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery cites this paper.

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.064730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T01:10:48.111387Z digest=sha256:9fe1271720df3b1ed9f4befb20e11429ec793804d1c6466ef42ea23b8f298230

Observation c8b0042f-638e-4de8-b0bf-1858489ed884 · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.358813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:1bfc04f8ca12b0124d31f41b7e8b115d2937a2113e9197519cc9242cb1e52131

Observation 9237db13-56b0-4f9c-8a7b-61c5ff07d408 · inbound

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control cites this paper.

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:30.361693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:47:16.205357Z digest=sha256:9f9d721a7152cabdca1c11eb426b66ce86fa9bbef80171c82ef2bb733e9b0fc0

Observation 2913645b-6fa2-4e10-8c4d-453eb1512fbb · inbound

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control cites this paper.

Hyper-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:25:09.759639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T00:18:14.307099Z digest=sha256:b80d4c77eccc46ddde97482476b7a4a111a8e998ba7130382526e2dd362fc109

Observation 30972e20-377c-40dd-b05c-ff820df2d8c4 · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:12:18.267361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:7bbc52ce8432992bcdf2c17933cd22946b1db148a72e1caaf8c4cf47cc76ae0b

Observation 0c5495cd-9e10-45be-8883-67a3eef00e56 · inbound

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA cites this paper.

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:07:23.737640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:57:25.010339Z digest=sha256:5282baf543474ce47db65e85eece0f08ef71e8d17db6bfe31ee1d51079e43d7c

Observation de9b08b3-5e48-45e1-9deb-259b4f49c5da · inbound

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA cites this paper.

Q-VGM: Q-Value-Gradient Matching for Off-Policy Reinforcement Learning of Flow-Matching VLA Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:13:46.004799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:13:46.004799Z digest=sha256:9b327883c7928cff2b52403a5f4f0aa82958886df6f36cf4c94e4bd7ee297b0d

Observation 5e3fd18b-156e-424d-98fa-ba4770f928df · inbound

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space cites this paper.

PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:58.507228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T00:57:58.312567Z digest=sha256:c69c8dcd1e20fd1985ecb4e616d9a1e4e0154a41e4bdbce01612adcc2045e874

Observation 53e4a2ea-4c7a-4e32-b782-a1ccb57258a1 · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.647699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:1e43dfb07afdd2707dae05e611013805070f14d81d438a023fbf76312473e0c8

Observation 8cf1726e-96e8-4ac7-8326-8919eb2502a4 · inbound

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation cites this paper.

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-27T04:40:32.814702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T19:29:00.117285Z digest=sha256:7c736411baf619c3d5870fa01b3dea2cec856d76b203510f7ef110ecec7b5f3d

Observation a96ac403-2fb8-43bc-bf7f-951da40ff893 · inbound

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation cites this paper.

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.756964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T04:48:28.868562Z digest=sha256:5b4527542a78348f6242c189589b5c3f8550f90f69f6d7568cca0a3b68f38936

Observation 1c1c081c-dc97-461f-81d4-34b61c42e28c · inbound

Recursive Self-Evolving Agents via Held-Out Selection cites this paper.

Recursive Self-Evolving Agents via Held-Out Selection Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.572040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T11:13:15.923479Z digest=sha256:01a4f76f302ae39aa9852e0fecccaa5c7202af9d636856cd367d122eb668d37d

Observation 6d0ce5b9-3a94-40f1-b56c-203336e675e1 · inbound

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning cites this paper.

Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.827744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T06:56:45.197407Z digest=sha256:eebaafc443c699aca62650b67e4e1baf619103804dab11eb919e10fe6e0d3b0c

Observation 984e3d47-f327-4463-81e0-474103ed5288 · inbound

ROSA: A Robotics Foundation Model Serving System for Robot Factories cites this paper.

ROSA: A Robotics Foundation Model Serving System for Robot Factories Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:16:52.678721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T11:11:39.584755Z digest=sha256:7236f53e6d14551cd0b3eba3c19153f6f3083735bacc2334227e087cd39f492b

Observation 67504498-1254-428f-91dc-015b7465a512 · inbound

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models cites this paper.

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T00:12:42.173815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:12:42.173815Z digest=sha256:9b7c3498cc6130e28e1452998356b5ee51cfd2d79b149a87633c4ab5344f59a6

Observation 73f704b0-a5e4-469e-b3e9-0953aa79b3a5 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T12:10:21.115628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:10:21.115628Z digest=sha256:ed51e1fb6ca23a328fda887f352c1d0863aa6e47b0b12ae59a68178b31b94527

Observation 5efd68fd-27c8-462e-ab83-5827f5d3a342 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.336447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.336447Z digest=sha256:5f5a8fa5d6e3200df0b46481a9ce32d495112649c49536aa5d7c0d0c362a4fb8

Observation 9c83eae8-90ec-454b-8d38-30cba65811db · inbound

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking cites this paper.

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:30:57.030649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:30:57.030649Z digest=sha256:c7150bc66948a3a64939eaf05b806e2f9d55d35899bcb6aee0a6723a5a7e7176

Observation d8f48e30-f03c-4298-aa2d-b3bd694feb13 · inbound

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation cites this paper.

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T19:56:57.191876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:56:57.191876Z digest=sha256:3f379773edfb4a3f7cf314f2484533dac714aae75dd1c4ad92580ea3f37b421b

Observation c3761f79-4d71-4820-b142-ab1456f4d407 · inbound

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection cites this paper.

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection Hume: Introducing System-2 Thinking in Visual-Language-Action Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:57.174397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:30:57.174397Z digest=sha256:d7fea81a706306019f140e9eeed5a9af95fcc4c3635ed4512d7c7fcfbce71194