Pith. sign in

Paper Citation Record · LEDGER

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

As of 20 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2607.07608.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07608 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T05:29:36.753268Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:48:28.878006Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:48:33.298230Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact26
  • verified fuzzy45
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c52b51d-4e6a-4e05-a87c-39912171f975 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.215054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:93a1d5bc1b5352598ab768589aa288d593c57bf3f9f16437aa279b719c594071

Observation d5ef3b4d-2d1b-4631-a151-577467e1d942 · outbound

This paper cites Openvla: An open- source vision-language-action model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Openvla: An open- source vision-language-action model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.892196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:bce4d2ad1552f8b0f612d19714a2dbff8bf73a78e340c12b6614b8855033293b

Observation 10ef95d4-f89a-4d43-82f2-99abe57c854a · outbound

This paper cites Rdt-1b: a diffusion foundation model for bimanual manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Rdt-1b: a diffusion foundation model for bimanual manipulation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.885881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:559a7436065411ae5e6a0869cc26779f5ddfdd6aa4d5511872f19e121e8eb65e

Observation ae71e712-f169-47c0-a38b-87a583995ef5 · outbound

This paper cites 3d-vla: A 3d vision-language-action generative world model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation 3d-vla: A 3d vision-language-action generative world model

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.884064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:47cafa45429cf2797dea5c2eaae036be467705b4596bc9f9381d546776b75ba1

Observation 7ea1439f-7640-420b-8de6-faeb8f0c0876 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.205303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:7ad6c8cb8969dbff760d477103c83412687eeda7dd834ba6ea2ae1a4215fb145

Observation 2f713d6d-1511-41af-b81c-79089e44ad12 · outbound

This paper cites Qwen Technical Report.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Qwen Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.149913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:f782f41c46b89155d1a681f6fdcdaf38ac759cdc09c9f916d67ffcb1737169a0

Observation ddb3cb7f-702f-465f-bcac-9ff60e8e383f · outbound

This paper cites On scaling up a multilingual vision and language model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation On scaling up a multilingual vision and language model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.882065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:e271b0d90c2ea2998318817f4993b0eaedcc28d7a31682f035c2a5e80c261043

Observation 68a23db1-9e35-4523-9769-6c36e233586c · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Diffusion policy: Visuomotor policy learning via action diffusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.887756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:6dff61bd78b6c1995d29653f5605be9a18a470aee69b98f37b8ae2a65e6e6ce8

Observation 76bbf25e-0765-4320-8738-a6e7b2cd5a19 · outbound

This paper cites Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.889950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:6c0bc7c9b5e615b1b7399596bc8ea0a18e101448febd4d24c8548896ebc5a5cf

Observation b21a8162-0191-4606-9f90-9b44558b8b7d · outbound

This paper cites Octo: An open-source generalist robot policy.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Octo: An open-source generalist robot policy

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.877560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:d09f2d30b3a82480cc524c3d0214dcf0bb2f8ccea44a7a070e50b763c2f50399

Observation 732f7e62-24a3-4830-960e-7ba329b89ff9 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.873488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:c44179e04ce34bb825cf12d4f9347cae5d9590648a9d6b51b90b345df2f5959b

Observation dc9d015d-be52-4f9b-b0f8-4be30e59484e · outbound

This paper cites Droid: A large-scale in-the-wild robot manipulation dataset.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Droid: A large-scale in-the-wild robot manipulation dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.871324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:ab820177f931451ea2aa9b8c1d0a8a65e73602bf0c67199be03daf10f8271e95

Observation 45149a7f-21de-4bb1-8eb5-e72e33119dcf · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Bridgedata v2: A dataset for robot learning at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.875592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:d1a2d8d8e6764396118cbeff252a5981b34e200a9e3024a0a5a4e9c53cbbce39

Observation 836d59b8-cfd3-4314-9482-742fd4ae5dce · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.247849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:6a06be36763f995d7dd4980581b60e9943719531dbc80a1554fe94e4b7ba14bc

Observation ad38fbb4-ca87-43bc-bb0b-8341a3aa9916 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.156824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:56215c86e8a96d23873efa1c67a1920a8bab8f97231e030b4b627ed1f33d6d18

Observation c71cc377-b180-4bb3-ab04-fce1ed191065 · outbound

This paper cites Language-Guided Object-Centric Diffusion Policy for Generalizable and Collision-Aware Robotic Manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Language-Guided Object-Centric Diffusion Policy for Generalizable and Collision-Aware Robotic Manipulation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.220052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:7162c08babe726caf909b4230495b83538f527761197a5224936c45490dfb639

Observation 171e4a30-9b9a-41c0-99df-8d821a46def3 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.RSS,.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.RSS,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.866669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:c912116900e79e1f9dd0bdee27c4c3b00773d8c716e59aedd532e9ceacc8314a

Observation 0dde5dc2-ca81-465d-b62b-251bf7a85223 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.254012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:b9771736117cd9661c8f1ee761e15ef23702ab1b6483076377b84b98ec9bf260

Observation 9a28d0b0-29a6-438e-88b5-fc6350fbc904 · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.869379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:774c76f0a6a94b190f7c4b21f2aa95d7ddbaee275130a0c595cc8fc380522e3d

Observation b8942b76-b11a-4ffc-a86e-747b3203cf0f · outbound

This paper cites CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling, October 2025

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:36:01.233117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:36e7d09f6ffc8204ed108b89d678893ece7358b1bce40c539768ffd29a772db1

Observation 60fb7c14-391e-44d2-bbf3-2902234e43dc · outbound

This paper cites HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.223722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:24a038c19fefc6bf0351f1cbf746ab6c2d4daab70bbbc7877e1034cfb2a5588d

Observation 8e81eda7-2bc6-47ad-8982-4bf3a8b33b43 · outbound

This paper cites Resolving state ambiguity in robot manipulation via adaptive working memory recoding.IEEE Robotics and Automation Letters, 2026.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Resolving state ambiguity in robot manipulation via adaptive working memory recoding.IEEE Robotics and Automation Letters, 2026

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.880111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:c5a5940a6078dcf977697748d5acd9b439bd27d9ccceb34f31682e8e578b888d

Observation 43cab735-4845-424e-a9f8-be293a75844a · outbound

This paper cites Memoryvla: Perceptual-cognitive memory in vision- language-action models for robotic manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Memoryvla: Perceptual-cognitive memory in vision- language-action models for robotic manipulation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.894901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:88e5ef363acc21bf46d1eb7f7ba8680fc1d5d7145b8e48f54cc13197bf894b09

Observation a15068c5-bfae-42dd-ac20-70ba63b8c7d6 · outbound

This paper cites Memer: Scaling up memory for robot control via experience retrieval.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Memer: Scaling up memory for robot control via experience retrieval

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:36:01.240704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:002d49997f1543bca504b8a38cd4688c8fb51e4f38ad583d3b474f62137af904

Observation 7673f631-5362-407d-be7b-42e6344a9c61 · outbound

This paper cites Global prior meets local consistency: Dual-memory augmented vision-language- action model for efficient robotic manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Global prior meets local consistency: Dual-memory augmented vision-language- action model for efficient robotic manipulation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.864791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:ba25e0340aebe2c1a06837566de9381f77715d1856487ed687b02599032d91ea

Observation f5d5fd2c-7633-470e-8f9d-032161386b75 · outbound

This paper cites Monet: Reasoning in latent visual space beyond image and language.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Monet: Reasoning in latent visual space beyond image and language

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.897015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:56472b9b25b1370337d7acfb5245e624a2445cb1858820a04dc06c2a0e2a03f4

Observation bc0e9a92-6160-4101-b9f9-71b992c0132d · outbound

This paper cites Machine mental imagery: Empower multimodal reasoning with latent visual tokens.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Machine mental imagery: Empower multimodal reasoning with latent visual tokens

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.820533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:baeee23a0ab9a4faa8fd534a18a00d73afaf31d27c0154dc557b0338594643c7

Observation d525808d-e432-411c-b45e-981637381a37 · outbound

This paper cites Latent visual reasoning.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Latent visual reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.815107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:1c18ef47bd829f9bbfcd7fc8d8c5a0b22d0596c79d3da592f9e94fdf27caaf2f

Observation ce34f478-4072-4502-9607-4f688fd9b487 · outbound

This paper cites Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.182613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:025725a27fe57a3bd06477f43215fc74d03b6648af59a59a9cfcfa1a24ebfaf7

Observation 3bd25b8b-617f-4680-a44d-1f06e6666639 · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Appagent: Multimodal agents as smartphone users

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T05:36:01.244436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:73e2998c85fdf7a184ee47988582c2df55bf14a8137361d801fd85c15b32dcee

Observation 2e8fcb0e-77bc-4122-a9b4-846a6de8baf3 · outbound

This paper cites Vismem: Latent vision memory unlocks potential of vision-language models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Vismem: Latent vision memory unlocks potential of vision-language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.825378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:9ed31c7b9c7f1546d396b319e848e318f47594b17390d7e6d538aa1a7c858aaf

Observation d35bfe89-30e8-49bf-8d1d-4373133994e1 · outbound

This paper cites The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.169585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:11fa5c1e894f09ef94bc695fdeb433351e82c78fe3b1f18d63b51058045e9175

Observation 468bb18f-4c58-4f70-9e21-cb2ffb6865dc · outbound

This paper cites Rt-2: Vision-language-action models trans- fer web knowledge to robotic control.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Rt-2: Vision-language-action models trans- fer web knowledge to robotic control

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.801736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:f0a2b6aa7ef769bd71f346ca8bb04dc750b3e8987d6edbf7e076e11c802d92fa

Observation 414f107c-9584-46bb-93ef-59ead2726ac6 · outbound

This paper cites Scalable diffusion models with transformers.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Scalable diffusion models with transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.798731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:99c9eaac32a8a425ccc00820f6c82174cafd302a64bda0295f37417439f5b049

Observation 6922aa31-d139-4c16-8f1b-f80ed48d36b5 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.827859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:bb044f7585c6569b116fdb302432da50aa5220fe8dfc48dea8e4eecb4c22121f

Observation 022a85cd-dd5d-4a90-af16-1dd47b8f63fb · outbound

This paper cites Evaluating real-world robot manipulation policies in simulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Evaluating real-world robot manipulation policies in simulation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.812848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:eb959e53e3ef8238a340b9e60c0098bf02e3cd1c6123f47afe6ae05584327d0c

Observation 0ce8fed2-df73-4d59-85c3-c9a8c74bf296 · outbound

This paper cites Manipulate-anything: Automating real-world robots using vision-language models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Manipulate-anything: Automating real-world robots using vision-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.807300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:42131a30804a25b46aee1f33c6e294b1c7eae2d4ecc8b46a3c3f9b78403c7bc8

Observation b74ff13f-9e9b-49f5-9cdf-c1fe852a794f · outbound

This paper cites Rekep: Spatio- temporal reasoning of relational keypoint constraints for robotic manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Rekep: Spatio- temporal reasoning of relational keypoint constraints for robotic manipulation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.823039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:f727a2ee939deac001949e16a72b6c6484ccb7d9759295e33dd0d7f6f3270a88

Observation d1e832fc-aa8e-4cca-be7a-a2c72d769426 · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.229345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:34baeae1b50efefb67838e27c804b926ecc3bc72cae19f3e4e864053fdaf16d5

Observation 340113ce-fed4-49dd-bc9c-ebb8b0558cbc · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision-language-action models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.862692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:ace14dc5c9877d7c2170ee0df2d5f603419ab01aae2ded668ebc7253033f1fcb

Observation a648fa49-288d-4f36-8e13-b005c59b38d2 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation WorldVLA: Towards Autoregressive Action World Model

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.236548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:42bd14f19c80754fac5676984337505253725e572acfd556f0efd61469bca78b

Observation d0a44ef8-cb7c-4f95-afad-eb9faee15b42 · outbound

This paper cites Unified Vision-Language-Action Model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Unified Vision-Language-Action Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.201663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:e6d4451d6152c87e30795844135817145c6ef9a8253217ab09fc2f0f960d22e8

Observation 0816f57c-57e0-4f93-ad67-b58a0d26215a · outbound

This paper cites VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.257153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:d0c1596e3b44319232ede36898f7db8f0a3bf755a0a2eecc0613bccb7b02f350

Observation 3c8c5ef2-b0c7-4663-bdb5-eb449eaa6a47 · outbound

This paper cites Spec-vla: speculative decoding for vision-language-action models with relaxed accep- tance.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Spec-vla: speculative decoding for vision-language-action models with relaxed accep- tance

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.854027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:3e9249176223d4bef956ea330d0fa6b2a3a87ec505bf510f5d1b7fd6b5a4207e

Observation e05629fc-de17-4052-ae1e-542a4e55aabf · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.209803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:082e059e45962c3d94ddd0e4793e04dbc81b9fe448220ae241266f2271a8f3e4

Observation 05869d1f-7bcc-4ac3-996d-b62b8260076f · outbound

This paper cites Mole-vla: Dynamic layer-skipping vision lan- guage action model via mixture-of-layers for efficient robot manipulation.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Mole-vla: Dynamic layer-skipping vision lan- guage action model via mixture-of-layers for efficient robot manipulation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.858596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:fa13b742f962dbb732f5531276e444dcc704e7fba01d6f665a83390bc258c210

Observation 6288911b-8a21-4a87-b61e-67126167e164 · outbound

This paper cites Bitvla: 1-bit vision-language-action models for robotics manipulation.arXiv preprint arXiv:2506.07530.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Bitvla: 1-bit vision-language-action models for robotics manipulation.arXiv preprint arXiv:2506.07530

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:36:01.198163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:71b0514f959970cf2ce948c7f62316226143b3298a5aaa466e3961436f2df156

Observation dd5bc57f-a59f-497f-8820-b375cb64340c · outbound

This paper cites Dexvla: Vision-language model with plug-in diffusion expert for general robot control.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Dexvla: Vision-language model with plug-in diffusion expert for general robot control

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.795515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:84f2429b600f2df79853c905e54eb8f4620c1855ceb2efc28646db1a5686920d

Observation 1989f193-05e5-4703-a156-ca23d2df6fd4 · outbound

This paper cites Vq-vla: Improving vision-language-action models via scaling vector-quantized action tokenizers.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Vq-vla: Improving vision-language-action models via scaling vector-quantized action tokenizers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.810323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:995c9a65dbbeac3fd09344eabef08d3f17dc0bceabc9cf5466c1842423b8175f

Observation 754d2efd-bc29-4b6e-989d-a66d06027dc0 · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.178426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:393304ebc9fbb190928968da5ead864990c0825764e9895cd11aa349c59fe028

Observation 5f9edbb0-e6c3-449d-b24d-34af7313939a · outbound

This paper cites Towards long-horizon vision-language-action system: Reasoning, acting and memory.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Towards long-horizon vision-language-action system: Reasoning, acting and memory

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:36:01.164800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:c4fc543cd6e3b0cf76ea329c76dd0e3fd0a830f3f93cc4d560373e38ef46b29a

Observation fa6721e3-beca-4f5a-ae9d-29d63d420c75 · outbound

This paper cites Vision-language foundation models as effective robot imitators.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Vision-language foundation models as effective robot imitators

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.842258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2c05ecbbddc49538796f065363d05609633f54c81011af05f4c7fede3c93e9b1

Observation d409a3bd-c294-43fd-a3fd-76e47311ec3a · outbound

This paper cites Tracevla: Visual trace prompting enhances spatial- temporal awareness for generalist robotic policies.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Tracevla: Visual trace prompting enhances spatial- temporal awareness for generalist robotic policies

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.847013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:088cfd1e5b9528edd20bfd9c6745f4db196a6b4ce326414680ea167ddba202d8

Observation 4df80ad6-5bf3-41ef-b3c9-3e6ee767774a · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.186238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2007b861c8aeb14c11414ea078a87e65ca442605cb54acd398a50853c946fe9f

Observation e7ef4cb5-c49b-4c92-a424-593a00401d6c · outbound

This paper cites BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames, February 2026.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames, February 2026

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:36:01.190060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:13a6e57793eb6416a3c26f2f68ac4fc37dbb130ac767027eaa4e7179ce6f8650

Observation bb75aa19-3024-4d5d-b01c-04c2d169cc8e · outbound

This paper cites Hif-vla: Hindsight, insight and foresight through motion representation for vision-language-action models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Hif-vla: Hindsight, insight and foresight through motion representation for vision-language-action models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.898934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:f9e231f2b5502d1e43eefda359c0b56da37ce5df0089e0e5886bfba9a3e9601b

Observation cb0975c1-a969-4700-8a33-d49983872e9a · outbound

This paper cites arXiv:2510.04246 [cs].

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation arXiv:2510.04246 [cs]

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T05:36:01.161110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:deeef12ba4504335d361f8a044efc72ee2f14605698284d5fed13aee0e9d254b

Observation 5ec692fc-4c66-489f-85ac-75d6e0d5b4ef · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.844664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2657c0122ce0f24bb200d78ccf6ef4e362130138274ef50183e645a8d487c1b6

Observation 485cccba-ec85-4443-bd0f-cbec72d7d522 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.Transactions on Machine Learning Re- search.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Dinov2: Learning robust visual features without supervision.Transactions on Machine Learning Re- search

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.849275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:7dbfccd86be858617244f22200a7be922a08d697786d04f15835881fbbc7c6a4

Observation b0082a14-6e89-444b-a1ec-8e845b226ab5 · outbound

This paper cites Sigmoid loss for lan- guage image pre-training.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Sigmoid loss for lan- guage image pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.804308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2469d28f5c1bc9c00a6ae91d6d43f061ba59b8a0f5af826c0e40737ff0c42d4d

Observation fd84c097-d777-44bd-b99c-826237ff192e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.193831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:b91788cc5b64a3757a06a84c62bfebe905f3da49e9742835b9184047a343b04d

Observation 0eed64a4-38fc-44e0-8b68-8dde505caf3d · outbound

This paper cites Squeeze-and-excitation networks.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Squeeze-and-excitation networks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.837429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:a735a33b49968f08f65fd0f7c332144d3de894fee715090a75e788af83f824b9

Observation 25f69a02-2527-47e3-ba38-e6ccf4f00733 · outbound

This paper cites Denoising diffusion implicit models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Denoising diffusion implicit models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.839873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:f8d94af690d7f96cd802baf3315cc70bbb17622b7c94ced96c6476c11750ec01

Observation 944ce056-85fb-4441-b63f-5fe283341e42 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.153187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2b840085d6515c0b0aa8d43455f0d84d9da9b63f197970a1c1b8ae267a2fbe31

Observation 8659a7e0-eaa4-4dba-8b02-0a0680087172 · outbound

This paper cites Magma: A foundation model for multimodal ai agents.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Magma: A foundation model for multimodal ai agents

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.830278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:695e7df0c62638483b350b29548c79df8b02293a3581f49069ffb1c4981616e0

Observation 500c7588-ba98-4c30-bd57-784ee93b88dc · outbound

This paper cites Dreamvla: a vision-language-action model dreamed with comprehensive world knowledge.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Dreamvla: a vision-language-action model dreamed with comprehensive world knowledge

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.832721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:009d885f95a9ad9a627e06590807e2461c0a93e84576fa7533bf13e5add885f5

Observation ab3c91c9-759b-48f3-9ad7-082467a4326e · outbound

This paper cites Thinkact: Vision-language-action reasoning via reinforced visual latent planning.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Thinkact: Vision-language-action reasoning via reinforced visual latent planning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.835088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:2f59f5844a2722b564651a43750ad5e449ac4921abd9ebfe303064cf0b4d9567

Observation 9a7d9ccb-ccaf-4f22-954a-b53cf0c2bd5b · outbound

This paper cites Towards efficient and robust manipulation via multi-frame vision-language-action modeling.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Towards efficient and robust manipulation via multi-frame vision-language-action modeling

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.851550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:500be12c26180f346c9431bdebb9c2b72627bef0fe3ed8443a6dca7a6a6d87cd

Observation 5d341f3a-cf3c-49df-b9ec-d08955f5e619 · outbound

This paper cites Semanticvla: Towards semantic reasoning over action memo- rization via synergistic explicit trace and latent action planning.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Semanticvla: Towards semantic reasoning over action memo- rization via synergistic explicit trace and latent action planning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.856231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:1d021e9a5e79ffc6fc164fd435aa5915932742f1c3814cf4745ff02a676d456b

Observation 0ed73f66-74cc-4d53-a293-3144688fec06 · outbound

This paper cites Universal actions for enhanced embodied foundation models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Universal actions for enhanced embodied foundation models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.860665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:498b0c012ecb0a5c5d2a4ddeecc71ea584a480ff9003aa9379ce2c65280c2a9a

Observation 563d5fa2-f223-4850-89cc-5d87bf820c2d · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.250815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:0e75597e0a610b78d2de6073cc733b7cc43106b97944a0e437307e47136bea35

Observation 0b4845de-6cef-4f39-a9e7-50ddfc43148c · outbound

This paper cites Fast-thinkact: Efficient vision-language-action reasoning via verbalizable latent planning.arXiv e-prints, pages arXiv–2601, 2026.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Fast-thinkact: Efficient vision-language-action reasoning via verbalizable latent planning.arXiv e-prints, pages arXiv–2601, 2026

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.818142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:eba95c6a6dac7bd54ce799ced55ca31ee8a6459970923027e5fd49cb246c785b

Observation 1a406fd2-1b25-4b6d-9d8b-a583830b5294 · outbound

This paper cites LARA: Latent Action Representation Alignment for Vision-Language-Action Models.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.174758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:052acd5b7e064708f10db0ab436e201de645a32abfc1e9d973d342f80cd11db7

Pith citing papers

Observation dd020876-d68a-4447-9842-82d6d3bf1c40 · inbound

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory cites this paper.

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:48:33.383515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:48:28.878006Z digest=sha256:3b0fd5273586b3945f09b1c06d80a4f4019861919dc0f60ceba402c695feda68