Pith. sign in

Paper Citation Record · LEDGER

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation

As of 5 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2605.11832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11832 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:25:18.120832Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact33
  • verified fuzzy57
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f2373c1-810e-4fa0-ac43-615bc624cc48 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.305442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:6662ff057b4b319b8c89d599b0042e2c7e36b7e5399ccbe8f62e55e6463cbf37

Observation b9ae34bd-6f91-44d1-99bc-2197ac885554 · outbound

This paper cites Visual instruction tuning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Visual instruction tuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.288472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:e4bbf44ad8b0ba6ae195ef23fb8f3be758a7f93ef6e3ca249ded1bdbb17259ee

Observation 2d6df557-0000-47fd-ae15-3ba9b5f41561 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.569048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:51e04ece144fa417abf2c94761c875650c69d2015475d637d41594e6564fce6b

Observation 7d7a2cc4-f18c-491c-a1cd-5d77ebe61b48 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.292662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:1d1ea9d8bc35e3ecb388273f15edfab92acec4669df22fe523e919f086f20061

Observation 8f3a12f0-a62a-403a-a83d-25a7ba7e1545 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.566164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:425a5dd0dd1efe8f56324c7ea24d6221e53da1686d59f94bf0bce4003a4e0f83

Observation 33f6e4df-ae6f-4116-865a-762f9b93b9f4 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.571718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:29e28a8d42f9952ee8aeb61ee17f198fdb5bcdd4518c9367d6b577a309a7ca46

Observation 0a8ecc01-d235-4f80-9c16-dbda403f7157 · outbound

This paper cites World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.555877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:c8f5a2b1d039a3a4f28061c3826364c7bb3a03189a94803a8981f37a217ee1f0

Observation db142b39-6285-4a31-b0c2-88094fc651eb · outbound

This paper cites Janusvln: Decoupling semantics and spatiality with dual implicit memory for vision-language navigation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Janusvln: Decoupling semantics and spatiality with dual implicit memory for vision-language navigation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.309592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:fabb0de07c52da63f8bd238bdd50e505e410946d2d7ce914e157b06235e30625

Observation 191a073c-a88f-4214-9dda-cbe1adb50dc3 · outbound

This paper cites StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-10T01:18:45.609520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:a9e76ebbdd777f2206545ab6d5a938746d0d3d87b01d2d22545b9d11051d3a17

Observation 18830f6d-259f-4fec-a569-e42c1b76770f · outbound

This paper cites Octo: An open-source generalist robot policy.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Octo: An open-source generalist robot policy

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.296967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:0a7c9f7c32f605f38310379e1ee6e96a08632b12140a4e101f5f62d2f87a2eb6

Observation 7fdd1dbc-37e2-42eb-8930-fe74a12b833d · outbound

This paper cites Fine-Tuning Vision-Language- Action Models: Optimizing Speed and Success.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Fine-Tuning Vision-Language- Action Models: Optimizing Speed and Success

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.280339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:1b594cbc08c04c273eb981e6b0b7dffe8f6b7ab653ba90863a7e2ebf44b19d1c

Observation 5c4c4424-fbea-4b52-b499-a6ab5cceb5bb · outbound

This paper cites FAST: Efficient Action Tok- enization for Vision-Language-Action Models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation FAST: Efficient Action Tok- enization for Vision-Language-Action Models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.284554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:053452c0791ec7eef88bd1d5339cbdac83277fbdd233052c8fcf79b5ebbfff7a

Observation 87aeed09-82b8-4aec-bf95-a8102cd9285b · outbound

This paper cites Pointvla: Injecting the 3d world into vision-language-action models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Pointvla: Injecting the 3d world into vision-language-action models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.317469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:a1ab781adc22f190326b0f6ab5c77dbc10c5720440ccd02ebac886ce2ffba7e3

Observation b0f6cd15-f67a-4474-8414-1a9447a21244 · outbound

This paper cites RVT- 2: Learning Precise Manipulation from Few Demonstrations.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation RVT- 2: Learning Precise Manipulation from Few Demonstrations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.313550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:83182f6afc56a8469c4e91b959db7333e3c8b942c7d20583d2768f2308b6b293

Observation 9225fb87-3914-4873-91a0-79b34f9a89a0 · outbound

This paper cites Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Evo-0: Vision-language-action model with implicit spatial understanding.arXiv preprint arXiv:2507.00416

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.559411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:57061667280c082b403fae1ad76daccbf08fad2f79af9ced0306c3275b6a6089

Observation 51f755aa-3322-4e6d-9818-ae1c76cce6bc · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision-language-action model.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Spatial forcing: Implicit spatial representation alignment for vision-language-action model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.197392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:b324011814e02d89e130a66c9414f8491e7df1f0d291a1902e6b2c6971ddc859

Observation 52ce7381-91d4-4f68-ac1d-a03b7ee1e083 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Vggt: Visual geometry grounded transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.157535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:61c3932e08ab79019e5c02d92ee2f4967c59cde064bdea72c8bd5a4305f07196

Observation f8284c38-bf39-4a1c-a294-a1821ab599eb · outbound

This paper cites Depth anything 3: Recovering the visual space from any views.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Depth anything 3: Recovering the visual space from any views

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.188676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:da3bdf1d8c3193752ee8eb208033068673133ae0054d2cec75b3773d438b852c

Observation cc119cc7-e19c-48e8-907b-9ddb18c8c265 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Diffusion policy: Visuomotor policy learning via action diffusion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.249341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:a55c3fb10b952e7078d3b908d8fdbddce6e664e0c84ca96ede26cd3e2c433f39

Observation 87f199a3-12e5-420e-8045-9d59d46d5d8b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.532163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:47f7c7bf7acaa5c541935e049f44e1b15e736de82fde1c43c4d2c8d3f22ef934

Observation fef89edf-faa5-4a28-ae26-b5dc647fc156 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.521709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:22dab1f1635a16c5f4fd919374ce13c9cd0183d4ed18e8e72e28d8f3147c9650

Observation a616a8cf-2349-4a45-861f-e567dd4672de · outbound

This paper cites Denoising diffusion probabilistic models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Denoising diffusion probabilistic models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.075769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:f2d04b77a852c3e77e1b3f675656c8854371cd004c6e3eb4e785416f86dd1ba4

Observation d02cea01-1b83-44c9-826f-5249dc0d058a · outbound

This paper cites Denoising diffusion implicit models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Denoising diffusion implicit models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.240886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:33b4e5ec4a591bd0306c360798f6b3b3802dfc5529b73a2ee93bf1a931f439ce

Observation e66d5391-c469-4206-8676-9f7e6bc9a517 · outbound

This paper cites Flow matching for generative modeling.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Flow matching for generative modeling

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.180152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:97103984752890fe09c2c6fd5bbbe3f75d382e4b5d34f422baa89cdab1062315

Observation 573dde78-35b8-40d4-8fa4-4bbe08bb6839 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.205576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:4846b202b26514bdfd2c6862b995b5e75ce3edd90ca4db3b1e90b2718e3854c5

Observation feb54258-3ab5-47fa-bcd4-ad67174d3dd9 · outbound

This paper cites Mean flows for one-step generative modeling.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Mean flows for one-step generative modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.079975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:b5797551be01693f0728edc7e7edb2efeb246d692fa785852956aa2bd3152985

Observation 1c80bdaf-5778-4b72-8259-936850baab92 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.201505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:718a3c8397ab57add2b65a8f1b77d0f495fd8210ae61513cf8e9b6e6abc20dc4

Observation 8803732d-4165-46eb-809c-c420832a64a0 · outbound

This paper cites LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.548705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:0aba8177c66d3062edef2244edc18836a49a57db9adfe388583cafa5c2f3fb8a

Observation 9ebe0adb-2601-41f0-b107-78d498d39560 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.475629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:bd8bfb028658c0df28b8c8d157bf815275e9493f8abb262f3940ecefe5cf7a24

Observation 63cadc00-50d7-4a4d-9cd3-c5420dca5566 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.184282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:3b0270e34f81e4815751efd3a35b7bf70e60d54423016d88e47a93f0215c19c9

Observation 545b4bd8-24ba-4190-b21d-b58ea673866c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.262897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:5c0d6021f8a43f9e10f8f5ff7235191629a093515c1eef81584137845905919e

Observation 7df243d0-936b-4619-a51c-ddedb1a12144 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Instructblip: Towards general-purpose vision- language models with instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.236976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:ac9b9076b891cb4f9b7ed25d4ec44658cc9157ac423689e5f893a11be997f6be

Observation 13a66d3c-996b-4e29-834b-8c30412c7481 · outbound

This paper cites Improved baselines with visual instruction tuning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.166510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:889f9563f993bd38e5307b4ff462758c3f264930f768ece6e1c3e7da0840a302

Observation e63635f2-176a-45c2-922e-a8832ac88554 · outbound

This paper cites Prismatic VLMs: Investigating the design space of visually-conditioned language models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Prismatic VLMs: Investigating the design space of visually-conditioned language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.275543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:8c6dc9eefec25e8388d36e5ed051d084ee33f2c729788c24ce09c88f4922e57a

Observation fa9aa304-2145-49a5-823f-9ec018dbdaa3 · outbound

This paper cites X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.232639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:1dc2352ff78c66b7ac3207df54e73fc92b1025efde8129a35ee6e1d2af734b1b

Observation b9f6804c-bce3-4ec1-8cb4-45ea467ddb3d · outbound

This paper cites St4vla: Spatially guided training for vision-language-action models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation St4vla: Spatially guided training for vision-language-action models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.253451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:1c2e688ec2489f441591a75cf2a1dee1ce2071e0ed17ab825e6bd7161aab98e6

Observation d779b73e-1eb4-43ef-a3ad-5fa9eacf084f · outbound

This paper cites Cot-vla: Visual chain-of-thought reasoning for vision- language-action models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Cot-vla: Visual chain-of-thought reasoning for vision- language-action models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.118512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:5859d71278ab173bfe1eb6c55caa7310e0d4942f892e1efb11cf047234e643a2

Observation 6b421c52-9705-4aff-9064-ee1219c92087 · outbound

This paper cites Univla: Learning to act anywhere with task-centric latent actions.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Univla: Learning to act anywhere with task-centric latent actions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.192799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:87fa6afad20ee5a84b8862395a0335988795b6128c1aff32f797737e55efc31c

Observation f6463a01-5e52-4108-85fc-4f0a404a8948 · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation WorldVLA: Towards Autoregressive Action World Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.509529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:ef7647e7c5d1ee747cdae20d619c7cd237fd9b7ba63b39ceeb79ec706742026f

Observation d3d95fd6-70a1-46aa-8be3-153b1ec476d9 · outbound

This paper cites Reconvla: Reconstructive vision- language-action model as effective robot perceiver.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Reconvla: Reconstructive vision- language-action model as effective robot perceiver

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.084080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:0a0768ca7414e7cce878e4ff22a2220ce0fa121e1e22a06ae7ce9d60d2b65c1c

Observation 8aa8db3e-76f4-4119-b759-b07151e93811 · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Interactive Post-Training for Vision-Language-Action Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.344448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:0fe7da131dc79f77392f45a3ee384ec54f38816815441595a8ea3eae32fde544

Observation 11121fa5-f607-4de9-ae82-b08f1e521eee · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.148678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:55b4ab25ca93addb1735aa366e5fdd08efdde77e82944dd84975c812b86f50d6

Observation f49523b3-8a00-4726-94dd-dc089cd3eeb5 · outbound

This paper cites Pre-Training for Robots: Offline RL Enables Learn- ing New Tasks in a Handful of Trials.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Pre-Training for Robots: Offline RL Enables Learn- ing New Tasks in a Handful of Trials

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.171255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:deb3f5b46a8adcf6dff5be0883006cfeb8c6e9c71fc2f4db0c4651ca5f61e4f9

Observation 6feb5edc-6d85-4b94-ae32-6f8fab2d30dc · outbound

This paper cites Robotic Offline RL from Internet Videos via Value-Function Pre-Training.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Robotic Offline RL from Internet Videos via Value-Function Pre-Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.552494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:baf35864c1a0119ba99890a2ffb3d174e26d1b65627d7a1ad62a0d5ed6c424bf

Observation 731b3f74-9101-4ab5-930e-4eeea282c11f · outbound

This paper cites Steering your generalists: Improving robotic foundation models via value guidance.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Steering your generalists: Improving robotic foundation models via value guidance

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.062630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:74b616f064e33a1a057766d0cea9792be9f6a839cb0f5c60731851efcc5d7a6c

Observation 8e231de0-43fb-4ac1-9d6d-c0af7ae5d9fb · outbound

This paper cites RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.223657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:39028fc6ccf57e1adc9c60697587ee2058de69b59ce1acf4cd7ec8604b40e9a4

Observation ac6df52d-89ff-44d2-bb7f-38590f7e4c60 · outbound

This paper cites Residual reinforcement learning for robot control.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Residual reinforcement learning for robot control

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.114028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:49509dd1803b96c9b70a3afeaf26f62f2ba98c961666e48100126a52d38272b0

Observation a8263df0-2b6d-435c-8b68-aa9a8721d04d · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.469693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:8da90d7feac579cd511bcfae3e97733a41c24667a22657b392458fd492d66a49

Observation 51e9b58d-dd57-4457-a838-e13fe2cadfbe · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:55:40.591564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:13f842f068673cd13c3d4741f26c101ab3f782e11d20ea76cb1ef168aa4e3079

Observation d7d366b8-3c47-4a69-b052-6e54f0bf5bdd · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.092705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:46411494144b90013e8f13d89d4c657cdc73ebfa7f4bc5d9bd9fae6fd4fc962a

Observation 334348df-bd31-4e85-b433-1013ac75ebc5 · outbound

This paper cites Simplevla-rl: Scaling vla training via reinforcement learning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Simplevla-rl: Scaling vla training via reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.101256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:e153e69d60bf61fbf76a8b62da4722d3e15c8b5adfe35a8ef24084af299f6281

Observation 73cffc9a-4790-4d53-b27b-42b0b638c888 · outbound

This paper cites Dream to control: Learning behaviors by latent imagination.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Dream to control: Learning behaviors by latent imagination

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.070955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:76203c2805667ed3c8e594bd19ad3e9a5f0273deb6d2e1c64de5b520ba8205ea

Observation a4194e46-fc59-432f-8cd6-8a72a9a94bff · outbound

This paper cites Mastering atari with discrete world models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Mastering atari with discrete world models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.152884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:5e3e6b1e7fdee3248692d5082760eede77b4725c20f2074821f9f2bdfa3b0720

Observation d43d812f-5efc-4de3-ad70-354597c0bf44 · outbound

This paper cites Mastering diverse control tasks through world models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Mastering diverse control tasks through world models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.301410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:fe03d610464cfa78288d48bf1f194f9873b917246e635d67e83f0c603982420b

Observation df70ea23-9bdf-46e3-a625-766c43ba3dab · outbound

This paper cites Td-mpc2: Scalable, robust world models for continuous control.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Td-mpc2: Scalable, robust world models for continuous control

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.105606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:fd5e8115db14189edb114ebdc3a6f90d28f7aa10c912281da118747e395f9466

Observation 58cb3ff7-21f7-4e30-bb98-b3a61cd37dd2 · outbound

This paper cites Wmpo: World model-based policy optimization for vision-language-action models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Wmpo: World model-based policy optimization for vision-language-action models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.144235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:16f7f673a7048a447f218dc190808f5430b17f1ca54c933ca4229fa58d9a8edc

Observation 6d8aa755-1f51-4f83-8dcd-60380d2305ac · outbound

This paper cites Srpo: Self-referential policy optimization for vision-language-action models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Srpo: Self-referential policy optimization for vision-language-action models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.466539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:0be701b4386559a909b46a484523d3d269e19e788a30911a6e9212cdb09ac92c

Observation 39255eaa-50b6-453d-b8cd-65ec9ec483a9 · outbound

This paper cites Vla-rft: Vision- language-action reinforcement fine-tuning with veri- fied rewards in world simulators.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Vla-rft: Vision- language-action reinforcement fine-tuning with veri- fied rewards in world simulators

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.488722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:c29363c328b7584bcb8914e0fa971048b8287d437f071a89db0b9783a0827cb7

Observation c23624bd-d09a-47b8-9311-99cc335c12c2 · outbound

This paper cites WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-30T02:16:14.370216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:d03a12a74da2bf36d0eaaa67c097b2d3a8efcbab58402fab81125702236bd74a

Observation 01dbde3b-dc5f-4777-9397-d2f5b845662b · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.244813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:77d7bfaa28208dbe32fe361b2defe8af35829e7ab86b84fb81066f94979992e6

Observation e7b87707-c817-4a48-812c-ccb782e5e886 · outbound

This paper cites Vihe: Virtual in-hand eye transformer for 3d robotic manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Vihe: Virtual in-hand eye transformer for 3d robotic manipulation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.219346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:21d86ed535edc5fa960ff9d5d27a7368fae7beed82d004e1c621beb8b6ce50df

Observation e2c7ca76-72ab-4945-b105-c9c064f5a39d · outbound

This paper cites 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation 3d cavla: Leveraging depth and 3d context to generalize vision language action models for unseen tasks

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.539303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:e8b83b5d5ff0db88d9f32319c67f1b97fd46f00f6131534bfb74d599c5da9a67

Observation e847dc85-45dd-46ff-b6b9-26c1ffcaa4ff · outbound

This paper cites Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Depthvla: Enhancing vision-language-action models with depth-aware spatial reasoning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.542759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:1edee9239a07e255f9ebb3f0fbf1768ef866d389e4b0a65b1a61cd25e5bdb616

Observation 4fcfcc17-bca0-4983-90e1-8e69f46bb3db · outbound

This paper cites GeoVLA: Empowering 3D Representations in Vision-Language-Action Models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.495377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:a1b50fe6cbcecb510d240b0cc6af5b0abbcca18722cf77853657f7cec2e21380

Observation 0a0f9c88-bad3-44e6-9be1-887eb936000d · outbound

This paper cites Spatialactor: Exploring disentangled spatial represen- tations for robust robotic manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Spatialactor: Exploring disentangled spatial represen- tations for robust robotic manipulation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.058494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:54fb7b810c1dc7e5e7c8a2e38b9713d9c7fc279870aa9d967c967f5d87239f98

Observation fd9730fd-98b3-4705-be38-93a5d8b9ef43 · outbound

This paper cites Voxact-b: Voxel- based acting and stabilizing policy for bimanual manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Voxact-b: Voxel- based acting and stabilizing policy for bimanual manipulation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.088480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:5db62c29c9f468f86913f941078657b28da67bd272862244c03fe2c8e28e8a61

Observation 0339400f-3587-41ec-9840-64a7155e4c18 · outbound

This paper cites Pointnet: Deep learning on point sets for 3d classification and segmentation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Pointnet: Deep learning on point sets for 3d classification and segmentation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.258600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:a0259d20afae7d10397453c89b4a5f441a45fbdcc6b278811627f78205afc2b4

Observation 293b8f55-0744-4bd4-9101-afdef174c933 · outbound

This paper cites Geoaware- vla: Implicit geometry aware vision-language-action model.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Geoaware- vla: Implicit geometry aware vision-language-action model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.515816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:94a9f405e6c0ee2eef15361a468bbc2fd3ed9e7dc368d552e157c0b18ec4649b

Observation c809f220-c69b-4ac3-b321-2e912a99a44a · outbound

This paper cites 3d-mix for vla: A plug-and-play module for integrating vggt-based 3d information into vision-language-action models.arXiv preprint arXiv:2603.24393.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation 3d-mix for vla: A plug-and-play module for integrating vggt-based 3d information into vision-language-action models.arXiv preprint arXiv:2603.24393

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.518797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:f0acf7e537f372784916491daf14f293b9ad7a3970010ad5aa552e1242bdef2a

Observation cd978677-b2dc-4d98-a906-02392c4811d7 · outbound

This paper cites Bridgevla: Input-output alignment for efficient 3d ma- nipulation learning with vision-language models.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Bridgevla: Input-output alignment for efficient 3d ma- nipulation learning with vision-language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.161974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:08ae26e8aa74ff4bab37a959e045d1839cda17f3d88df239e6e98ab62ca184ff

Observation 38cc84cf-425c-4ad6-9538-06ebc126ce24 · outbound

This paper cites Learning to see and act: Task-aware virtual view exploration for robotic manipulation.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Learning to see and act: Task-aware virtual view exploration for robotic manipulation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.228003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:6ac7f13a3922abad3df575e7abfa45e7a851ec9d0c2e43d274ac8d70ed138278

Observation f20e311a-94f7-4b80-9e3e-0882d69c25f7 · outbound

This paper cites Qwen3-VL Technical Report.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Qwen3-VL Technical Report

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.535241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:6196b60e14baf0651daa8334a397aab81e15259f81f94b248c74924ab995e6dc

Observation 838d5c97-3ccb-4a72-aab0-da79130e38bc · outbound

This paper cites LongCat-Image Technical Report.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation LongCat-Image Technical Report

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:04:13.097348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:17e19564000b181717b604fea4c1178112355e53eb70533efb9fbbd7bc86ed0e

Observation 8511fb50-234e-41c7-86d5-0c67e5419ac3 · outbound

This paper cites Attention is all you need.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Attention is all you need

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.066744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:785b1c487c03eed06a779f506c6b1dd5f984d76af00e06596c004a8aa80b1b0e

Observation 7ed0102e-927a-4bb8-ad9a-cb1cb1ce139e · outbound

This paper cites NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:53:29.501619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:f23e5c4bea770816f0b22d9e6b758c0c0824be4cd53d4518e5b6af2e24a439be

Observation 1af9a063-7198-4022-81bb-20dce2602072 · outbound

This paper cites Mergevla: Cross-skill model merging toward a generalist vision- language-action agent.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Mergevla: Cross-skill model merging toward a generalist vision- language-action agent

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.214989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:c55074e528137334875ce95bc2d4eca1a09732fa216e490c82dcbc581db6db87

Observation 326e6fed-5784-42d5-aa94-4e7bb8b46c5c · outbound

This paper cites Unifolm-vla-0: A vision-language-action (vla) frame- work under unifolm family.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Unifolm-vla-0: A vision-language-action (vla) frame- work under unifolm family

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.109920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:f29e1733f7fcb4bed0546b1369f322e583614d09133c3236b66b244f63b61356

Observation 9a4e9853-a15e-460e-9667-7c41370cc1c8 · outbound

This paper cites Spatialvla: Exploring spatial representa- tions for visual-language-action model.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Spatialvla: Exploring spatial representa- tions for visual-language-action model

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.210394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:b21c6ee2d505f5b3878b09d4aae5742b699c5e600192e104d8a5f6a8beeb6c73

Observation 9d64095d-d893-4373-a8d5-e355477963ba · outbound

This paper cites F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:42:47.099823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:a400e15961e9252350c8c158eba4cb2b7b31e597faeafafd762ca7fcc3eeccea

Observation 04e208cc-5ecd-4083-8d89-6b9d5e55d37e · outbound

This paper cites InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:40.076178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:2b8fe0a226d71a44a5fa8521badc83003c74a4ba70f1b5661f3d0edcf153287d

Observation 51c0309e-8e16-4b81-be9f-6f25b96e42c1 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-06-02T03:04:00.649074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:5aa46a737eb178a08bf5b4d42dfe7055734e9a6f09ed690358dd2d68fe02d7f5

Observation 24268b22-8938-4898-839a-a5a2c3fa97f3 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.545536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:3a477e82dd773fff4b545b481b20f62d8c3a0a9c54308a089d50093e0c5a9302

Observation 5cff0f09-a29d-49bf-a7f2-63e8e3f96287 · outbound

This paper cites and Chater, Nick , year =.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation and Chater, Nick , year =

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:18.305742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:8647de50301224ba0bcd27fae6a899dd5cfefc0baba2500b17027303b939ee08

Observation e0fb0e1e-bd71-4b47-93ea-120209225648 · outbound

This paper cites Topology and data.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Topology and data

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.096669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:b18d8784476fc93bfd56769335649053e3d085e1ba872298f04f6bf714e2e129

Observation 2f78275e-6a36-4e55-9add-a5a97c83b205 · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Back to Basics: Let Denoising Generative Models Denoise

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.491910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:c288a796036f3c4c8f19fb66693f85f82502eebe4be226d5b093097df0547552

Observation db854668-6e36-4969-b74c-f80a75aed200 · outbound

This paper cites Stacked denoising autoencoders: Learning useful rep- resentations in a deep network with a local denoising criterion.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Stacked denoising autoencoders: Learning useful rep- resentations in a deep network with a local denoising criterion

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.267346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:775fd84ec8d37210cfd723cf2e9314b86c0f816dcf8d26f18496a921896eff8c

Observation 7ce0e2ca-735d-4389-a930-ec31bff82385 · outbound

This paper cites Scalable diffusion models with transform- ers.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Scalable diffusion models with transform- ers

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.053959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:c36857173ee9849f15fc7355f46c1521747cd7bf0feafe668c571ec1795bf6fd

Observation 1a86b9b1-97f8-4144-a485-80407dab895d · outbound

This paper cites StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.512326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:bebd46a9767da5be6a15b7f3f29f3453c13c075beaad79f298b2e552de11bd5d

Observation f45a6ce5-b7a9-44da-a8e6-009ed8f3f772 · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:27:18.525531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:cead0f6536d83329ca2e7648af623f3a0c85ad0c75dbaae31005e956ec4a8f5e

Observation fa25e1e8-cbd5-4ef9-b9fc-1c8d87df606c · outbound

This paper cites Decoupled weight decay regulariza- tion.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Decoupled weight decay regulariza- tion

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.175533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:55e125cc1738e6bf8e1cb961613dcad172dbc4cedc0060600d2bb59d81f3abb8

Observation 361fb4d1-2fbd-41bc-a621-d77e7d2e671c · outbound

This paper cites Vision transformers for dense prediction.

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation Vision transformers for dense prediction

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:32:38.271379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T05:25:18.120832Z digest=sha256:12a01d3043d46c4dec5ade48a162033424e56d3f8219415d7a81c7c29a04ebec

Pith citing papers

No inbound Pith citation observations are available.