Pith. sign in

Paper Citation Record · LEDGER

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

As of 20 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 7 inbound Pith citation observations for arXiv:2502.09268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09268 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:10:13.474149Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:26:40.586348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:16:58.078690Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77857f45-ef83-4b28-9f0c-327c15996f64 · outbound

This paper cites MCIL trains a single goal-conditioned policy by mapping various contexts, such as target images, task IDs, and natural language, into a shared latent goal space.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation MCIL trains a single goal-conditioned policy by mapping various contexts, such as target images, task IDs, and natural language, into a shared latent goal space

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.954329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.439006Z digest=sha256:10f21f01b7337c6d03f6fcd70d76f0656450144851b5da77799659b5565c153c

Observation 0907ae65-e66b-4c30-b0bd-77c67ab2996c · outbound

This paper cites MdetrLC uses text query mod- ulation to detect objects within images and demonstrates strong performance in tasks like visual question answering and phrase localization.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation MdetrLC uses text query mod- ulation to detect objects within images and demonstrates strong performance in tasks like visual question answering and phrase localization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.938766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.442947Z digest=sha256:fdb6414896af6e9eac69af0ffe224d651a4c3ddf357aaa35956a0c4549e5a141

Observation 02ca56fb-f985-4f80-af86-2c27cf2bb354 · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.923892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.446944Z digest=sha256:efa3959258189d72f15b9e7031efd68f17b0c61e4962ee1e78f8ef280d35ac34

Observation 0885ae0c-9d06-4c79-81c9-32b1fec08f36 · outbound

This paper cites Multi-column deep neural networks for image classification.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Multi-column deep neural networks for image classification

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.329891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.329891Z digest=sha256:c937447a8f28d9d60f574d34f99b0f868d30c29096b87ce354681b6166470d59

Observation f1fd9f1e-c76a-437e-8e54-b594bbd20827 · outbound

This paper cites QUAR-VLA: Vision-Language-Action Model for Quadruped Robots.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T22:10:13.737757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.340871Z digest=sha256:d61e88afe4b7e53fd120dcf094743269a57e9922353c7d7ad41eff3790701ad1

Observation 2516f0e0-03cd-4a76-abb8-cbd4490c7864 · outbound

This paper cites Learning universal policies via text-guided video generation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning universal policies via text-guided video generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:14.015370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.346506Z digest=sha256:c13fac8a41a257455a3d1aff4be9b2e012ea7351548b1088cae055e9feb3e12a

Observation 6fc987c4-dec2-4ff9-b155-f9ee0bbd2348 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Denoising Diffusion Probabilistic Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.356571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.356571Z digest=sha256:86b4f8c8938d32d3a8ce6cde43daa596c1b47c8ebafc309fce98307d0b36700e

Observation b85776dd-69b6-4ec7-b660-b452a7b2d027 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.371215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.371215Z digest=sha256:8133a9c185798c9bd6c117a6e46534d1f1c0afbc01f102f578a073464cdb305a

Observation 70c4fcec-1f1e-4baa-af95-4b606682321e · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.376171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.376171Z digest=sha256:731d587cf3c5bf24d1a881977530036280fb2a751f93acbf7388c317bac1378d

Observation c690418d-fc6f-48c7-9c61-2b72c2a73ad0 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.381137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.381137Z digest=sha256:e78381340f8aa9b4e35baea91110098918c48de1a81cde832eec64ddc52f6c3e

Observation 4cdcda9e-d353-4bc5-9939-bcd46c632d83 · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Language Conditioned Imitation Learning over Unstructured Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.386053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.386053Z digest=sha256:4fdb377e50c1e9dc1240b5c85495d2a4bd7f89d5e3883aa5f48d472e93febdd2

Observation 7bfa824a-96b0-471d-9955-e41c25ebed75 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation What matters in language conditioned robotic imitation learning over unstructured data

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:14.000161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.391135Z digest=sha256:afb2c039e31c55e549fe89d7b0533862eaac7d3835e33d53a03a1b9e7cfe6b64

Observation 3266e597-4108-4216-911f-0a1331347839 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation UL2: Unifying Language Learning Paradigms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.406375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.406375Z digest=sha256:7a35e569d33015ec09e313fd6607cddeb3a4e411e4368b86627bb0ac318120fe

Observation 23772e90-b9b7-47b2-a94c-cbe4ca309cae · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.411250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.411250Z digest=sha256:3a90a7ed5a02437665a8e36befdd1a799ed4b39d17d588f4b89e15c05d10f143

Observation 0a303011-a4a4-4ae6-ac89-690b7b35470c · outbound

This paper cites Learning Interactive Real-World Simulators.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning Interactive Real-World Simulators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.416313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.416313Z digest=sha256:2aa50204f2b3634b07c1aaf9719534a8884c8a8f88fa6bbce9b98cc95f8f91db

Observation ae8ed9be-3096-4d47-9f3a-a247369a29f0 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.421325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.421325Z digest=sha256:ce121bd84eca74158ec46a14f7b673d261dc7ad722ec98e2c4e7f37c47a71dc0

Observation 8fd42b4e-9e4f-410f-954d-d9f8695a8d1e · outbound

This paper cites Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline Data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.425586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.425586Z digest=sha256:de67de68792c6a30a17d1cc540cd2d6bb01839507b16f520404cf9e2d0ed20c5

Observation a5b8a444-4ad4-40c5-a1e2-e671fe3d0028 · outbound

This paper cites RoboDreamer: Learning Compositional World Models for Robot Imagination.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation RoboDreamer: Learning Compositional World Models for Robot Imagination

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.429917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.429917Z digest=sha256:d4f531e1fedfe2d0767b967c3fc0af5386bc0c42798bca68f065dbdcc5921467

Observation f38028bd-bebc-4c38-bf91-1ff54971880d · outbound

This paper cites CALVIN Datasets.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation CALVIN Datasets

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.969688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.434172Z digest=sha256:555d1d292099e5874308688054e4f64c2d375cc0e8d2c091f427c6370749b21c

Observation 916af6bd-da65-4567-99fb-987036e90c28 · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.908798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.451086Z digest=sha256:9128f4e4b1da78b10b6907dc6b08f5023ab00d75e152ae1d8fc2c4b3816aedea

Observation 61136e6b-baca-498c-afb1-d98e0e5d757c · outbound

This paper cites This approach decomposes tasks into high-level planning and low-level action generation, improving task execution in com- plex scenarios.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation This approach decomposes tasks into high-level planning and low-level action generation, improving task execution in com- plex scenarios

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.892714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.455199Z digest=sha256:bf3dde6436029364fa9813bc013563d94ed5c23c0348e829256ccd1fba59ed07

Observation 7d8850be-c30c-42d4-838c-19eb376365b5 · outbound

This paper cites SuSIE integrates a large-scale internet visual corpus during sub-goal generation and achieves these generated sub-goals through a low-level goal-oriented policy.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation SuSIE integrates a large-scale internet visual corpus during sub-goal generation and achieves these generated sub-goals through a low-level goal-oriented policy

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.878128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.459443Z digest=sha256:86f28adf6017f8f33d15feca3deff7316d81da0530fc4d33b077cf8a75b4b9f0

Observation b1f26547-a910-458a-903e-73f8612dae7b · outbound

This paper cites time [s] 1 2 3 4 5 A vg.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation time [s] 1 2 3 4 5 A vg

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.862349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.464439Z digest=sha256:7bb05d4ddd827bbc5a0499f1f454d0934415082117592a9502e18deaf31bb3ad

Observation 1447bb0c-990c-49be-a7c6-a4b93879d94f · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.831243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.474149Z digest=sha256:088010120921de103167a51488661044544d753924019a8b3ef250e950c6407b

Observation 79826b32-2b58-4ea7-bbf4-a72f253c24f5 · outbound

This paper cites an unresolved cited work.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Unresolved cited work

Reference 256

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:10:13.846647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.469084Z digest=sha256:68642b1a35c6148b96e8ac80d4ff42a8e9908556c93e6da77ae1d4cff336b86d

Observation f9a55c44-7a6c-4c07-846f-047ab736febd · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.396381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.396381Z digest=sha256:3417dbd1e07affe675501762e44fc2462ec324f095c04d5171b68b7c8ff159ab

Observation 6fe32628-a306-4271-8b7e-b46bcac190a1 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.366340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.366340Z digest=sha256:f97046b16dfb4cef9d6639c369dbdb06985fb80e24481e48cc6c390c9ef1fd2d

Observation cbe8c123-ccae-4043-be10-6e7100ab296b · outbound

This paper cites Learn- ing to augment synthetic images for sim2real policy transfer.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learn- ing to augment synthetic images for sim2real policy transfer

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:10:13.984953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T22:10:13.401709Z digest=sha256:8031b6e9b74beb0b16b039b37de1a3c821a634d57485fe05237be2773ce15320

Observation 1c6bc823-d6bb-4212-a499-aa72abd1804c · outbound

This paper cites High-Performance Neural Networks for Visual Object Classification.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation High-Performance Neural Networks for Visual Object Classification

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.335182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.335182Z digest=sha256:c01bb6ef2a739cf257440f9a08d1cb32b2dba18870bd59d3cfd9a67770e03677

Observation c4e6c8f0-c41b-44da-a0b2-5ef64e776293 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.324245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.324245Z digest=sha256:08a9a9d607a01169882866ce9ae30c7948a64fa4be21c19a1292ae32ecbf070a

Observation 096025b6-c5eb-4783-af4b-cc2115be1cba · outbound

This paper cites How Far is Video Generation from World Model: A Physical Law Perspective.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation How Far is Video Generation from World Model: A Physical Law Perspective

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.361758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.361758Z digest=sha256:284ec4f448cd0c7db606a021e133eb7722902954ca08ab5b81c629c8851b42a5

Observation 34b33d7c-b41a-42bc-8e51-54355e420024 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.306435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.306435Z digest=sha256:2c1c07470434e8ffb145cf28f321f0f924938c175b6dcdff9a120c20c39c57b7

Observation 392c608a-c55d-489c-9f06-83501e476f14 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.313181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.313181Z digest=sha256:f2bd7f289620efec2832b32ef9b9627ab65202f8c8fbab21a9d174c742a34ad9

Observation 1ea1bcbc-4357-45ea-923b-c6f37ff14cc0 · outbound

This paper cites Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.318398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.318398Z digest=sha256:da2744384583d84c82ae2132c9a3b45b4a3b972f6386d9a7a9315811d2a64461

Pith citing papers

Observation 087655ed-d631-4ecc-90a3-cb09436d4293 · inbound

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation cites this paper.

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:40.586348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:26:40.586348Z digest=sha256:4d19a53e43f1cc20efc3f3025998a9ad1713eb71a8e06cebc37b1e461768aefb

Observation a1b77edf-4751-4032-837f-11d91ee29833 · inbound

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning cites this paper.

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:25:59.007338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:12:08.970164Z digest=sha256:04171b59371fcab2ca0ff24cb4cf81f310cdd6c890eb1694ae3b8ccdfe59b10e

Observation dab481a5-713d-4ec9-b353-a4d72be0babf · inbound

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence cites this paper.

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T00:03:53.609175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:03:53.609175Z digest=sha256:5f21f8d29f17d6d5b868856c1c4d5ad97f9d21ed71eb6288ad4d7340852c2afd

Observation e74777fd-9995-46df-84c4-9d70a46718ba · inbound

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations cites this paper.

STRONG-VLA: Decoupled Robustness Learning for Vision-Language-Action Models under Multimodal Perturbations GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:41:00.027103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:33:16.554905Z digest=sha256:028d7ba482a5dce24c6b7a0dce342b5f480f47db268943eedf7644b0dae222bd

Observation 324a9723-1a61-4abb-b640-8aeace9bd246 · inbound

World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems cites this paper.

World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:30:19.071025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:25:57.218073Z digest=sha256:801c27557f9582ab734d87ec856f9b7b88630ef65105f7f6cb0288452725d0d2

Observation 1693a345-199e-4622-a697-bec1e2ae1258 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.080059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:67f71b357cbaf28dc67f2cb9d950f93f42983548d21824b6a55db8996829c906

Observation 8a465ad2-d513-480b-ab07-c2d1cc5d1253 · inbound

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation cites this paper.

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:32:58.543303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:32:58.543303Z digest=sha256:00a325f0dfadec8e9781799b076a0ba07b9e13ed99003365680665e84e61ed66