Pith. sign in

Paper Citation Record · LEDGER

Geometry-aware 4D Video Generation for Robot Manipulation

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 16 inbound Pith citation observations for arXiv:2507.01099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01099 v4

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T00:12:39.787489Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:12:54.494472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:09:14.650014Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact33
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 53cffb0b-c61b-41c4-bf6d-3fea8a147ec5 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Geometry-aware 4D Video Generation for Robot Manipulation VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.919148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:42ecd8a698af13449267318da09e0d52b76b66d6a20d00c59a3ff5b743aca75d

Observation ac4bf494-836a-4ec4-b010-8df19bdb0a11 · outbound

This paper cites Video diffusion models.

Geometry-aware 4D Video Generation for Robot Manipulation Video diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.050856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:a83bb2e80aee0cb1387d3e35e0ab3f1a238c5bd324fe36c9b0da08be8f0d459d

Observation 282ec790-7e75-4ae9-962a-ccb6ed8741b3 · outbound

This paper cites SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency.

Geometry-aware 4D Video Generation for Robot Manipulation SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.784521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:e665c64f0615b8e5eb49bcdf7da499ff05a13201c3697cdc2ec76ae98ad1129d

Observation 776ac925-ee96-4cde-88f6-cb58b823bccb · outbound

This paper cites Vivid-zoo: Multi-view video generation with diffusion model.

Geometry-aware 4D Video Generation for Robot Manipulation Vivid-zoo: Multi-view video generation with diffusion model

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.046757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:298a45b7282e67254866c7f17bea230bc373357adc50e18260970985d5ab8f2c

Observation 7924ca27-2e5d-4eda-a1d0-13c6b686de57 · outbound

This paper cites 4diffusion: Multi-view video diffusion model for 4d generation.Advances in Neural Information Processing Systems, 37:15272–15295.

Geometry-aware 4D Video Generation for Robot Manipulation 4diffusion: Multi-view video diffusion model for 4d generation.Advances in Neural Information Processing Systems, 37:15272–15295

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.043148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:7494dc28a46dfed3bf3dd497706fe8b62d3354a2029c43e82f1aa16af3157bf3

Observation a9aa21b8-949f-4bed-ae0c-1bfb8d9e4dc6 · outbound

This paper cites Dust3r: Geometric 3d vision made easy.

Geometry-aware 4D Video Generation for Robot Manipulation Dust3r: Geometric 3d vision made easy

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.038181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:7346b494046bace32875ed819aa3c25aa2da8ec1f89e6c8c3d16ff50ed262b9e

Observation 42c0b8b2-8c5b-4519-84f9-4f843365a309 · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

Geometry-aware 4D Video Generation for Robot Manipulation Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.034300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:5900ac643972522632ef8b9ad039e84992d85de678494ae17ef1557da0b6dbff

Observation 8b98f810-1122-4878-954f-ce81858b3223 · outbound

This paper cites Unsupervised learning of video representations using lstms.

Geometry-aware 4D Video Generation for Robot Manipulation Unsupervised learning of video representations using lstms

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.030736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:5b3eb405bdcff25c1990a6ea5dc682d08f44b549d6faf7571eccc2dfb4caa75b

Observation 86333821-dce2-480c-a4fa-1dc5b76dead8 · outbound

This paper cites Recurrent Environment Simulators.

Geometry-aware 4D Video Generation for Robot Manipulation Recurrent Environment Simulators

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.871276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:976a6d9035deed4539c484fd985f9a5014030c108c2c82456caa692c6bc17f1c

Observation abcd5f32-e313-4b8f-a048-ef27d102450c · outbound

This paper cites Generating videos with scene dynamics.

Geometry-aware 4D Video Generation for Robot Manipulation Generating videos with scene dynamics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.026536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:0561658b0f2ef95edb3b89c2e5a48ffd67c482a73ee4f9c793347da32280edf6

Observation 35c5d183-c00f-4c58-b0db-5aacd25b020d · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Geometry-aware 4D Video Generation for Robot Manipulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.846421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:4b2eee66a395b5a498ae77efe42b226f5845eb10afb017031d3c7d78cba8c0b0

Observation f92cfa37-e3ce-47c6-8180-4db4227a202c · outbound

This paper cites MarDini: Masked Autoregressive Diffusion for Video Generation at Scale.

Geometry-aware 4D Video Generation for Robot Manipulation MarDini: Masked Autoregressive Diffusion for Video Generation at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.886735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:6417d2163412139aeb51b4b1f716b5df0c7fbf16e534b3e0ee8ccede1e1eb5db

Observation 3ce58965-c534-4215-9a5e-2878ca5fe0b0 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Geometry-aware 4D Video Generation for Robot Manipulation Open-Sora: Democratizing Efficient Video Production for All

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.888618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:91f3527362634b8c1a29601476525d2b9199de1011fc2b5fb83a03cd1ee4f27e

Observation bd79ae07-3fff-4f44-a3eb-f18531803e8d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Geometry-aware 4D Video Generation for Robot Manipulation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.875634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:6a3babd3c9926ab9157e2fe44e196333309bca65d53d58ad430007bf8f1b04ad

Observation 6f16495f-6ce2-47db-9004-608a320039f3 · outbound

This paper cites Monocular Depth Estimation using Diffusion Models.

Geometry-aware 4D Video Generation for Robot Manipulation Monocular Depth Estimation using Diffusion Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.907942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:3efa7f5ef0e4d495780b3575e8b3088f2709e20d6ad38c0283c793f89cb2a5a0

Observation 5ab287ad-9094-4f9e-b2d4-e73163baea35 · outbound

This paper cites DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos.

Geometry-aware 4D Video Generation for Robot Manipulation DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.814877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:d8a1a218afbba77a29a5860eff00342aa8da22bfe53b477706e93e87a77a0159

Observation 98fde7d7-a816-4244-9ba4-206e26a065fd · outbound

This paper cites Learning Temporally Consistent Video Depth from Video Diffusion Priors.

Geometry-aware 4D Video Generation for Robot Manipulation Learning Temporally Consistent Video Depth from Video Diffusion Priors

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.927985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:d20500c56ebea8a16972c6d14a53402c77961cb5629b32ec2fd3e8efe96bce9a

Observation c2a5b748-0b16-4153-994a-21507831e5ac · outbound

This paper cites Pointmap-conditioned diffusion for consistent novel view synthesis.

Geometry-aware 4D Video Generation for Robot Manipulation Pointmap-conditioned diffusion for consistent novel view synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.852285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:6c4c0b5ccf1568d098e64026a781795065e2211ed93e0591789e563fa753c116

Observation 33371518-6719-403e-a43f-bf30293ff11f · outbound

This paper cites Generative camera dolly: Extreme monocular dynamic novel view synthesis.

Geometry-aware 4D Video Generation for Robot Manipulation Generative camera dolly: Extreme monocular dynamic novel view synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.022981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:a4e56bbc0e6fbb7bd98faeb139d75daa2b989a7c2e804b65704f3d403e09557b

Observation 35b11df5-3c89-4303-8cde-943849bb4d36 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Geometry-aware 4D Video Generation for Robot Manipulation CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.947604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:f5186ea867c0a1673670bd043304a9171f338e28b5c3e617a711544d1ec9304e

Observation 20b4a955-a76d-4ecf-b51a-de8ead676bf0 · outbound

This paper cites Collaborative video diffusion: Consistent multi-video generation with camera control.

Geometry-aware 4D Video Generation for Robot Manipulation Collaborative video diffusion: Consistent multi-video generation with camera control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.019764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:33c1f194398ffec34e2bbc8ac52f620c132768b43275b4f6784d2275a751cacc

Observation 58a7ca4c-0757-46c3-baff-bb2e48051243 · outbound

This paper cites Boosting Camera Motion Control for Video Diffusion Transformers.

Geometry-aware 4D Video Generation for Robot Manipulation Boosting Camera Motion Control for Video Diffusion Transformers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.882162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:1e9221c109d9a6c6ee0b97d4647f9984399fc67e7ecfc96ab27fdf810ae410e1

Observation 6c089247-936e-4758-88a3-7d551b34f1ee · outbound

This paper cites EG4D: Explicit Generation of 4D Object without Score Distillation.

Geometry-aware 4D Video Generation for Robot Manipulation EG4D: Explicit Generation of 4D Object without Score Distillation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.923404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:ae52f558951c6c29bc0254d01bc2093c50b2466307095b270ce993f9ff243776

Observation 482f4f86-5beb-4bfd-9536-a1172a48a60d · outbound

This paper cites Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels.

Geometry-aware 4D Video Generation for Robot Manipulation Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.831284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:a8ee9a36f17d0992c7897a056858e8cb0f3b9a5fa101dac1cf86778e8176674f

Observation 485c2479-39b6-452b-85e4-fb7e43431184 · outbound

This paper cites Diffusion$^2$: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion Models.

Geometry-aware 4D Video Generation for Robot Manipulation Diffusion$^2$: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.932836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:f77759a4a1cbc3b59b4583cb1ce172eae14951f2d4c0966ddec91835a9836475

Observation f7188171-1d57-41f8-a4b6-204ea149d6e7 · outbound

This paper cites Learning universal policies via text-guided video generation.

Geometry-aware 4D Video Generation for Robot Manipulation Learning universal policies via text-guided video generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.016551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:b96b6c2992a2b741fbcb6d9419672ce5c75eecf401f1f51288bac74728b7ecbc

Observation 55ecb31a-45c2-4e5d-93c9-27233adf3d82 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

Geometry-aware 4D Video Generation for Robot Manipulation Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:38:25.165602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:0c0aa6fb822aaed62fc144580b8ccb7babf01ca05528b4bf25ad46de0ffd2f84

Observation 8ff1a0d8-8c02-4d00-8320-c8fcdca8e743 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

Geometry-aware 4D Video Generation for Robot Manipulation TesserAct: Learning 4D Embodied World Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T00:14:27.897807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:b226b0b314b9fe9beab62bd0c0c087ee5aac69916742600c83d1ab964f37cf76

Observation 8829f7e2-2a08-4528-89f5-c414f9aa0190 · outbound

This paper cites Flow as the Cross-Domain Manipulation Interface.

Geometry-aware 4D Video Generation for Robot Manipulation Flow as the Cross-Domain Manipulation Interface

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.867242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:8ee1e9a820809bc12eaa4c2ab20eec87c43fb359e7d8db26250de695657e48ed

Observation 68102722-7369-4727-9a9b-da6f7232ef04 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

Geometry-aware 4D Video Generation for Robot Manipulation Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.918399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:9e44af13ae9312a101277a32bd4d92baa75b95dfb8af7dc40e6975e801c05c51

Observation 766cab2a-1d52-4710-b0c0-08c656489fc2 · outbound

This paper cites Enerverse: Envisioning embodied future space for robotics manipulation.

Geometry-aware 4D Video Generation for Robot Manipulation Enerverse: Envisioning embodied future space for robotics manipulation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.913920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:49640a3dcc6b0f642c2092a40fa6ecd6ca8559507994c2bb6c1362041267cf18

Observation f79772b2-0054-4973-82f3-263257369076 · outbound

This paper cites Dreamitate: Real-World Visuomotor Policy Learning via Video Generation.

Geometry-aware 4D Video Generation for Robot Manipulation Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.884033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:c0a10807f4b07f74329ca76b78bfffd0f696ef635e84d2a3296fce5603ec6754

Observation a7b8eae1-d4d4-48ba-91e4-9d491e308384 · outbound

This paper cites Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets.

Geometry-aware 4D Video Generation for Robot Manipulation Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.909928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:2e7d19d9436f0bf6c1a2cff392d1dff000f97499734cacf3e43621a1bc670ac6

Observation db5616ea-89d5-4be6-83c7-dd1edfafe0d5 · outbound

This paper cites Unified Video Action Model.

Geometry-aware 4D Video Generation for Robot Manipulation Unified Video Action Model

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.819690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:bfd870288b0ab5bb87253968ecb87313df3886ecace6ce78fe86c9c9d756ad06

Observation de99005a-2dcc-4422-b314-9753c4d54569 · outbound

This paper cites Auto-encoding variational bayes.

Geometry-aware 4D Video Generation for Robot Manipulation Auto-encoding variational bayes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.013736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:fb58ba9d3a5ac4794db4acf3878d2da8ce640a9e7f4be4b307ad5828eda52d1c

Observation 27866a51-beda-455b-80d3-4fa29af03b59 · outbound

This paper cites Denoising diffusion probabilistic models.

Geometry-aware 4D Video Generation for Robot Manipulation Denoising diffusion probabilistic models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.010697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:c72b7d588043620cd43418f5ae25781ca4f008630bbf00ef97f40a82a8bc4410

Observation a98b81f6-f58d-445b-b56e-5e2cff3fc537 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Geometry-aware 4D Video Generation for Robot Manipulation SAM 2: Segment Anything in Images and Videos

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.850803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:d77560b88b4d90d5593097b2c0671b3f1f77eda67aa28307a428e931780df905

Observation a4b54c6b-2f2f-4f91-a7bd-7e9fb69698d5 · outbound

This paper cites Lbm eval: Drake-based lbm simulation evaluation suite.

Geometry-aware 4D Video Generation for Robot Manipulation Lbm eval: Drake-based lbm simulation evaluation suite

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.007134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:cdc876167c34868b627f9f9765b885345aada96a0090bd776913eb20f0414953

Observation 2af2c1d4-520e-4eb6-8271-4d5193ea2ff3 · outbound

This paper cites Drake: Model-based design and verification for robotics.

Geometry-aware 4D Video Generation for Robot Manipulation Drake: Model-based design and verification for robotics

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.003952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:48b5797837ad8b5da19a9da650f7340d1127684d1604da98d5f2fb72dc1b713c

Observation c9db65d3-bd9a-4b2d-b122-1571bd43e7c2 · outbound

This paper cites Shape of motion: 4d reconstruc- tion from a single video.

Geometry-aware 4D Video Generation for Robot Manipulation Shape of motion: 4d reconstruc- tion from a single video

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.836845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:df838263446e9dfe252485fe97a2caef020c0bf97a650d39ac26707e807d13e0

Observation 924c111e-e842-4932-be53-cf421368a57a · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Geometry-aware 4D Video Generation for Robot Manipulation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.923660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:383d8db16247c0040f002a07a5f3201807444687a04c29fe53dc57118d879b42

Observation dda0635f-4b09-4033-ac11-64d8f295dd9f · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Geometry-aware 4D Video Generation for Robot Manipulation Diffusion policy: Visuomotor policy learning via action diffusion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:29.000667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:119a38013aa4bd03245122b02d8618e5ad52b9b30f620ada0a5164421867de47

Observation 7ba82f52-973d-4d1e-b1d3-dec217fc3a23 · outbound

This paper cites MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare.

Geometry-aware 4D Video Generation for Robot Manipulation MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.928249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:37812d70861aefae7e1c4b13cb562260f4930a2b14d0b4196db8b647ed2a4703

Observation d0329822-8eff-487b-bb99-4f7e1eab28da · outbound

This paper cites Learning transferable visual models from natural language supervision.

Geometry-aware 4D Video Generation for Robot Manipulation Learning transferable visual models from natural language supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:28.997703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:8f1064278b770131795f161af30268088986ca925157bca8bd168103eef29c01

Observation e6fcda68-b45b-45f6-af27-3344da5c4c62 · outbound

This paper cites FoundationStereo: Zero-Shot Stereo Matching.

Geometry-aware 4D Video Generation for Robot Manipulation FoundationStereo: Zero-Shot Stereo Matching

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.932995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:a6fdcfd67eefb16363b4a9716dc9b80aef8ba97c9801392dee5c922135f845a8

Observation 7e36acaf-7971-4ad2-8fd4-cfaf96e7a220 · outbound

This paper cites Efficient video prediction via sparsely conditioned flow matching.

Geometry-aware 4D Video Generation for Robot Manipulation Efficient video prediction via sparsely conditioned flow matching

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:28.992171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:c1b4718b434da5d42ea5b98726341869208c4c065f546bc2b01485405aa6e396

Observation 36bfb67f-5508-482c-84e7-80ca88898b61 · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954.

Geometry-aware 4D Video Generation for Robot Manipulation Pyramidal flow matching for efficient video generative modeling.arXiv preprint arXiv:2410.05954

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.942858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:a8c0963a5cb10d0398e363a3665ecd74cdcf976cefbaa9cb2bd7c5d469e8ec45

Observation 9a930659-c399-46ff-945b-82b73107f590 · outbound

This paper cites From slow bidirectional to fast causal video generators.

Geometry-aware 4D Video Generation for Robot Manipulation From slow bidirectional to fast causal video generators

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.799024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:ecd2c387e6023202ef9180cee41f9b4e56ec5c064e121a31b2b9617bf684a97c

Observation 2d1a1932-63a9-4f17-afee-c4073be0e256 · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

Geometry-aware 4D Video Generation for Robot Manipulation Autoregressive Video Generation without Vector Quantization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.855764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:787cbf8dcd89d65a0f79028893466f6b9a3e2f1ccf162a6fda40b980cf58a8e3

Observation 71310e1e-4e10-4181-b05a-a37b0ef42456 · outbound

This paper cites ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation.

Geometry-aware 4D Video Generation for Robot Manipulation ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:14:27.809466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:823d499325666e187f9063050fe3491308d67f98feac77e5893ba5264cd3c6e8

Observation e00cf64f-7b1b-4096-bb85-d76c5f62a9e0 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

Geometry-aware 4D Video Generation for Robot Manipulation Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.847601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:93fe5d15bff306a25b4b6064801bb8b0d193fccc6819f86a33a6fc4ce9664f05

Observation 7fc3ebbc-fd3b-4be2-b191-c69e8cbd6ed3 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

Geometry-aware 4D Video Generation for Robot Manipulation Elucidating the design space of diffusion-based generative models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T00:14:28.988800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:a47758695b4b5b9620131ab14f968d6ff8d03c39215ffd337f843da6b0e60a35

Pith citing papers

Observation b1bcbaed-01d8-48b0-a5ca-1eb891ba66f7 · inbound

A Comprehensive Survey on World Models for Embodied AI cites this paper.

A Comprehensive Survey on World Models for Embodied AI Geometry-aware 4D Video Generation for Robot Manipulation

Reference 215

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:54.494472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:54.494472Z digest=sha256:56b43ab67eaa500e2260b9764cc0213f7feedd9fa6fa68b5a5bc52666dbef893

Observation f2f0c198-bbbb-458d-a84f-fd548184518b · inbound

GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation cites this paper.

GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:11.530715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T21:26:59.341715Z digest=sha256:588da0c6899f12d79100ef0e2bf374580f0e1bf9ec1965bbcfb0bb14eaed0ac7

Observation ac8960f6-1069-456d-8fc3-34b85e1fe6e2 · inbound

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation cites this paper.

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:45:07.583622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:45:07.583622Z digest=sha256:432fca0dd4c8d978bba2c6e89a5310ca911e34463d6ff57a6f4e8034b28e3851

Observation 22fae36d-ade7-4181-bdcd-5b18bbc08358 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:04:11.530715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:0033c0dbbacbcb433e8528a61c696c8d7353dc3bd7ad292bcb0f0c58ac52988e

Observation 2866a64b-9007-4a49-84aa-ce01f84f878d · inbound

ShapeGen: Robotic Data Generation for Category-Level Manipulation cites this paper.

ShapeGen: Robotic Data Generation for Category-Level Manipulation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:11.530715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T10:17:05.939312Z digest=sha256:a4b4e900ec85c107f4d6c824e6bfd9fcbd919c3397e5ad5827c627fe97fb5256

Observation b05f868a-cf9b-4dfe-86dc-4b04d68ea36c · inbound

VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis cites this paper.

VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis Geometry-aware 4D Video Generation for Robot Manipulation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:11.530715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T21:18:12.192842Z digest=sha256:b96b364ee25373db39b8862902e209e1a55ad4b64e3e039fc8223e4c1f443ab9

Observation 331752ca-14b8-45dd-8dcc-6ecb1985b2b9 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Geometry-aware 4D Video Generation for Robot Manipulation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:11.530715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T10:40:04.767657Z digest=sha256:f4aeaa4c4d5def67c3017113b4e44978f11b32627473b9bf7235d8b59465d303

Observation b47d4d92-06bb-4240-ba0f-f4d6efb0d4b9 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Geometry-aware 4D Video Generation for Robot Manipulation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:11.530715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T03:20:04.435904Z digest=sha256:396008d51ca98d5446dd0de59a56ee97d1570392c1af1e1a101d668432ee4f55

Observation 6950829d-e897-478c-86c8-51e8e5981ff8 · inbound

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors cites this paper.

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors Geometry-aware 4D Video Generation for Robot Manipulation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:04:39.509788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T06:02:40.043871Z digest=sha256:c07d4c493626f7afe09b02ae0a373ad1dd13b349644acc00a6a61d4a3309a307

Observation c24504df-a69e-40e4-9332-b841e18998bc · inbound

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors cites this paper.

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors Geometry-aware 4D Video Generation for Robot Manipulation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:00:23.174342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:59:52.760486Z digest=sha256:f8eb0d48f4a639703edec39b8bd61117d606db3ba10ae437196d487e25cb1847

Observation 3b3c036d-6d5a-4cf5-b745-3560ba3cbfef · inbound

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation cites this paper.

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:36:39.742035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:34:03.684055Z digest=sha256:ffecae91bc85af29063827965b9e330a4d274c3cfd486371872fe5d83831c552

Observation a4f0a876-4357-4e8f-ab1d-955b9f91d0aa · inbound

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation cites this paper.

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T16:54:59.342170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T16:45:15.955954Z digest=sha256:5358259f3ae8476ad4c54356d2fcf8a3a0c269825a8c30f894f37b3bea7ca524

Observation e8548b04-6d39-42a5-8651-85b6531a5d1c · inbound

Towards Consistent Video Geometry Estimation cites this paper.

Towards Consistent Video Geometry Estimation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:13:16.062309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T08:03:13.579650Z digest=sha256:6080db46d6c5f76a1e3e7a291212e50af4699b82ae519b6227ad9600913897f7

Observation a4405053-c274-4b8d-aba6-38826adb8cf3 · inbound

Towards Consistent Video Geometry Estimation cites this paper.

Towards Consistent Video Geometry Estimation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:16.470127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:16.470127Z digest=sha256:6e5003b48e6ae56db870312def621c7b384c3c2f75132d06571df106e4074c60

Observation f7cb2cd3-0b22-4871-83e8-23bd9c7614f0 · inbound

CP4D: Compositional Physics-aware 4D Scene Generation cites this paper.

CP4D: Compositional Physics-aware 4D Scene Generation Geometry-aware 4D Video Generation for Robot Manipulation

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:57:30.503918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T16:52:36.288955Z digest=sha256:cdee6c3397f27ae8042af69521e16f6426f4e337da4af939433fe63e8bab1918

Observation 3596d35f-8fe9-4e4e-8237-197fb3c10558 · inbound

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction cites this paper.

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction Geometry-aware 4D Video Generation for Robot Manipulation

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:09:14.651825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T21:27:47.702578Z digest=sha256:e3b1b68d6a11d2bfd4f69051465ecb83de6c107a0e3e394ad3b8a7198aa0abeb