Pith. sign in

Paper Citation Record · LEDGER

World-consistent Video Diffusion with Explicit 3D Modeling

As of 20 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 5 inbound Pith citation observations for arXiv:2412.01821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01821 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:59:22.654750Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:13:40.201882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T14:24:45.166579Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a40f213-86c6-4610-9a16-fcefa4de87c1 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

World-consistent Video Diffusion with Explicit 3D Modeling Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.313777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.313777Z digest=sha256:23e70f28954cf940ec71d27bd5c92fdc043ec2f8e23bc0925b77037215c0316e

Observation ec2399cf-bf68-446c-bea6-32c26bcd16e6 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

World-consistent Video Diffusion with Explicit 3D Modeling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.321343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.321343Z digest=sha256:447cb2690f404a64d0a07647953b52be023b91e97df8c207038e0dc0fed4a383

Observation cbe03498-7435-489d-b3f5-1430a889cc3d · outbound

This paper cites Video generation models as world simulators.

World-consistent Video Diffusion with Explicit 3D Modeling Video generation models as world simulators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.326463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.326463Z digest=sha256:81f847f349cc242a4bdc808c15f4ecff7b8f6ac1775c56135099671fcc061ba2

Observation fab3a95d-971e-4805-a868-738577e5d40e · outbound

This paper cites Video generation models as world simulators.

World-consistent Video Diffusion with Explicit 3D Modeling Video generation models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.716775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.332881Z digest=sha256:72944e7b9437991d9d15976916e54e89b80561add25c96a8523cd394f56eb026

Observation 91753383-d7f3-46ac-863f-be0700137d14 · outbound

This paper cites Efficient Geometry-aware 3D Generative Adversarial Networks.

World-consistent Video Diffusion with Explicit 3D Modeling Efficient Geometry-aware 3D Generative Adversarial Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.338302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.338302Z digest=sha256:be903d051bf24cb1cf8d31ee13646275fb8919d4185aeed7218839e21f64c6ba

Observation 896ce24f-a253-4f15-ad1d-ff0351e98e21 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

World-consistent Video Diffusion with Explicit 3D Modeling PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.343626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.343626Z digest=sha256:9908408b26d12872e996970cefa5ce90ecb10ae581bb71de36061f87e059e8b6

Observation cf8d4779-4dd3-4e05-a466-9b0d20963b5d · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

World-consistent Video Diffusion with Explicit 3D Modeling Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.348841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.348841Z digest=sha256:a57f2865af7308b17be471886acc252e4fc6b8deb2443676a81e4f90fbff5971

Observation 09f1c202-05df-4126-be9f-c8a7f318b5d8 · outbound

This paper cites Depth-supervised NeRF: Fewer views and faster training for free.

World-consistent Video Diffusion with Explicit 3D Modeling Depth-supervised NeRF: Fewer views and faster training for free

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.353768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.353768Z digest=sha256:dc13cf032530bd8f1416579c8359c9ae2fdbef74c755bd6c9c9eb628c77b7df8

Observation 48c0f20a-6148-402b-b669-3c37d268ed8f · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

World-consistent Video Diffusion with Explicit 3D Modeling Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.358745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.358745Z digest=sha256:02f817688f1b59f57debcb4a06f5f7d4a21f6f228f8614eb72c3be52a572ec09

Observation 4ce86b0c-b623-41a4-965a-34d5a9d90557 · outbound

This paper cites Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, 1981.

World-consistent Video Diffusion with Explicit 3D Modeling Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Communications of the ACM, 24(6):381–395, 1981

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.364440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.364440Z digest=sha256:85768dba4b308a214a599d33bbab1fd427748113da40855018f058a7458bd35f

Observation c3a0d97d-87c2-4bf8-9b98-29e89f28876d · outbound

This paper cites CAT3D: Create Anything in 3D with Multi-View Diffusion Models.

World-consistent Video Diffusion with Explicit 3D Modeling CAT3D: Create Anything in 3D with Multi-View Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.368795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.368795Z digest=sha256:a185b7f7a600a1ebdf98417792cf8a415d7df262ecf54052e3c3acd8cbb3a2b0

Observation 9b9ddeb9-6603-4097-b682-afc21f3c449a · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

World-consistent Video Diffusion with Explicit 3D Modeling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.373007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.373007Z digest=sha256:2c5a68add793ef17170186d4c53f67304f2f68e038619b908c2836a45b3d1cff

Observation 33ba5c2d-344d-46af-87e2-3dfb71dac17e · outbound

This paper cites NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware Diffusion.

World-consistent Video Diffusion with Explicit 3D Modeling NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware Diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.377521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.377521Z digest=sha256:caf9105cd51ef6272b917b26cdf7a02c92669dea938384784342794a73054fb9

Observation 127b668d-3215-4bf9-a42f-b31e3b6ba2ce · outbound

This paper cites Control3diff: Learning con- trollable 3d diffusion models from single-view images.

World-consistent Video Diffusion with Explicit 3D Modeling Control3diff: Learning con- trollable 3d diffusion models from single-view images

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.658917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.382215Z digest=sha256:406d4d6572979052b00ab75b47dfca5199a921c668fb7b745808169924715de4

Observation 6129322b-e9b5-4db2-b582-edc1e173180e · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

World-consistent Video Diffusion with Explicit 3D Modeling AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.387324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.387324Z digest=sha256:f312c6a2c14a2548f99a4bd975d61b608a5d0d3a9d7051f7a6d2a31ed19f1e07

Observation 9293ea54-ca52-4457-99c4-1baaa74ff268 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

World-consistent Video Diffusion with Explicit 3D Modeling CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.392978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.392978Z digest=sha256:162b24d912adebaa7e0f2f644adeb0e5f4eb5611edac6dfd4c3516c9b182863d

Observation cdc74aa5-cd08-4d1b-b18c-565979ef928b · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

World-consistent Video Diffusion with Explicit 3D Modeling Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.397706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.397706Z digest=sha256:323fdc489064c26889d71a1559a809000216ccaa646ee7399da47ced239e3225

Observation 213824c2-1b9d-41e4-bf1d-709406f286fe · outbound

This paper cites Denoising diffu- sion probabilistic models.

World-consistent Video Diffusion with Explicit 3D Modeling Denoising diffu- sion probabilistic models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.629525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.402948Z digest=sha256:2beedc2d9d0ff9cd826219d8b5d8b846d868829901e3f330647a8b3dee1d50aa

Observation fc87e58d-6030-4525-8c72-1f10833d7c8f · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

World-consistent Video Diffusion with Explicit 3D Modeling Imagen Video: High Definition Video Generation with Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.407423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.407423Z digest=sha256:bbc376b0c1b445879e6b4168128f0c93def4fcfcf6e932a3fecc21cf8ba53651

Observation 0e20f642-fd3a-413b-9de6-a6e9e6bd37b9 · outbound

This paper cites Video diffu- sion models.

World-consistent Video Diffusion with Explicit 3D Modeling Video diffu- sion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.611028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.412844Z digest=sha256:2d17d419e37ba789c10aa416592c25f60c8b099caab41f2fcfd9a01ba216bcc5

Observation 1ce77cb0-62c5-4408-87cd-869a604a0670 · outbound

This paper cites Viewdiff: 3d-consistent image generation with text-to-image models.

World-consistent Video Diffusion with Explicit 3D Modeling Viewdiff: 3d-consistent image generation with text-to-image models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.592906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.417494Z digest=sha256:319c58364f2b91fd36f648768078554fa89804d9d75cac34497522ae0f0d4c49

Observation 139530ad-7517-4da4-b89f-f6097094dd90 · outbound

This paper cites LRM: Large Reconstruction Model for Single Image to 3D.

World-consistent Video Diffusion with Explicit 3D Modeling LRM: Large Reconstruction Model for Single Image to 3D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.422531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.422531Z digest=sha256:a9ce873b50d0adeffd8d071bb712d4f01e09143d3be3b72851a11241d64f21ba

Observation 941151d8-815f-48da-b669-4bc0a882d662 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

World-consistent Video Diffusion with Explicit 3D Modeling Vbench: Comprehensive bench- mark suite for video generative models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.573363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.427156Z digest=sha256:bd0564821992efaae83809370b3f435433b3dee171a80212823e8fdfc04830a2

Observation 83c05cd4-c6f9-4bdf-b616-f21ac42b4b0c · outbound

This paper cites Auto-Encoding Variational Bayes.

World-consistent Video Diffusion with Explicit 3D Modeling Auto-Encoding Variational Bayes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.432329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.432329Z digest=sha256:1fb7c5b09ed1b60381e042a863820f129c0db5f103ead1c294b4d3056de2f46b

Observation 5d47135f-b931-49f2-a3f3-a368d7ee94aa · outbound

This paper cites Ground- ing image matching in 3d with mast3r, 2024.

World-consistent Video Diffusion with Explicit 3D Modeling Ground- ing image matching in 3d with mast3r, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.436788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.436788Z digest=sha256:f014c45cbe36b02939e9b5db723aef7d9642be97d10aadedb5b6762eb3368495

Observation 81ad3779-aee3-46c9-9837-aaa5a2389d71 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object, 2023.

World-consistent Video Diffusion with Explicit 3D Modeling Zero-1-to-3: Zero-shot one image to 3d object, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.544611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.442173Z digest=sha256:d414465100faa3a1e5528d701fb058eeddd51744da9fb4a1a14e7a651d618eb9

Observation d9e72d13-c0cf-4c82-8869-91ed41236709 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

World-consistent Video Diffusion with Explicit 3D Modeling SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.446660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.446660Z digest=sha256:228ea3d02749b70b8eefa9b2c2a06228f57b629ed284eb80a283d829b6cba117

Observation 56ffc5a0-648a-4fe1-b7e6-1cad262bdabf · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

World-consistent Video Diffusion with Explicit 3D Modeling Repaint: Inpainting using denoising diffusion probabilistic models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.527487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.451606Z digest=sha256:8b1f0435f3fa33465354eda70918fd1be25dcc6cbc4d12f92eab9a3e8b526e3f

Observation b8316ea5-4ac2-4a8a-a23d-fa48951d0e08 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

World-consistent Video Diffusion with Explicit 3D Modeling SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.456118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.456118Z digest=sha256:23e25f513ff969e5695f39880efa1029b0f72ded29d1ce79078224c1c9bfe356

Observation 842ad03b-bcd3-42e1-af5c-759d19c896e2 · outbound

This paper cites Gnerf: Gan-based neural radiance field without posed camera.

World-consistent Video Diffusion with Explicit 3D Modeling Gnerf: Gan-based neural radiance field without posed camera

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.512926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.460554Z digest=sha256:052820c739cca61045f9ed0212a113005d881b31c3a85402cb42c8aa6c2ad097

Observation c34951db-2055-44b7-a6f4-c53e941666e8 · outbound

This paper cites NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis.

World-consistent Video Diffusion with Explicit 3D Modeling NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.464449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.464449Z digest=sha256:28666c9c96034cf8b14ab8c017693da988218f2d755b244461e1c5e995057e7d

Observation 7f2bbfd2-a178-4401-a5ae-c55148462ac9 · outbound

This paper cites GIRAFFE: Repre- senting Scenes as Compositional Generative Neural Feature Fields.

World-consistent Video Diffusion with Explicit 3D Modeling GIRAFFE: Repre- senting Scenes as Compositional Generative Neural Feature Fields

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.497367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.468427Z digest=sha256:8d37738e0e4889f3a1b39bbfb35f21ee1c94f199296b63d9b5b24083a40ba003

Observation b6749ee4-a787-4887-8be1-d39feb144952 · outbound

This paper cites Refusion: 3d reconstruction in dynamic environments for rgb-d cameras exploiting resid- uals.

World-consistent Video Diffusion with Explicit 3D Modeling Refusion: 3d reconstruction in dynamic environments for rgb-d cameras exploiting resid- uals

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.472242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.472242Z digest=sha256:afb2e4f7c8479f46c0cc3903ce722aade8dd0e68052f7dfe3f4c96aa6747c947

Observation f79a3bf5-b65a-48e1-912b-d0c8d69302c4 · outbound

This paper cites Scalable diffusion models with transformers.

World-consistent Video Diffusion with Explicit 3D Modeling Scalable diffusion models with transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.476212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.476212Z digest=sha256:30c00e1995da08a4ea0590381c22c472c8a4d15638c94f7461eb928072f46799

Observation 66ed9a93-edac-4d6d-a062-3f1df3dc3850 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

World-consistent Video Diffusion with Explicit 3D Modeling SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.480255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.480255Z digest=sha256:b76d2abe2fec59212c91b26fc2844722270ac7c4de2b3c16d49b7c25f2759921

Observation 5516aaa0-231c-459f-96b7-488c099b6de8 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

World-consistent Video Diffusion with Explicit 3D Modeling Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.461781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.485361Z digest=sha256:e61fa0df6259ef5945575abb3071dda81d4d5139f6ba72b4383b8e74795692dc

Observation 0cb21094-00d2-4a5f-9dbd-aba91cea4c70 · outbound

This paper cites Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction.

World-consistent Video Diffusion with Explicit 3D Modeling Com- mon objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.445968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.489871Z digest=sha256:e5530101be5c011928d5eaa54d9a80cf646b56c09c94da1c30259ba9ac848f2e

Observation 2324abcb-fa68-412a-8008-f1ff9fb805e7 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

World-consistent Video Diffusion with Explicit 3D Modeling High-resolution image synthesis with latent diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.494813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.494813Z digest=sha256:4c844cfbc23f414d915f836918a7710b0d3a81cac132860ad8ce7c110750a67b

Observation 6a5b5715-0424-4914-aab4-0e61e2eb4f27 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

World-consistent Video Diffusion with Explicit 3D Modeling U- net: Convolutional networks for biomedical image segmen- tation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.419181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.499407Z digest=sha256:54e80866f056e46d9ecc5b0e037b7a880d1d2ad9b4fbaeff315ef3d500c65414

Observation be327aac-0eed-4f3a-a5d1-66ba79f8d094 · outbound

This paper cites Habitat: A Platform for Embodied AI Research.

World-consistent Video Diffusion with Explicit 3D Modeling Habitat: A Platform for Embodied AI Research

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.402012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.505291Z digest=sha256:ce2bb352eae9efa25dc2fd4f2d0325bc6bdcab50580932eb941e4735bcc31079

Observation 81961f21-0ae7-47d3-891b-d5d4930ab393 · outbound

This paper cites Pixelwise view selection for un- structured multi-view stereo.

World-consistent Video Diffusion with Explicit 3D Modeling Pixelwise view selection for un- structured multi-view stereo

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.510467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.510467Z digest=sha256:c501e8a0d25f34ec251fe64e921bce0ab23f6d81246a8329cbfe8f4c44c2ab3e

Observation aecc5aaa-856b-4946-b8a9-2bb092dd7153 · outbound

This paper cites A benchmark and a baseline for robust multi- view depth estimation.

World-consistent Video Diffusion with Explicit 3D Modeling A benchmark and a baseline for robust multi- view depth estimation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.374568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.515619Z digest=sha256:de1ee5babe824fca57575b79c4991ee8ba65eda06a032d74bd5e543fbc6c9d91

Observation 6bb927f5-ece4-4729-ac5a-4e578824560a · outbound

This paper cites MVDream: Multi-view Diffusion for 3D Generation.

World-consistent Video Diffusion with Explicit 3D Modeling MVDream: Multi-view Diffusion for 3D Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.520412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.520412Z digest=sha256:ed0ab76bc05a424ca72fd762d075b4b1151ed66ebaeee939d945d057cde8640b

Observation ef8665a9-f7bd-49ed-a54d-fe81c311333f · outbound

This paper cites 3D Photography using Context-aware Layered Depth Inpainting.

World-consistent Video Diffusion with Explicit 3D Modeling 3D Photography using Context-aware Layered Depth Inpainting

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.357664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.525490Z digest=sha256:a6c2afc58838d2bd2a8fbde2fc85ea21f6fc1c77078fb6e5b1ab3f3d0acd4ea6

Observation 0528201c-c770-4638-9f68-57fac0fd99d8 · outbound

This paper cites Indoor Segmentation and Support Inference from RGBD Images.

World-consistent Video Diffusion with Explicit 3D Modeling Indoor Segmentation and Support Inference from RGBD Images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.340600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.530167Z digest=sha256:f088e396588b092b3bf6d79f9dd8edf371932f7d797be983273238c1d350b51b

Observation bdc6617a-d761-47d0-af65-b83021125512 · outbound

This paper cites Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering.

World-consistent Video Diffusion with Explicit 3D Modeling Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.535320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.535320Z digest=sha256:81a089a1d62abf08103d7bb94021238c35cb05322406fbdd5e1bb3ae9f9d1ab7

Observation 4b5cc56a-e1d2-41d2-8faf-2c1fccb37afa · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

World-consistent Video Diffusion with Explicit 3D Modeling Deep unsupervised learning using nonequilibrium thermodynamics

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.540560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.540560Z digest=sha256:972f18b386b68533d722d911bc0d4e3c6e0cd5050863be127b31ac8b2b6bc1a9

Observation 1f8384e1-d6ba-4514-97d3-32512b8c2c1e · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

World-consistent Video Diffusion with Explicit 3D Modeling Score-Based Generative Modeling through Stochastic Differential Equations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.545495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.545495Z digest=sha256:a3c55d68b96158b7bc020585fadfc98ddffa7e3189f890f13d15dd08b0d3b8c3

Observation 91e4a68c-ba22-4fa2-853a-4fcd5f795813 · outbound

This paper cites Kick back & relax: Learning to reconstruct the 10 world by watching slowtv.

World-consistent Video Diffusion with Explicit 3D Modeling Kick back & relax: Learning to reconstruct the 10 world by watching slowtv

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.313086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.550242Z digest=sha256:2d8a14a8a2964961d9b3961afecca2fcad473bd57cf27ccb479cc8902b1679f7

Observation e16ca5f5-1ff3-41f5-a2eb-6735857dcfb7 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

World-consistent Video Diffusion with Explicit 3D Modeling Roformer: Enhanced transformer with rotary position embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.555714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.555714Z digest=sha256:3d9a20eb27f175795a401c1247984f6221bc147436ce0483216edac3fccd8bee

Observation 3d4ae820-3139-46f9-8e1c-962b73805ebc · outbound

This paper cites Loftr: Detector-free local feature matching with transformers.

World-consistent Video Diffusion with Explicit 3D Modeling Loftr: Detector-free local feature matching with transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.287273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.561084Z digest=sha256:1d4b640998e4c44cc39746c60c72d6f6617e46a34d819473bbf961b15108d197

Observation f84f4506-1b4a-4e58-bec9-4400675d9a8c · outbound

This paper cites DeepV2D: Video to Depth with Differentiable Structure from Motion.

World-consistent Video Diffusion with Explicit 3D Modeling DeepV2D: Video to Depth with Differentiable Structure from Motion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.565704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.565704Z digest=sha256:7c1350c866b045f6cff8867edccef82f3798ba0edc58e108f4c2b904e37a9013

Observation b6d6f419-23ee-47d1-8935-dc60abc0fecc · outbound

This paper cites Demon: Depth and motion network for learning monocular stereo.

World-consistent Video Diffusion with Explicit 3D Modeling Demon: Depth and motion network for learning monocular stereo

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.570401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.570401Z digest=sha256:5c12e9ce51cda2d405b4b93eb244685dc7169d552e97c1ed91d251e0dcd9483e

Observation 0f2f7b85-b8c7-4f33-a37f-8b99dade6290 · outbound

This paper cites Vggsfm: Visual geometry grounded deep structure from motion.

World-consistent Video Diffusion with Explicit 3D Modeling Vggsfm: Visual geometry grounded deep structure from motion

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.574553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.574553Z digest=sha256:052e8b3267f01cbbea949d2f64ae5e25546d12a81dc98733c2395097abbf4734

Observation c4d7b6a2-affa-46d6-a743-d41ea0286182 · outbound

This paper cites ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation.

World-consistent Video Diffusion with Explicit 3D Modeling ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.578661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.578661Z digest=sha256:32748693817f6401dc54507f725590d185e3e5159175952dbc11038292290cfb

Observation 53697f09-d39e-4936-a04c-3a56b2206e0e · outbound

This paper cites Dust3r: Geometric 3d vi- sion made easy.

World-consistent Video Diffusion with Explicit 3D Modeling Dust3r: Geometric 3d vi- sion made easy

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.248635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.582619Z digest=sha256:e203690d322ad4edcae04aed19ef054abeb8312d55325e536bd5b619f9dd4ec3

Observation 95373792-f642-45a2-b23e-34fa89119db4 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

World-consistent Video Diffusion with Explicit 3D Modeling LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.586563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.586563Z digest=sha256:2dcfe57976aaa58b2a1ee8a5aae2bce9ba2e2064f8e1a86a5a74b397ae2eb1b3

Observation 23ddbe7d-ff02-4306-af1f-11bdc1c80743 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving.

World-consistent Video Diffusion with Explicit 3D Modeling Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.591426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.591426Z digest=sha256:9806277b3a377848848b823b1fc1cc0d29deb5ae55eaf94138c44a0af7c3040e

Observation 6ead9686-bdc8-4f0f-a095-51fe27d2ec71 · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

World-consistent Video Diffusion with Explicit 3D Modeling Motionctrl: A unified and flexible motion controller for video generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.221941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.596002Z digest=sha256:6bb96801cf6b5a27a8c9f521f2bf16b3d542ed599228b8f504bca92d2c054cfb

Observation 6b6e4247-3dc0-4900-a3e0-b44f7692a8ee · outbound

This paper cites Controlling Space and Time with Diffusion Models.

World-consistent Video Diffusion with Explicit 3D Modeling Controlling Space and Time with Diffusion Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.600677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.600677Z digest=sha256:237c19d1902572be74a9fcc78bb1ba8360b7931aadc03d4f2a3a72953b0c47a0

Observation 9c908c1a-4947-4128-92f2-bd648979eaf3 · outbound

This paper cites CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation.

World-consistent Video Diffusion with Explicit 3D Modeling CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.606479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.606479Z digest=sha256:2790a9020ee7b583c9109451506f27a9673ba86d0f483bc8233618aea4117f7a

Observation 65d8e9a2-2035-4030-8087-9a677ee8cabb · outbound

This paper cites 3d-aware image synthesis via learning struc- tural and textural representations.

World-consistent Video Diffusion with Explicit 3D Modeling 3d-aware image synthesis via learning struc- tural and textural representations

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.204683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.611275Z digest=sha256:c897a45e725bdb73875740760bfd51a739693a244d9af942b4be5eba9cff5bd7

Observation b2ca9657-cd1f-49fd-98f7-9271d716fd7a · outbound

This paper cites Consistnet: Enforcing 3d consistency for multi- view images diffusion.

World-consistent Video Diffusion with Explicit 3D Modeling Consistnet: Enforcing 3d consistency for multi- view images diffusion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.188927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.616642Z digest=sha256:4696dad63d87a9ce526ca558afcb4b3e64204ae91633cb6c8f3756ed32a3e619

Observation bf1f76fb-be3f-45b2-af73-72c608bb8ff6 · outbound

This paper cites Mvs2d: Efficient multi-view stereo via attention-driven 2d convolutions.

World-consistent Video Diffusion with Explicit 3D Modeling Mvs2d: Efficient multi-view stereo via attention-driven 2d convolutions

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.173612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.621393Z digest=sha256:c96229faf7b64c7161ab262dd355c0e8d09bff5d4dadbb686023dce650b5c066

Observation 7c339a15-5a63-4155-94c9-06ed48a4f344 · outbound

This paper cites Mvsnet: Depth inference for unstructured multi-view stereo.

World-consistent Video Diffusion with Explicit 3D Modeling Mvsnet: Depth inference for unstructured multi-view stereo

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.156877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.626445Z digest=sha256:04173d0efa1c7d22fb3648579f6210704462a8264448a215ba3729545efc90be

Observation 6a942ca6-8fe8-4d4c-a3c9-b796c443d308 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

World-consistent Video Diffusion with Explicit 3D Modeling Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.140069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.631169Z digest=sha256:76d1429ea7d68037514fa569a69ef761e345b0632e35d06554c095aaa0b8a79c

Observation fa8ad4e5-f085-49ee-bf5c-99c6bdd09f9a · outbound

This paper cites Mvimgnet: A large-scale dataset of multi-view images.

World-consistent Video Diffusion with Explicit 3D Modeling Mvimgnet: A large-scale dataset of multi-view images

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.123626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.635840Z digest=sha256:eaff5ce9f290fc685bcb2c4558674286b0bfc55569d362c462b1d1713abeba3a

Observation 9d5b0382-f830-46bb-aa85-ccbd429ea3f8 · outbound

This paper cites Root mean square layer nor- malization.

World-consistent Video Diffusion with Explicit 3D Modeling Root mean square layer nor- malization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.640119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.640119Z digest=sha256:0aa41d69a5d3017d6a0d9c48dda99fe2162a494896bcb2e939354c23da18f77d

Observation fe658124-dec4-432a-a465-0d6148dbd5cc · outbound

This paper cites Vis-mvsnet: Visibility-aware multi-view stereo net- work.

World-consistent Video Diffusion with Explicit 3D Modeling Vis-mvsnet: Visibility-aware multi-view stereo net- work

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.097569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.644658Z digest=sha256:f772e2c115af90ad194aeff7bb7cc1d6fee92df4c8f227546f9558b2254ff044

Observation 696b4bcc-096e-48c7-a4b4-873a5b789aea · outbound

This paper cites Stereo magnification: Learning view synthesis using multiplane images.

World-consistent Video Diffusion with Explicit 3D Modeling Stereo magnification: Learning view synthesis using multiplane images

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.082059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.650084Z digest=sha256:32eefdb9a84ea6632003006e343f51777cc3f2d7ffe6627bde5d855868fef294

Observation 45a56b71-fdc1-4570-b2f9-d437619c2f6b · outbound

This paper cites Stereo magnification: Learning view syn- thesis using multiplane images.

World-consistent Video Diffusion with Explicit 3D Modeling Stereo magnification: Learning view syn- thesis using multiplane images

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:23.067723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T00:59:22.654750Z digest=sha256:4e3f079eeb59c6e36f6c5d98d09d0fed3f13b7c39632d29064e741abb3e66014

Pith citing papers

Observation bab06ec1-95da-4441-aae9-cdc3817c1259 · inbound

Emergent Temporal Correspondences from Video Diffusion Transformers cites this paper.

Emergent Temporal Correspondences from Video Diffusion Transformers World-consistent Video Diffusion with Explicit 3D Modeling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:13:40.201882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:13:40.201882Z digest=sha256:81b67fd39625e7456bd06e8fd526c9b569a86333e02fb5ad619f08758ec7bdd5

Observation 8ff09493-8686-47ca-9a6d-86ba1e844feb · inbound

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations cites this paper.

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations World-consistent Video Diffusion with Explicit 3D Modeling

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:37:07.380617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T06:36:13.144868Z digest=sha256:15e5d89a464f5a76303f9bd03322acfa9505ca9fd5518f1b3a858094f9cc5c6b

Observation b0781bd0-2d4a-41bb-aa5e-314da47346e8 · inbound

SeqTex: Generate Mesh Textures in Video Sequence cites this paper.

SeqTex: Generate Mesh Textures in Video Sequence World-consistent Video Diffusion with Explicit 3D Modeling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:56:37.407284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:56:37.407284Z digest=sha256:5dd6b84ce9532cea4e8a35df051ac5acdbca05c93a611012881c41ae8bf4486f

Observation 4501095a-4b6d-4bae-b511-54d43a1e959f · inbound

Epipolar Geometry Improves Video Generation Models cites this paper.

Epipolar Geometry Improves Video Generation Models World-consistent Video Diffusion with Explicit 3D Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:20:10.193109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:20:10.193109Z digest=sha256:9f3d292936fa7c2ec6ffefdfc71a82103ba84bc6ce1f9f8a3a000d4bc67a8713

Observation 402390c6-7c05-4c0b-a9cf-af707026ac1a · inbound

Unified 3D Scene Understanding Through Physical World Modeling cites this paper.

Unified 3D Scene Understanding Through Physical World Modeling World-consistent Video Diffusion with Explicit 3D Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.168083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T14:19:03.682306Z digest=sha256:9948b2c10c305d8ea5bba9ea68f43a68448732c9b7dc2d1ea13636d5729a9d06