Pith. sign in

Paper Citation Record · LEDGER

VideoPoet: A Large Language Model for Zero-Shot Video Generation

As of 5 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 83 inbound Pith citation observations for arXiv:2312.14125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.14125 v4

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T17:51:05.465548Z

measured 138 of 138 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 83 of 83 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:43:40.189837Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch14

External citation measurements

19
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8ab6cf90-919f-4578-9576-703355b0a8e0 · outbound

This paper cites MusicLM: Generating Music From Text.

VideoPoet: A Large Language Model for Zero-Shot Video Generation MusicLM: Generating Music From Text

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.678105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:e3cadd54b3ef8150d3b06f0854927603efcd75ce46efc6a961d49b82f2638c52

Observation d72324c4-0441-4096-931b-d4a3a2ea8736 · outbound

This paper cites Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.522392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:96dd2a2955ee9a7c78cf99294fbbfb6d811e2ff6d837bcffde99266ef72bbc3d

Observation f8532c7b-a952-495a-959b-ec029b5aa662 · outbound

This paper cites PaLM 2 Technical Report.

VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM 2 Technical Report

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.529112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:654b4483216023e5f007c69244b1ffbdfa366dad0359c98e2cd7370b19a51127

Observation e041e5bf-2d20-4c9d-867d-66925fee679e · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.536609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:3b95fc29c18fac414e0cd6ae4165fd8612e75cd4891e0f76fd8e0e18b6d96c8e

Observation 677f517e-ed14-4944-b95a-9fa4b713fb03 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.543020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:346b86fbd13e0edb16256bb997f0b99eb26731fbe4c82b1ec67e43ca99d109ff

Observation 38fdb4f7-7d26-427a-9614-d8a7eeb63858 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

VideoPoet: A Large Language Model for Zero-Shot Video Generation D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.758609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:a29da89c38648e862c0d4dbcc505284d3d69ee8da25992754d3db4d86c76ea7b

Observation e801f874-f46e-48ee-9f0c-c769c3e73901 · outbound

This paper cites A Short Note about Kinetics-600.

VideoPoet: A Large Language Model for Zero-Shot Video Generation A Short Note about Kinetics-600

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.641950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:6db70f8d25194c158d170411cbf0dca2db156a6028c771109db57d65c89026a8

Observation c40f8024-858c-4265-a131-657112c79235 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.648073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:91a21d4e3d5d573dacdfced34cd5a766f99ced5c0e08056772a725c954747281

Observation 426170da-27f0-43e9-a152-835b47dfd271 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.654063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:7101a866f7d65b1028a327f8a53bfcd15f8032086dae7f9a6752c9bf51e4a499

Observation 857360ad-dc8d-4c2b-9692-c79f79f51d0c · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM: Scaling Language Modeling with Pathways

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.660350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:09cc5f81f854127c7e2eb166ebbab9b89c2a2d76812b887ee35d9c3e79afce17

Observation d2bb2690-eed2-4093-ae7c-c962639163d2 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.666225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:41efb5850776cdd4a576cd0e673c60d6acaea2adbf26c444b02795767412e78d

Observation 526db017-b20d-44d4-b90d-84f51d2c4f29 · outbound

This paper cites CCEdit: Creative and Controllable Video Editing via Diffusion Models.

VideoPoet: A Large Language Model for Zero-Shot Video Generation CCEdit: Creative and Controllable Video Editing via Diffusion Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.673022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:48f1dbbf49e8d1abcb3a55d344dbd65844f9016834fe90f20f23bfffd223a23e

Observation e0c5f9ec-7039-4b58-85d5-1414656f5da5 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

VideoPoet: A Large Language Model for Zero-Shot Video Generation TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:17:47.115702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:398a8c5013eb5f911fe89ac39dc2f71aaff0c4a4c8c2748b7754c58c0f2a2c95

Observation 49d9dea7-ebb3-44c1-a303-60614fdea031 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

VideoPoet: A Large Language Model for Zero-Shot Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.684061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:e7301178f2f0c4476de20d0b91aef98a71397ed2708e86a41127a13b91faa93a

Observation 58b1e455-212d-4e5b-8210-49090e264496 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

VideoPoet: A Large Language Model for Zero-Shot Video Generation MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.690665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:7c1836251df04515f8b4f8cb36dbf96ddd76d4e02e461048f0d1071df3a34150

Observation f3ac6d4a-7d3f-4712-a79a-fcfab2ff46af · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Photorealistic Video Generation with Diffusion Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.696684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:156daf6d8b5deea08817fe756b054acdbe820621db5edd5f56c353e343c765d3

Observation 261cd0ef-3c62-48da-b345-be20ab3e4c9b · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.701564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:d1f8b0749f408b6ca4eb8cd7a37448bdd27610e986d51737d625684525cd5476

Observation e5162db9-394b-4500-8c3a-f2b9b4f8d1a3 · outbound

This paper cites CNN Architectures for Large-Scale Audio Classification.

VideoPoet: A Large Language Model for Zero-Shot Video Generation CNN Architectures for Large-Scale Audio Classification

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.707028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:b6aa90a612b36a398aa67c3116e2814d2c467a3b0290e54ee986106663e288fd

Observation 52d30ccd-6698-475e-a49c-1ea087c9c81a · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.713137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:2f63e2696dee43e84e7d1766f66c137fc4b2ee3cfb099d50fe5f472e48f40567

Observation 2d3b364a-44c3-4335-b960-52bddf13a906 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VideoPoet: A Large Language Model for Zero-Shot Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.718331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:b0555de2f5e9dfdcceaf96c659e12173f0aa2b090a509f29973542c4a85967bb

Observation f2a79ac9-2499-46af-993f-b37500f8ab94 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

VideoPoet: A Large Language Model for Zero-Shot Video Generation GAIA-1: A Generative World Model for Autonomous Driving

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.724037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:35856339920dac0f40ba8ef2e3a43de5a45feb895f32ff034bb59941a6ff4c75

Observation 8a4cb642-62f5-4bd6-8982-89ef070e3ec1 · outbound

This paper cites StarCoder: may the source be with you!.

VideoPoet: A Large Language Model for Zero-Shot Video Generation StarCoder: may the source be with you!

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.729578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:f778cf729ab2515af4f70f5031228a17e0735fdc18dbc6f6bd7b5c86cf700441

Observation e4b4618f-8d54-44d2-9ed0-59eb43989095 · outbound

This paper cites MagicEdit: High-Fidelity and Temporally Coherent Video Editing.

VideoPoet: A Large Language Model for Zero-Shot Video Generation MagicEdit: High-Fidelity and Temporally Coherent Video Editing

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.736244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:24a2af2a67177f57d772261bace3c2768bea8f5aa25d867d38937d662cfa3084

Observation 1786137a-258d-4e4a-932a-48301ab9a16c · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

VideoPoet: A Large Language Model for Zero-Shot Video Generation SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.742000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:c57b1d8805647ea3b14d184aa4a45b3a7703d27451050e0934e3a1c3584c7c7c

Observation b57423d2-4534-421e-bf27-bd75324c3949 · outbound

This paper cites Transframer: Arbitrary Frame Prediction with Generative Models.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Transframer: Arbitrary Frame Prediction with Generative Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.748188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:bf0439288a2bc331cc23ef053a9fc35b0180a50820049148250ddaa2da6647dc

Observation c19e3ccf-6048-44b2-b9d8-43e9ed9e1eaa · outbound

This paper cites GPT-4 Technical Report.

VideoPoet: A Large Language Model for Zero-Shot Video Generation GPT-4 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.753771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:861c33fe399f828881eb314736411b523a7c4a888a3eda325a6ffc625ac2ef5e

Observation 53d008ba-ffa9-4723-aa49-19573a81d6f0 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

VideoPoet: A Large Language Model for Zero-Shot Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.549868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:cc016e580fe8c5f3cb31740bbbb3c953c79e2f425641e9ffc6acfbec072583f8

Observation 99106562-d136-4f96-a72f-81dee26e93d4 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Zero-Shot Text-to-Image Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.555998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:cd577a119e89db9528da625302a12ad6ef52bdf5b0fd8b7fdea03ad43d461f50

Observation 876925b9-0608-42c9-a67e-f86e9aca2254 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.561487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:80e4815674ffbb26235aca93297edb03d4b16347dc6e7e697c62bfe6e001b378

Observation 9bae3512-608d-42df-8cee-aca6de0ea124 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

VideoPoet: A Large Language Model for Zero-Shot Video Generation AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:07:57.965390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:728d7a72e4c7e683b75476d2be2698a4f6afe07f4235e24ca01c12ce8e3858e2

Observation 81968d48-1fc2-4979-8168-60b154da969e · outbound

This paper cites A step toward more inclusive people annotations for fairness.

VideoPoet: A Large Language Model for Zero-Shot Video Generation A step toward more inclusive people annotations for fairness

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.815683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:290e57686503686bc9a5ea35b3621eaa5e4f122ada92589d02949f5a7a763644

Observation 6c5d05e8-be76-4106-82d8-d3bcd6dc658e · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.572419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:75b59be03ff88f9f59f32915d945b58478ae8ff3114a9b7b5138fcd78aed93a3

Observation b57c2841-be37-484c-adce-62c917fc5b98 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

VideoPoet: A Large Language Model for Zero-Shot Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.577395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:a4eac836abed645d738295a9d51618dfbf703abb1d2fed68bfa1ac9a41493d09

Observation b0f4a301-317f-4477-a003-e8fad0a29748 · outbound

This paper cites Any-to-Any Generation via Composable Diffusion.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Any-to-Any Generation via Composable Diffusion

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.583568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:003d0ef10c090657da2e51cfa20ee890b7f9e73c218857956f24cb9719c87120

Observation 720d0770-2896-4a25-ae46-c13ea65cf235 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.589795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:6989e8e0320e95e10d6c0a6984f160af717e942c5d65761ff42331af6d4909a9

Observation bbc9e50c-d28f-4d90-9c32-e51673828cc5 · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:43:34.268786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:b0de6980155be9bcc68192743f508c98e7c104eb99945508437f62c6c00320eb

Observation 707238eb-cf21-4b6d-843e-becff3507e23 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VideoPoet: A Large Language Model for Zero-Shot Video Generation ModelScope Text-to-Video Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.602447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:58439b5aed53ef917eb14b320d417c20e2642eab66bd2376b8aae24aacb0cc8b

Observation 744824f4-a4c0-4150-87e6-6a53b95b9172 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

VideoPoet: A Large Language Model for Zero-Shot Video Generation VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.608254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:b0231be62d3f41103c1b6077d780764c4ec745caed79d5b7845e3badd637e16a

Observation 817bb176-c1ba-4fce-9048-b41dca849e31 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:51:05.614606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:6a73f7f1f0510fd25d2ae439367fa0b8be5640e07d8c1dd869f6a4b209df7c1f

Observation 76de91a3-5e90-42e5-8f2c-ab66f4329dd0 · outbound

This paper cites Make Pixels Dance: High-Dynamic Video Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Make Pixels Dance: High-Dynamic Video Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.623145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:0ec654ef90032f36e58eb2ec1808817b857babbdae111c204f0d0381f9de8c88

Observation da46a16a-343c-46e5-934c-826c4a7a0248 · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.628714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:a7ac9951e2b3dc423cf0529ca164b4e8409ea3e172bbc6decd5b9beadb1b6cbb

Observation 8777de45-7e06-48ae-aafe-406197f2a587 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

VideoPoet: A Large Language Model for Zero-Shot Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:47:57.649874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:1de3bf2036bc14d17bf6624b44f18a02293e9a58a31ed4c86313cd5845ee4a33

Observation be72d8fc-555c-431c-aa22-57d4eff183e1 · outbound

This paper cites a {profession or people descriptor} looking {adverb} at the camera.

VideoPoet: A Large Language Model for Zero-Shot Video Generation a {profession or people descriptor} looking {adverb} at the camera

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.806670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:771f453478dd4d2bc57e926728629b937e21117ec4743aa9e3264918922d6792

Observation 70197330-f764-4ba8-b493-0ee893d5aad3 · outbound

This paper cites Both FVD and FAD metrics are calculated using a held-out subset of 25 thousand videos.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Both FVD and FAD metrics are calculated using a held-out subset of 25 thousand videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.811615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:689c1caf5a9ad7ec89236f1bf7a38b1045261efc8fa26ea8af344a463ed20458

Observation f84e8beb-ce07-449f-9ab1-6411976d5f6a · outbound

This paper cites one by one.

VideoPoet: A Large Language Model for Zero-Shot Video Generation one by one

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.820478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:60206b96eb5f59bfeb790289a546d40ee0991f4a843e135a616640e36d5a3c8a

Observation 645acb88-5270-4429-a465-1afe1b80b170 · outbound

This paper cites content” or appearance of the output and the optical flow and depth control the “structure.

VideoPoet: A Large Language Model for Zero-Shot Video Generation content” or appearance of the output and the optical flow and depth control the “structure

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.824953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:ddb892a6aeb3c15547dd9e72728d59bd3ed0aa40ec1dcd5c794511748322f397

Observation 167c74e9-2f36-4773-978c-5bda83bda56e · outbound

This paper cites an unresolved cited work.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-15T17:51:05.829140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:0e209dbb893f55e034e5ece7fd419b4fc46670e4cea118eeadef3da879f638f7

Observation 887fda48-f5ca-4927-b848-b537429ebe6c · outbound

This paper cites For more details, please refer to Appendix A.5.7.

VideoPoet: A Large Language Model for Zero-Shot Video Generation For more details, please refer to Appendix A.5.7

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.763387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:35c813ce6d250bc532c9ade7df1874be4fc5f76f8d696bd0dcae883df834bf93

Observation bde42484-c61a-4667-bc1f-7bad23e08674 · outbound

This paper cites an unresolved cited work.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-15T17:51:05.768055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:fdef5f1f5fac6c70ddf96079e58994a0448ce5d32b6983cda5c144c3db90ad6e

Observation ef73be0a-66be-4975-9bee-bae75e00ae6b · outbound

This paper cites scale (Ho & Salimans, 2022; Brooks et al.

VideoPoet: A Large Language Model for Zero-Shot Video Generation scale (Ho & Salimans, 2022; Brooks et al

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.773516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:c378cfe3e9335564a095c41c7ab52a6b58ec4339fee8a9c7533ab85aead29f02

Observation 4068d1c1-d4d4-4ade-a664-c5c2c802f4c3 · outbound

This paper cites (2022), measure FVD (Unterthiner et al.

VideoPoet: A Large Language Model for Zero-Shot Video Generation (2022), measure FVD (Unterthiner et al

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.780341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:d74da9d2b4a461c3fd6c5e5f9d0f080822c0a09167fb8fa6313a56953b45e0da

Observation f1a4c409-3332-4218-a0b6-d09f7bf6c897 · outbound

This paper cites a still shot of an ugly cartoon, slideshow of an empty scene, low resolution, distorted and disfigured.

VideoPoet: A Large Language Model for Zero-Shot Video Generation a still shot of an ugly cartoon, slideshow of an empty scene, low resolution, distorted and disfigured

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.785995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:9e979f3c6c44555889ed3714ef2915f9e52a7de930bee985ebd9946a5b9242c2

Observation b79ce6f3-a329-416b-8d2a-1fbd06c2d2a9 · outbound

This paper cites Zero-shot MSR-VTT.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Zero-shot MSR-VTT

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.790804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:13b13546c3138ebeb32fab2a1fa9d9e19424c38c0abefa974ec9461fa5d26804

Observation a70028ee-56f3-42dd-ba7d-1f0a054023c8 · outbound

This paper cites To compute the FVD real features, we sample 10K videos from the training set, following TGAN2 (Saito et al., 2020).

VideoPoet: A Large Language Model for Zero-Shot Video Generation To compute the FVD real features, we sample 10K videos from the training set, following TGAN2 (Saito et al., 2020)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.795768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:0767dad2cf333036bea876d41a751d33a5348c502f008841a47ce3513a726f23

Observation 11117250-3312-49e9-8ebe-55e4fd72021e · outbound

This paper cites We follow MAGVIT (Yu et al., 2023a) in evaluating these tasks against the respective real distribution, using 50000×4 samples for K600 and 50000 samples for SSv2.

VideoPoet: A Large Language Model for Zero-Shot Video Generation We follow MAGVIT (Yu et al., 2023a) in evaluating these tasks against the respective real distribution, using 50000×4 samples for K600 and 50000 samples for SSv2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:51:05.801548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:6f2cc9f49e5d9b5977d09f2a01df72ca811c8f8e35c9854271b187ce1e509844

Pith citing papers

Observation 8e454b83-b0a6-4acf-97c9-d5fa3c0306e7 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:6e03965d406630e9d8c2f57530b64592c40ae82a3d5432e597ac98709382f6b4

Observation 9e79960b-2b2d-4427-b250-c01b636a250f · inbound

CameraCtrl: Enabling Camera Control for Text-to-Video Generation cites this paper.

CameraCtrl: Enabling Camera Control for Text-to-Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T02:06:23.410241Z digest=sha256:c2f9e1e70ff452ec43f629cdcd7045d6cd1235be8f72cee607d8f82053220c56

Observation 51943ed4-e0a5-41c3-8e45-7bf8104702d3 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:34:37.790501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:d276c407918667e3c718da4ad6e008c41c27e6d62665a29e2f391d7f9836baa8

Observation c2a2363e-e428-4d25-a201-718edf0b38e4 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:088397caaf6096aed9fad34cc3eda1aa6e73d73cf3785e94e2837f7842f43892

Observation 9ff7ada9-c30e-489a-8d14-57165783bfd1 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:c4f0afef403f0fe0fd82af37168d30f58070383f0fdfb0f6902e619238de0d31

Observation 4c9b4171-7598-425f-a432-73f44bacb461 · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:7440c37b3d1ff8f20731412cc25555a42c5b64c7698d19ec946391d48baed83f

Observation 2d71e7d6-c074-4006-ac1a-b7669ce3efa9 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:def4408faaed5873a5069ea9c95b64132346c6623d88bba6b4c08aa02d8115b4

Observation 92e00695-367b-4952-9570-f4dc5192e3c9 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:2b9428c02deb0edcdb3a01ffa0abf5f5a312f255703da4223313001e2efab2a4

Observation 526fe2fe-c0d8-4dc2-a911-acac4423c6c0 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.848021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:c8b21f2ff1428ec996d4a26b0e9d3b6f39d1e1edc822f9117c02abeb87897a68

Observation 94fdc615-fe7a-445e-85ef-4b717f84e870 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 209

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:5696845b0d1006567b131f967409be712d9a6a2430c11c83e60a73a1cda2d333

Observation 9ee0dc3a-b722-4908-92d4-a0ba3f123e9a · inbound

Long-Context Autoregressive Video Modeling with Next-Frame Prediction cites this paper.

Long-Context Autoregressive Video Modeling with Next-Frame Prediction VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:05:17.338204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T23:05:17.201790Z digest=sha256:76d59d71249edce7165fe1b7b6d12ae1ce2c84c66c82475b3f322c7943f959b3

Observation 147c329a-8e88-49c8-80c4-334865117b8e · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:21:45.070218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:879e118c4cec6cda5ee26de042a6234d2e26b9649332345218fc629260479f34

Observation 3fd18643-4b44-48ce-9684-60c483cc1370 · inbound

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment cites this paper.

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:51:54.639429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:50:59.797593Z digest=sha256:3f39ec84088b83dbc1e093d3beb0f0b23843b847d90d74874775915228973881

Observation 9ae9a470-cd0f-4065-836a-6a790a35dde2 · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:3238cbe5df9e0b2bf69b1e8dabcfc0a376f533b3ab17b67a9cc92f8abac6a3da

Observation 6d296a95-8439-42f2-b733-edea7623bfea · inbound

MSDformer: Multi-scale Discrete Transformer For Time Series Generation cites this paper.

MSDformer: Multi-scale Discrete Transformer For Time Series Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:44:52.852564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T13:44:08.891339Z digest=sha256:f3c9a4ae29c59d3a30f304872a09af9009b328741796d29348353d4dd750b183

Observation e9eef0f7-ce2b-4a08-8c80-2884929218b5 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:1bde5e2d0c5f7cf2af97adde1126dbff79069b8d35b2777e53788a8133562396

Observation 1f71f0a6-667b-468d-9809-fad01b8fa81d · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:2e1d1e3c973533ec28f7725c3d288eef5a0a0abbded2d8865fc399be363f7caa

Observation 1682882e-4061-41e8-93b9-731cc7c8b120 · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:42:57.276017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:7194891493f01ccb827b628f37098ea90d4bd9a775a90e8a03a1dc088f707b5f

Observation 6365ec3f-b529-4515-afe4-b52724596acc · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:42:43.884153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:bd62393cafe31d2e6ee6b9e12a6114b82b30373880654317121b18956b4dae9f

Observation da9d7a49-a7d4-4fa3-9d49-742379d01979 · inbound

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution cites this paper.

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:51:20.336839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T11:47:22.725217Z digest=sha256:6d19a6824e41948c33f3f8c76ca6eef47283e7680a20d4ea70619d6250532626

Observation 350695a1-585f-4bf6-879c-02c2fbc262f0 · inbound

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time cites this paper.

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:15:29.140906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T11:15:29.102090Z digest=sha256:90ba0378a5479e312009fe80d2e695026d5916071ae4fa85d39977381055ecfb

Observation 8de9f563-1905-46e7-8306-0e1ababb2f2a · inbound

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation cites this paper.

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:43:40.189837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:43:40.189837Z digest=sha256:83d4fd6f277c651cfedf8073fb928ea926845b3e9c49c820eecc99e998afcba6

Observation 0854a0a2-e6ca-4e51-b458-91b31e781a56 · inbound

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection cites this paper.

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:48:58.444755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T03:48:26.807495Z digest=sha256:22fc3eae273b0a23cda950ca16e0080fa5ea666685075b87dd12a25b65877cd8

Observation 98559ac9-9b94-4336-b852-3d323f2b331b · inbound

VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation cites this paper.

VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:01:21.184319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:01:21.184319Z digest=sha256:3aef365f150afee540127a8be5241edfcba2a58bf33d5cdaae93e980e012f27b

Observation 21ef90e5-088c-44c6-be63-7e55b6c3f12d · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:21.717266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:21.717266Z digest=sha256:dfc5079fb41745dadaafb1786f43548d6d435f970cd3cdf9f3f21e7febec9039

Observation 0b4d6e9a-1ce7-491f-898f-cf766bb21e39 · inbound

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation cites this paper.

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T08:16:22.386457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:16:22.386457Z digest=sha256:1ec7800d1605ea858e6b54f647f453347c374787d2979eda51f427e54df00233

Observation aa323421-0d13-4461-a656-801cae91e0a8 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:32:01.754595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T17:32:01.642256Z digest=sha256:914931d3a4ee40f5bc5b4ae57582c622d2999a970c2ada00b03f9b256304036f

Observation 39bc8bc5-6125-4b19-b68b-fec4391c411c · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:51:29.910992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T11:48:00.633421Z digest=sha256:5514d36d22c07e397aed44002209c732d92be4bf28d78c3865694426bc1852b9

Observation fd2f9a66-309a-424c-9797-aca88258a8b0 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T05:30:26.243823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:30:26.243823Z digest=sha256:0996c511988b01b9093534ad67b5c3366569542f667ffdcd837a1393e6530b27

Observation 1d7d2b76-04d3-42ec-a25f-dea3a65d84b2 · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:07:29.913988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:af47e09e95635b8a92f9d27ed8bd30867e257166a5be16a68d8f08fc7619447a

Observation 7d5bad3e-2114-4a3e-a60b-72ed088366dc · inbound

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild cites this paper.

Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:24:10.623170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T12:23:02.046757Z digest=sha256:4f83aca7f560afa79ba434b4dc7c8446d2abf34bfe82097f5a3e443b7f59b461

Observation 44896242-d21c-4e96-aec6-7940e67e4f78 · inbound

FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation cites this paper.

FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T13:27:23.641294Z digest=sha256:b2298fb1d44140367f3b2753b34575ed8f5fb6b1e6ea87206522f4c5fd5bbdfe

Observation 374813fc-773d-4a47-9646-51a05d427c6e · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:05.484034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:05.484034Z digest=sha256:0e677fa051b0c194508f9423ddb471fc2a3e11e065294f71798f6b643745aeb7

Observation 11dd78fd-c35b-4708-a638-13d4ed6a3e51 · inbound

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? cites this paper.

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:10:54.362107Z digest=sha256:821a660e4b7259a5c44adf49caa28866514af91c248398daa26a4ce7d25dbc8e

Observation 277bc4e9-0e94-4c3f-a0e7-de9a8b6d9d25 · inbound

DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer cites this paper.

DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:19:55.598710Z digest=sha256:f65b18b3b02f698727d197411413c6cb8fb5eb26e974cfbc5df1c7441b6a4db2

Observation 0d47a4ea-c273-4d8b-a059-0e24ba64dab6 · inbound

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production cites this paper.

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T09:10:31.245928Z digest=sha256:f1691801d681f3ded824e08e60a5b9f28b7de8f36b0ed7b0a77ce60f6e4b4bd3

Observation d8789e48-97b0-4e7a-b65d-b281bf12358b · inbound

Latent-Compressed Variational Autoencoder for Video Diffusion Models cites this paper.

Latent-Compressed Variational Autoencoder for Video Diffusion Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:23:19.583713Z digest=sha256:68d13a0c4a4ae8fee13248d497982fc06894263e09ad3a4b69bbb078a41e1e5e

Observation ee434b15-5d32-4f30-b3be-10927fec7e46 · inbound

Animator-Centric Skeleton Generation on Objects with Fine-Grained Details cites this paper.

Animator-Centric Skeleton Generation on Objects with Fine-Grained Details VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T22:50:07.371117Z digest=sha256:4711688e7fc58fc1a93a9f397e7154a8de7908aa14d94b88bfd2501b617422fe

Observation df1188a5-0226-4137-af9c-92abd584f8ca · inbound

Stream-T1: Test-Time Scaling for Streaming Video Generation cites this paper.

Stream-T1: Test-Time Scaling for Streaming Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:16:14.985693Z digest=sha256:26a3e1d05a98a08c67e142400a9f4cefc95ea87014ad379bece571b772e61b77

Observation 09e6acbd-796b-420c-a216-a593fc77d27b · inbound

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models cites this paper.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:37:24.612175Z digest=sha256:8c65cc3219653cc140e3487487367e1ab8de4af314b91b2dc37b5ab58920f254

Observation 103c0351-9261-4bfa-bc3c-2eae9b239183 · inbound

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models cites this paper.

CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.682740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T22:05:05.400221Z digest=sha256:627bc3ec3c826facb59ca51070dbece47e6942b9131cff02cee6026dd81fbecf

Observation b052e83f-3a42-4d39-9b7a-3979560af37a · inbound

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation cites this paper.

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.399144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T20:59:34.847496Z digest=sha256:842c145182522e3f41dacf1b3009ec2ce779f8aacd9750350800ccf914504922

Observation 214fdd48-07c6-4a65-8631-cbb16196c450 · inbound

PanoWorld: Geometry-Consistent Panoramic Video World Modeling cites this paper.

PanoWorld: Geometry-Consistent Panoramic Video World Modeling VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T15:42:38.272571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T15:41:32.971032Z digest=sha256:141d13e05ef4ef0fb52a192240f409c4ed0282c384f7edcc60d8034e9b8c5368

Observation 5fa8a814-d5a4-4ea4-a5e1-7535e5a3c170 · inbound

See Before You Code: Learning Visual Priors for Spatially Aware Educational Animation Generation cites this paper.

See Before You Code: Learning Visual Priors for Spatially Aware Educational Animation Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T19:18:54.433227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-20T19:18:05.091160Z digest=sha256:3ab36cca7181a88702657f6deec170c2f3563e5f0200ce06d74d0867b36b1eab

Observation 2f9404d6-a817-4f95-8655-4807fe54d58c · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.547588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:8c845af16370a2cc604b1e979e1bcde09805ad801c56bff363286c66be97e3b5

Observation cd686623-2c12-458a-8078-4be655357532 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:25:00.528874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:26a2b95150476a9ab40a180564baff9595e63c6aef6a17864286011ae413f1ab

Observation 5eca1c2d-7e7d-47df-b23e-46862c44f4cd · inbound

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation cites this paper.

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:03:43.538608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:02:45.505404Z digest=sha256:eeb6bc108cf8bd1a26eb9b2c944e85eedef7bc7139cdda9ba948b8b700bcf72c

Observation 493fec11-dc07-43ae-81d6-b6ac6cbda965 · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T17:58:49.542361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T17:57:00.909897Z digest=sha256:13d9be963dfb3aa87b61c4d79f82258c3ed9dad64b8495552255bc4bb299cea8

Observation 5c5ccc08-2965-4900-b399-995bd0de47db · inbound

GeoWorld-VLM: Geometry from World Models for Vision-Language Models cites this paper.

GeoWorld-VLM: Geometry from World Models for Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.568353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T19:02:05.125937Z digest=sha256:4850adfb30d95088c99303d86ad049e42b96f6e3022afc7f84a01de90e0642a3

Observation 907d6c54-9b59-4573-9c18-602b2b3396bf · inbound

Efficient 3D Content Reconstruction and Generation cites this paper.

Efficient 3D Content Reconstruction and Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 120

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:43:14.788662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:38:48.194538Z digest=sha256:340601860e1f0665d09835c23e5e046fe32dc45f205fb729fe7fa7ed18356199

Observation fdd36a72-e2ea-4afc-a7b3-1fbd954de047 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:48:15.028123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:71411f7307f1a97cdf38932ef3ebd1f0b1350eb7b6719d8e3d92af65121cf21c

Observation a5986faf-0494-466a-9eb0-a7170ce8fea6 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.713727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:056810161c3853f89d462803ba177e2162b5c6e9db521dd22932871f0cbc92f5

Observation d0af000d-86ca-487e-89e0-67a1c0ed6e89 · inbound

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV cites this paper.

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:00.781386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T22:52:38.330851Z digest=sha256:acab45b1fc8075062a10368411233e9e9364c4c88490085c8cda9f4e78519e50

Observation 49f6aea2-8df3-45e8-ae61-fa2345974d7a · inbound

Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors cites this paper.

Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:23:20.852373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T11:19:42.981402Z digest=sha256:a8111e3364ecef841772f5fe58583db482bf75e4e857f8439f6889b7723f52f1

Observation d043362b-edb4-4f0f-aca9-d47e52c5dfb9 · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:13.807618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:995098dc5c0560d5f42369438db7412569e95fd6b6c6935975e8c6e6ffd9823d

Observation 8ed48a88-b16f-4e57-aa11-e5a8136ba46c · inbound

YoCausal: How Far is Video Generation from World Model? A Causality Perspective cites this paper.

YoCausal: How Far is Video Generation from World Model? A Causality Perspective VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.532361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T08:27:03.674229Z digest=sha256:400bb86c40d1adbc660211c8911540f72cc8ed60f6a6d1826477c8af08ffbda0

Observation 30c66c7a-eddb-4173-8018-536f167db83d · inbound

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory cites this paper.

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:42:46.814700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:38:10.473453Z digest=sha256:544c54b72cbe0bf0cf1687cf75d5c2c714f36055e1f2f2f4a670bce04435eda2

Observation b3920632-604f-4c6d-adf6-b6b8c6ca31e0 · inbound

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models cites this paper.

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.520649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:56:21.783415Z digest=sha256:1cf8c23eb05561a3244ede84fba0d50ba3d7b2f963ee6b122d01f5171e5cacea

Observation 13e1c57a-b89f-40f4-bd48-fb8c092e4078 · inbound

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation cites this paper.

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:16:17.064109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T15:28:41.932618Z digest=sha256:9f0c1e43dbcab5ae74a0266c788aa80a940afdba723099fb29c1fc99a55e3a54

Observation 280afcbd-9cfd-4c24-a13e-52b6731975ec · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:07:12.392743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:902391868908dfec1632d755db2908db6fad2c00e07063d3c31063bbb6e443eb

Observation 6bd3d277-85df-4bf1-84c7-3bf649177d01 · inbound

DisCo: World Models with Discrete Camera Motion Control cites this paper.

DisCo: World Models with Discrete Camera Motion Control VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:37:22.488975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T20:21:20.257532Z digest=sha256:d998dfcf19251431ea6c7e2109b0dd668890232bdca7ab59f199c8aaec16b977

Observation 87b1fd09-ef7f-4892-9577-90b65a0696d6 · inbound

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation cites this paper.

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:57:23.126166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T20:04:10.711238Z digest=sha256:22aceb11bbe626ac437325c7e8a53fce264519f72ca8023bc28df896fd7921b6

Observation 7b5b4f9d-799f-401d-a394-04b3160022c0 · inbound

BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension cites this paper.

BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-27T18:41:08.050272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T18:35:14.331659Z digest=sha256:00978d0e46d1363f2ab4243ee84757ec98918d7ed758d1b293b8bc6c8250d5ef

Observation afdf7d71-0fe2-4ca5-8148-f4f7d7e5ae13 · inbound

FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion cites this paper.

FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:37:37.517195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:44:44.550033Z digest=sha256:7c88032409ad8ee0c371e26bac27c4f8e26a1dee84cd8e22dc5f77ee0f6d375d

Observation 58240efe-f89e-4d1a-a256-81d7542829bf · inbound

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI cites this paper.

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:28:44.918718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T04:02:53.110012Z digest=sha256:15fc2ba282ff9fa0c3bdf6439298f3c7fe760746bc38a1a1cd158cc0c4088615

Observation bd6588b5-2608-4360-b040-605a5ed687ef · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T04:09:35.226682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:4c85ba388dd021a08bfef999e34985177c39803629e8566ad9e1f0b97d1c01ef

Observation 0eeeb9d6-c1d6-4ec4-b33a-006514a0384b · inbound

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing cites this paper.

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:09:50.816535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T05:28:20.527237Z digest=sha256:f6dd66ad5f8e7d067bf42c7702232c6c37aee627cb34143943eaad2484a2efb0

Observation adf36259-efe3-48f3-aae9-0c7d8b2c0269 · inbound

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing cites this paper.

PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:33:45.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T05:16:26.506412Z digest=sha256:4ee300c1c19912b68b08ff8090aa2657ae67a8eac6c88d15595f77fd79a79067

Observation 6b541de9-6699-4ee5-a039-c4e3c4f86411 · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:09:53.841149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:e1ae5433dac011c09b857b4d38c8b34abdd5582271a76cd0d8eb75a82a97f591

Observation 817a37fb-350d-46a6-bd02-822033d8af5a · inbound

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control cites this paper.

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:33:51.327714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T05:03:56.624536Z digest=sha256:809c48e3ec04eff8eb8da2d27bc9f792581dca636ab4655c0f8bc4694b167264

Observation 860ec1ad-416e-49c1-a838-577590f10939 · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T08:14:26.766231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T06:07:44.343539Z digest=sha256:34d04eeb0489b52d9da443e829f2d7c1f705f3d0bc6e7028b49f728d63efa4e9

Observation 56306ded-0c26-4f58-b6af-2abe477d7d65 · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:37:20.285337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:37:20.285337Z digest=sha256:348594fa214d8ab52d2789445f1ca9cebb8e74d091f3c8caceaa64bd23a9e714

Observation 8cdce0f2-e95f-42ef-9a04-bb3aa6fd1a02 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.351137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:d16374af7e01c7e4737a79b247a11d103d04f90ced47af60c249993405225044

Observation d09c0f50-9b73-4199-9535-72598503a29a · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:25:41.910873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:e46b5a0089594f666c008b3d7513c7601ef2369d08bf252b8462e05222d1dbae

Observation 7376c00f-d6db-408c-84a6-e1a648edce93 · inbound

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models cites this paper.

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T10:17:43.230980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:17:43.230980Z digest=sha256:ae07c2675b58bc81e56d92c7d8f7060f7d32fe7e331b17c1adfe5da0f42e58db

Observation 27f04a61-2c8c-4a3b-94df-7c053bfcaee5 · inbound

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models cites this paper.

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-02T07:15:24.131370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:15:24.131370Z digest=sha256:4f30c1247cdbff5502b1b8cd580453592f6a4b7a53cabb90647feccc7608c0c5

Observation 827494e7-2d97-42c6-a2e7-265c38e680a0 · inbound

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination cites this paper.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:53.067819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:53.067819Z digest=sha256:e55895d51a1ffb2da2eb7fe36d25274d134cbc5c5347f0addb063be85676bc07

Observation 251e1a99-dd8e-4408-81e2-3fe5fd703cbd · inbound

GS-Agent: Creating 4D Physical Worlds With Generative Simulation cites this paper.

GS-Agent: Creating 4D Physical Worlds With Generative Simulation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T07:14:42.439165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:14:42.439165Z digest=sha256:7a1a0d707ac23e73c70fc1d19363fdbb1df47057d98044916ea548f949fdfad0

Observation 0b9a291d-83b3-4bcd-ad54-24f687eab5db · inbound

Unified Video Dense Prediction from Disjoint Data cites this paper.

Unified Video Dense Prediction from Disjoint Data VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:03:31.467431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:03:31.467431Z digest=sha256:3ec371f2553c566dd88a763603a5d3c8c6c6582cd78b33703dd6e3318749300b

Observation 2eb4b961-5215-4e54-bab7-079bd8f8616c · inbound

Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression cites this paper.

Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:40:12.451852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:40:12.451852Z digest=sha256:c744169ae73c6839b70e31c5bdeec25d93ff175c75096c8b0d83b46c494797b1

Observation 57efef5a-3e34-4d59-8684-9d2da48cd22b · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:13.802617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:13.802617Z digest=sha256:e48e21faa31729f43ba1ad25fabe7cc47ebcbb1e19befc7da046ac0d918fe3c8

Observation 4e3ddda5-a8e2-47b4-a486-cdf97f064abb · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 253

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:31.951096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:31.951096Z digest=sha256:e7a8fc62c888972b3a719cc1b03e7862a74d47a046e0428cbae7fe20e06eff21

Observation 21f913c4-90c0-4b78-a6e2-e0d922063140 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:12.202084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:12.202084Z digest=sha256:3baad9177085127f29e9812fb091e07265969082972e85b77cf7c0f603e1ce2a