Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T17:51:05.465548Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 83 inbound Pith citation observations for arXiv:2312.14125.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T17:51:05.465548Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:43:40.189837Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
55 of 55 outbound references displayed
External citation measurements
19
pith, observed 2026-08-05T02:28:24.338817Z
Observation 8ab6cf90-919f-4578-9576-703355b0a8e0 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d72324c4-0441-4096-931b-d4a3a2ea8736 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8532c7b-a952-495a-959b-ec029b5aa662 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM 2 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e041e5bf-2d20-4c9d-867d-66925fee679e · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 677f517e-ed14-4944-b95a-9fa4b713fb03 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 38fdb4f7-7d26-427a-9614-d8a7eeb63858 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e801f874-f46e-48ee-9f0c-c769c3e73901 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation A Short Note about Kinetics-600
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c40f8024-858c-4265-a131-657112c79235 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 426170da-27f0-43e9-a152-835b47dfd271 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 857360ad-dc8d-4c2b-9692-c79f79f51d0c · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM: Scaling Language Modeling with Pathways
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d2bb2690-eed2-4093-ae7c-c962639163d2 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation PaLM-E: An Embodied Multimodal Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 526db017-b20d-44d4-b90d-84f51d2c4f29 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation CCEdit: Creative and Controllable Video Editing via Diffusion Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0c5f9ec-7039-4b58-85d5-1414656f5da5 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation TokenFlow: Consistent Diffusion Features for Consistent Video Editing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49d9dea7-ebb3-44c1-a303-60614fdea031 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58b1e455-212d-4e5b-8210-49090e264496 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation MaskViT: Masked Visual Pre-Training for Video Prediction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f3ac6d4a-7d3f-4712-a79a-fcfab2ff46af · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Photorealistic Video Generation with Diffusion Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 261cd0ef-3c62-48da-b345-be20ab3e4c9b · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e5162db9-394b-4500-8c3a-f2b9b4f8d1a3 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation CNN Architectures for Large-Scale Audio Classification
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 52d30ccd-6698-475e-a49c-1ea087c9c81a · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Imagen Video: High Definition Video Generation with Diffusion Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d3b364a-44c3-4335-b960-52bddf13a906 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f2a79ac9-2499-46af-993f-b37500f8ab94 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation GAIA-1: A Generative World Model for Autonomous Driving
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a4cb642-62f5-4bd6-8982-89ef070e3ec1 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation StarCoder: may the source be with you!
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4b4618f-8d54-44d2-9ed0-59eb43989095 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation MagicEdit: High-Fidelity and Temporally Coherent Video Editing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1786137a-258d-4e4a-932a-48301ab9a16c · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b57423d2-4534-421e-bf27-bd75324c3949 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Transframer: Arbitrary Frame Prediction with Generative Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c19e3ccf-6048-44b2-b9d8-43e9ed9e1eaa · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation GPT-4 Technical Report
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53d008ba-ffa9-4723-aa49-19573a81d6f0 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 99106562-d136-4f96-a72f-81dee26e93d4 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Zero-Shot Text-to-Image Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 876925b9-0608-42c9-a67e-f86e9aca2254 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bae3512-608d-42df-8cee-aca6de0ea124 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 81968d48-1fc2-4979-8168-60b154da969e · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation A step toward more inclusive people annotations for fairness
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c5d05e8-be76-4106-82d8-d3bcd6dc658e · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b57c2841-be37-484c-adce-62c917fc5b98 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b0f4a301-317f-4477-a003-e8fad0a29748 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Any-to-Any Generation via Composable Diffusion
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 720d0770-2896-4a25-ae46-c13ea65cf235 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bbc9e50c-d28f-4d90-9c32-e51673828cc5 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Phenaki: Variable Length Video Generation From Open Domain Textual Description
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 707238eb-cf21-4b6d-843e-becff3507e23 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation ModelScope Text-to-Video Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 744824f4-a4c0-4150-87e6-6a53b95b9172 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation VideoGPT: Video Generation using VQ-VAE and Transformers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 817bb176-c1ba-4fce-9048-b41dca849e31 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76de91a3-5e90-42e5-8f2c-ab66f4329dd0 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Make Pixels Dance: High-Dynamic Video Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da46a16a-343c-46e5-934c-826c4a7a0248 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8777de45-7e06-48ae-aafe-406197f2a587 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be72d8fc-555c-431c-aa22-57d4eff183e1 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation a {profession or people descriptor} looking {adverb} at the camera
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 70197330-f764-4ba8-b493-0ee893d5aad3 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Both FVD and FAD metrics are calculated using a held-out subset of 25 thousand videos
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f84e8beb-ce07-449f-9ab1-6411976d5f6a · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation one by one
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 645acb88-5270-4429-a465-1afe1b80b170 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation content” or appearance of the output and the optical flow and depth control the “structure
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 167c74e9-2f36-4773-978c-5bda83bda56e · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 887fda48-f5ca-4927-b848-b537429ebe6c · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation For more details, please refer to Appendix A.5.7
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bde42484-c61a-4667-bc1f-7bad23e08674 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef73be0a-66be-4975-9bee-bae75e00ae6b · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation scale (Ho & Salimans, 2022; Brooks et al
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4068d1c1-d4d4-4ade-a664-c5c2c802f4c3 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation (2022), measure FVD (Unterthiner et al
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1a4c409-3332-4218-a0b6-d09f7bf6c897 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation a still shot of an ugly cartoon, slideshow of an empty scene, low resolution, distorted and disfigured
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b79ce6f3-a329-416b-8d2a-1fbd06c2d2a9 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation Zero-shot MSR-VTT
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a70028ee-56f3-42dd-ba7d-1f0a054023c8 · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation To compute the FVD real features, we sample 10K videos from the training set, following TGAN2 (Saito et al., 2020)
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 11117250-3312-49e9-8ebe-55e4fd72021e · outbound
VideoPoet: A Large Language Model for Zero-Shot Video Generation We follow MAGVIT (Yu et al., 2023a) in evaluating these tasks against the respective real distribution, using 50000×4 samples for K600 and 50000 samples for SSv2
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e454b83-b0a6-4acf-97c9-d5fa3c0306e7 · inbound
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e79960b-2b2d-4427-b250-c01b636a250f · inbound
CameraCtrl: Enabling Camera Control for Text-to-Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51943ed4-e0a5-41c3-8e45-7bf8104702d3 · inbound
VideoPhy: Evaluating Physical Commonsense for Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2a2363e-e428-4d25-a201-718edf0b38e4 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ff7ada9-c30e-489a-8d14-57165783bfd1 · inbound
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c9b4171-7598-425f-a432-73f44bacb461 · inbound
Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d71e7d6-c074-4006-ac1a-b7669ce3efa9 · inbound
Emu3: Next-Token Prediction is All You Need VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92e00695-367b-4952-9570-f4dc5192e3c9 · inbound
Movie Gen: A Cast of Media Foundation Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 526fe2fe-c0d8-4dc2-a911-acac4423c6c0 · inbound
Autoregressive Video Generation without Vector Quantization VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94fdc615-fe7a-445e-85ef-4b717f84e870 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ee0dc3a-b722-4908-92d4-a0ba3f123e9a · inbound
Long-Context Autoregressive Video Modeling with Next-Frame Prediction VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 147c329a-8e88-49c8-80c4-334865117b8e · inbound
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3fd18643-4b44-48ce-9684-60c483cc1370 · inbound
DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ae9a470-cd0f-4065-836a-6a790a35dde2 · inbound
MAGI-1: Autoregressive Video Generation at Scale VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6d296a95-8439-42f2-b733-edea7623bfea · inbound
MSDformer: Multi-scale Discrete Transformer For Time Series Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e9eef0f7-ce2b-4a08-8c80-2884929218b5 · inbound
MMaDA: Multimodal Large Diffusion Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f71f0a6-667b-468d-9809-fad01b8fa81d · inbound
Show-o2: Improved Native Unified Multimodal Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1682882e-4061-41e8-93b9-731cc7c8b120 · inbound
Frozen Forecasting: A Unified Evaluation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6365ec3f-b529-4515-afe4-b52724596acc · inbound
FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da9d7a49-a7d4-4fa3-9d49-742379d01979 · inbound
Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 350695a1-585f-4bf6-879c-02c2fbc262f0 · inbound
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8de9f563-1905-46e7-8306-0e1ababb2f2a · inbound
RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0854a0a2-e6ca-4e51-b458-91b31e781a56 · inbound
MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98559ac9-9b94-4336-b852-3d323f2b331b · inbound
VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ef90e5-088c-44c6-be63-7e55b6c3f12d · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4d6e9a-1ce7-491f-898f-cf766bb21e39 · inbound
The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa323421-0d13-4461-a656-801cae91e0a8 · inbound
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 39bc8bc5-6125-4b19-b68b-fec4391c411c · inbound
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd2f9a66-309a-424c-9797-aca88258a8b0 · inbound
Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7d2b76-04d3-42ec-a25f-dea3a65d84b2 · inbound
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d5bad3e-2114-4a3e-a60b-72ed088366dc · inbound
Are Multimodal LLMs Ready for Surveillance? A Reality Check on Zero-Shot Anomaly Detection in the Wild VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 44896242-d21c-4e96-aec6-7940e67e4f78 · inbound
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 374813fc-773d-4a47-9646-51a05d427c6e · inbound
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11dd78fd-c35b-4708-a638-13d4ed6a3e51 · inbound
Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 277bc4e9-0e94-4c3f-a0e7-de9a8b6d9d25 · inbound
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d47a4ea-c273-4d8b-a059-0e24ba64dab6 · inbound
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8789e48-97b0-4e7a-b65d-b281bf12358b · inbound
Latent-Compressed Variational Autoencoder for Video Diffusion Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee434b15-5d32-4f30-b3be-10927fec7e46 · inbound
Animator-Centric Skeleton Generation on Objects with Fine-Grained Details VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df1188a5-0226-4137-af9c-92abd584f8ca · inbound
Stream-T1: Test-Time Scaling for Streaming Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09e6acbd-796b-420c-a216-a593fc77d27b · inbound
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 103c0351-9261-4bfa-bc3c-2eae9b239183 · inbound
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b052e83f-3a42-4d39-9b7a-3979560af37a · inbound
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 214fdd48-07c6-4a65-8631-cbb16196c450 · inbound
PanoWorld: Geometry-Consistent Panoramic Video World Modeling VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5fa8a814-d5a4-4ea4-a5e1-7535e5a3c170 · inbound
See Before You Code: Learning Visual Priors for Spatially Aware Educational Animation Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f9404d6-a817-4f95-8655-4807fe54d58c · inbound
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd686623-2c12-458a-8078-4be655357532 · inbound
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5eca1c2d-7e7d-47df-b23e-46862c44f4cd · inbound
Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 493fec11-dc07-43ae-81d6-b6ac6cbda965 · inbound
GeoWorld-VLM: Geometry from World Models for Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5c5ccc08-2965-4900-b399-995bd0de47db · inbound
GeoWorld-VLM: Geometry from World Models for Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 907d6c54-9b59-4573-9c18-602b2b3396bf · inbound
Efficient 3D Content Reconstruction and Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fdd36a72-e2ea-4afc-a7b3-1fbd954de047 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5986faf-0494-466a-9eb0-a7170ce8fea6 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0af000d-86ca-487e-89e0-67a1c0ed6e89 · inbound
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49f6aea2-8df3-45e8-ae61-fa2345974d7a · inbound
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d043362b-edb4-4f0f-aca9-d47e52c5dfb9 · inbound
Archon: A Unified Multimodal Model for Holistic Digital Human Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ed48a88-b16f-4e57-aa11-e5a8136ba46c · inbound
YoCausal: How Far is Video Generation from World Model? A Causality Perspective VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30c66c7a-eddb-4173-8018-536f167db83d · inbound
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3920632-604f-4c6d-adf6-b6b8c6ca31e0 · inbound
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13e1c57a-b89f-40f4-bd48-fb8c092e4078 · inbound
LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 280afcbd-9cfd-4c24-a13e-52b6731975ec · inbound
Streaming Video Generation with Streaming Force Control VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6bd3d277-85df-4bf1-84c7-3bf649177d01 · inbound
DisCo: World Models with Discrete Camera Motion Control VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 87b1fd09-ef7f-4892-9577-90b65a0696d6 · inbound
VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b5b4f9d-799f-401d-a394-04b3160022c0 · inbound
BioVid: Autoregressive Video Generation with Biological Behavior Semantic Comprehension VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation afdf7d71-0fe2-4ca5-8148-f4f7d7e5ae13 · inbound
FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58240efe-f89e-4d1a-a256-81d7542829bf · inbound
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd6588b5-2608-4360-b040-605a5ed687ef · inbound
World Action Models: A Survey VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0eeeb9d6-c1d6-4ec4-b33a-006514a0384b · inbound
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adf36259-efe3-48f3-aae9-0c7d8b2c0269 · inbound
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6b541de9-6699-4ee5-a039-c4e3c4f86411 · inbound
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 817a37fb-350d-46a6-bd02-822033d8af5a · inbound
Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 860ec1ad-416e-49c1-a838-577590f10939 · inbound
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 56306ded-0c26-4f58-b6af-2abe477d7d65 · inbound
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cdce0f2-e95f-42ef-9a04-bb3aa6fd1a02 · inbound
Bridging Video Understanding and Generation in a Unified Framework VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d09c0f50-9b73-4199-9535-72598503a29a · inbound
MemLearner: Learning to Query Context memory for Video World Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7376c00f-d6db-408c-84a6-e1a648edce93 · inbound
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f04a61-2c8c-4a3b-94df-7c053bfcaee5 · inbound
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827494e7-2d97-42c6-a2e7-265c38e680a0 · inbound
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251e1a99-dd8e-4408-81e2-3fe5fd703cbd · inbound
GS-Agent: Creating 4D Physical Worlds With Generative Simulation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b9a291d-83b3-4bcd-ad54-24f687eab5db · inbound
Unified Video Dense Prediction from Disjoint Data VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb4b961-5215-4e54-bab7-079bd8f8616c · inbound
Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57efef5a-3e34-4d59-8684-9d2da48cd22b · inbound
TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3ddda5-a8e2-47b4-a486-cdf97f064abb · inbound
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 253
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f913c4-90c0-4b78-a6e2-e0d922063140 · inbound
Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 234
Source-reported events for the cited work
Unavailable: canonical work link unavailable.