Pith. sign in

Paper Citation Record · LEDGER

Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.15818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.15818 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:03:03.827609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T15:07:39.832649Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da46a16a-343c-46e5-934c-826c4a7a0248 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.628714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:5ef8300c944c8b7cca1662e04e2cf2283b1f9c85cdf50831a4bbc6328d2ab8c0

Observation a2eefc06-abbe-4521-bd74-678f58431d37 · inbound

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation cites this paper.

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:34:53.242156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:34:53.141898Z digest=sha256:fcac887028c241c0328019e17a654a17337163c45a99194faa24adf45c899a11

Observation c8a26450-2e36-4c1b-a2a1-dc8ee2314689 · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:26:22.362066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:39e12cd4a2981921e78a5ff2832e6d684f82efdc693b4f631356ce04bfcd5797

Observation 23c178f8-e3c1-4e2d-a33d-713b0800ab21 · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.503525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:25e7ea19b981c72a1a8f15cb979f1106900ca5476346fc5001f0eb214da2e579

Observation 31424b6a-0733-4bc8-b814-63af7358e723 · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.836921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:f6c0fb974ec5046a50d22c2c1bdfb8e06d929510ead489e71cf7f19d011babe5

Observation 79daf691-163a-4b8d-a35a-e57b70bf965e · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:42:03.022934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:b4aaf08bc9ca882dc07804a3a5afb8ac853882beb19101168d981a285e369b8e

Observation 10f83b87-7cdb-434c-b36f-e1beda082db1 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.937297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:f68b1b419deef273767cb4c1795cb04d734de6059fb1f08f00847a990186af31

Observation 11b1ca19-bb6c-4ff8-b3bf-9e08a21dc877 · inbound

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation cites this paper.

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:03:03.827609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:03:03.827609Z digest=sha256:02d2df566de5dc6b789fea628f8d818084195d8a60280b0a885f577df966a5fa

Observation edc4c914-eb9e-4cdc-a452-ca4cdb0daf9f · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:30.153376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:30.153376Z digest=sha256:fa34b5b5a13ba850b903161e9aa3f45774b7b08cadc3f0feec6a15cc0c35bbc1

Observation 8307ad76-5cfa-48f0-92c3-b7aed07c9196 · inbound

Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model cites this paper.

Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:14:41.760354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:14:41.760354Z digest=sha256:7135b3b6477a0745862185d4db2eed8990c4988ac49de1e48042a91762d404b4

Observation 93bbbced-40a3-4aa5-852d-047e6952a905 · inbound

Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity cites this paper.

Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:52.252058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:31:24.784166Z digest=sha256:a646b1826557da7f4858503d862e5f758fce9cdc2e70d2b62cb31992b1ea0603