Pith. sign in

Paper Citation Record · LEDGER

Autoregressive Video Generation without Vector Quantization

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 58 inbound Pith citation observations for arXiv:2412.14169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14169 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T15:07:39.718555Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:42.997910Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.772589Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact26
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76ac30db-26ed-403f-9540-774f28894fd7 · outbound

This paper cites PaLM 2 Technical Report.

Autoregressive Video Generation without Vector Quantization PaLM 2 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.750148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:5f6d1202d5ef81503cc442195c5a0fb9a5b4e3deaa8e3d363e5a37cd59605b28

Observation f1a46148-ad26-410a-95d2-9700b4f269f4 · outbound

This paper cites Imagen 3.

Autoregressive Video Generation without Vector Quantization Imagen 3

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.755520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:9d8ada33ddb41ad31d0a49e93f0a5efec32b9731fc5023a835c5633aea04da08

Observation ba50ecb8-fb95-471d-8cc3-a381649a280b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Autoregressive Video Generation without Vector Quantization Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T15:07:39.760358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:f62d0f45522e33eadd7ee6134033b81ed26ec6ff47a9daeea9d3b3e16655b6b5

Observation 0d06b667-717c-42bf-bcb0-afba223e2c30 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Autoregressive Video Generation without Vector Quantization Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.765826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:110f2c4ab34bd0b3199fb023eeb344cd80cfac6b5b3fcdef57efec541683bfcf

Observation 66e0d9ea-dd01-44ae-baca-dd2d23af5fab · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Autoregressive Video Generation without Vector Quantization Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.770391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:38e10a4c89d5312f699c8c85ca9d5f8bcd19280659a7404ce9b7cd24d9a8a905

Observation 7d57e557-7d78-405d-89fe-b6177df03839 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Autoregressive Video Generation without Vector Quantization Unveiling Encoder-Free Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.775119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:ec40b86d8ff4218478c0973743332347151a0be2b7ff9f867e52ac2a1e76fb75

Observation 19a5f75a-1906-4525-9d36-5a66a2499902 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Autoregressive Video Generation without Vector Quantization Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.780066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:c483926d6cc0855fdcfd14cd4eebe7959afb1429126c7c1faec787c4d2d04bd6

Observation 808f4892-2bbc-4654-ae2f-466784efe87a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Autoregressive Video Generation without Vector Quantization AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.785283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:58717409259bd043ba32e900f7632cece89605971cf709f0ba24e0b1fdb3bf32

Observation b441d8ef-4e23-4bf3-bdad-a6e99cb71ba5 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Autoregressive Video Generation without Vector Quantization Classifier-Free Diffusion Guidance

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.790517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:8eb5015f60bcd834f448826c8175c831336a0282a1083fd2bf1ac8347a03412a

Observation 6eb5ac61-ff59-4442-a916-68628b2d8529 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Autoregressive Video Generation without Vector Quantization CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.797214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:b150bafc886772ab0d81d181064941d9eb1130f2c3e66e2ba22124c76a8e0499

Observation d813ed29-913d-45ae-9ba9-1207ee15fd13 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Autoregressive Video Generation without Vector Quantization ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.802397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:619fd36c6bd3805bc4a0b9b561857349bbd7571696ed7a544661c35956e87afd

Observation e05e5b2e-1d99-4845-81d4-a13b7aeb280d · outbound

This paper cites Deep networks with stochastic depth.

Autoregressive Video Generation without Vector Quantization Deep networks with stochastic depth

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:07:39.921188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:da43ea336bb6167cbed4a4edc2f1d90afd16742bb738c5e5d734f143f4080a4a

Observation 526fe2fe-c0d8-4dc2-a911-acac4423c6c0 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Autoregressive Video Generation without Vector Quantization VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.848021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:061d55affef20a6c326d229720b002806162e3f3f41a908fa567034f10850bce

Observation 7abd1a6d-593b-4dc2-bddd-8e0c71fff58f · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Autoregressive Video Generation without Vector Quantization Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.853919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:40ec7e1b7191f41f6127d20ad461af47837ea72345adc0fe46347e65f7e505c8

Observation a64fa0b6-6fba-4423-8591-c0e820380f6f · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

Autoregressive Video Generation without Vector Quantization Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.859781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:afdc2d4af9580d476d3fe2e685c4eacb816e2df91c233761a5c0b568c6532aca

Observation 51b11311-6f92-4453-8a85-380af600e016 · outbound

This paper cites Decoupled Weight Decay Regularization.

Autoregressive Video Generation without Vector Quantization Decoupled Weight Decay Regularization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.864263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:e05067a796fb0406d9f8003e6b301a587b153dc53c8f8396dba85dfd840d256a

Observation 6a6deca0-96af-4a28-8433-c62921ad5f90 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Autoregressive Video Generation without Vector Quantization SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.869713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:949a5364962219f37e6f5541cdbdc0e99b341c3fe1ee45553760ee927cc9658b

Observation 2140ae52-36a3-45e8-9069-f2ca91432107 · outbound

This paper cites Transframer: Arbitrary Frame Prediction with Generative Models.

Autoregressive Video Generation without Vector Quantization Transframer: Arbitrary Frame Prediction with Generative Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.876874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:b907f5ed4ab93f34dd37e2b12568a5cae55d15b0f0430ae3b540a0c331ca133c

Observation 90101cbb-03d1-4d0d-ad9e-3f4600f48ca3 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Autoregressive Video Generation without Vector Quantization GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.882909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:99f48568cd5f473588f6d47fb9e7330d496fd7a2c389359d1d2a94a75637e658

Observation d3e7fe0d-8c7c-4038-a3d1-9c29714185f7 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Autoregressive Video Generation without Vector Quantization SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.888848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:943c1cc997c3ea160086d749d5dbec5f10b19366b8913a237d6700fb91fb5752

Observation 59847b72-a959-4125-ba4e-ae98b9fc612e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Autoregressive Video Generation without Vector Quantization Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.896283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:8da6354448834bf83bcf54c9faf83d56c7735d7eda72af61d464bbfb0ddcee6a

Observation 5d36d21f-3e7a-48e6-8b93-471f14b1a2f1 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Autoregressive Video Generation without Vector Quantization Score-Based Generative Modeling through Stochastic Differential Equations

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.902161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:1d12c413e86594c651fb492592c198b95575a44e1edc61f31e5c4652be43c077

Observation eac14bfc-28e3-48c3-8ac1-5053b2cfc05c · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Autoregressive Video Generation without Vector Quantization Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.907061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:07e58b4b09e296cb0e159865a9f90a5559533a15a2494b0890f44e81950f5eb9

Observation 1c6270e0-1e58-47e8-9950-89e91be7415e · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Autoregressive Video Generation without Vector Quantization Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.912883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:874bf340172849222f987003c8a9a90d1019d5e3f3bccb1a13c8a42badf8f956

Observation ac86e80f-5e89-4fa1-9a3a-120a335acb8f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Autoregressive Video Generation without Vector Quantization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.917441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:bb2448df0807ce06419c4a131efcb09ce202d1c275825fbee8a435fec9837684

Observation db8a4d00-dc3e-4601-86c0-1f31e894ec11 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Autoregressive Video Generation without Vector Quantization Emu3: Next-Token Prediction is All You Need

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.810241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:5c9cdc98f562b02a98d7aeb10d00e26849c335c4febeae544e5b3316b2e00c6c

Observation 43ecb43a-4735-4712-b75f-d7a2087f9dc3 · outbound

This paper cites Loong: Generating Minute-level Long Videos with Autoregressive Language Models.

Autoregressive Video Generation without Vector Quantization Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.816112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:9b95a270a276551a88d5de511ff49aeda89b50cef28cb4808358b6eff6788825

Observation ff83a8d2-cdcd-49d6-806b-3deebe214092 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Autoregressive Video Generation without Vector Quantization CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.822727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:589c599c5079d96826176c8ed26317cdadb86bfcad28ca7b59883570c603fec9

Observation a131a26e-66f1-48e2-a545-bc3cf6dc6b16 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Autoregressive Video Generation without Vector Quantization Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:07:39.829441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:1b01dfc4181ae3275452ff1ee5de2eeff64343d25cec26bd99219915beb7ba09

Observation 31424b6a-0733-4bc8-b814-63af7358e723 · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

Autoregressive Video Generation without Vector Quantization Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.836921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:a39a8b69e7072fb6bb3d2657c7a62619755ae0bc915a668eb422b8e90b27d3ad

Observation 6dcbc11a-5b14-4933-8af5-0903233ddb4b · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Autoregressive Video Generation without Vector Quantization Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.842642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:3e2ca4c6f24d4c848717286accb2f1af456021c9dee784bc00d562608f69b9e4

Observation bd32df8b-70c8-40e0-ad63-00943cb04283 · outbound

This paper cites Here, more implementation details and ablation experiments are organized as follows: • Architecture details of Scaling and Shift layer (Sec.

Autoregressive Video Generation without Vector Quantization Here, more implementation details and ablation experiments are organized as follows: • Architecture details of Scaling and Shift layer (Sec

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:07:39.938730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:fc69583c241ebfdc1c5c708c97c6994400a336b489e7e75b79fe594821350883

Observation 66e6bdcd-f00d-4246-acab-b239aa686ccc · outbound

This paper cites UpProjectorDownProjector LayerNormScale, Shift <BOV>outputs Temporal outputs Indicator Tokens Figure 11: Scaling and Shift layer.

Autoregressive Video Generation without Vector Quantization UpProjectorDownProjector LayerNormScale, Shift <BOV>outputs Temporal outputs Indicator Tokens Figure 11: Scaling and Shift layer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:07:39.924731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:80d19d008d1f285f4b1de52469edeebba97e91d157110b125b0edbb1dab9687d

Observation c190cdcf-6a73-432b-9591-74dd12e6d4bc · outbound

This paper cites While NOV A is already efficient in text-to-video generation, there is potential for further acceleration in the spatial layers.

Autoregressive Video Generation without Vector Quantization While NOV A is already efficient in text-to-video generation, there is potential for further acceleration in the spatial layers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:07:39.928240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:eb87bcdd8c6b11151335dcdac96def244e1b05e97b7807d1002db0e36c601891

Observation bc11d232-6cc4-48cf-80eb-8f3aea805107 · outbound

This paper cites This limitation may be attributed to our reliance on extensive web datasets, such as LAION and DataComp.

Autoregressive Video Generation without Vector Quantization This limitation may be attributed to our reliance on extensive web datasets, such as LAION and DataComp

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:07:39.931696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:23ef0fa469a6ce7eb7761160bdd22e69236e873323056283fbcf038b1fdda040

Observation 5480f284-237b-41c8-a924-9253e95f5750 · outbound

This paper cites In the foreground is the detailed, head-and-shoulders portrait of an elderly man with a long white beard.

Autoregressive Video Generation without Vector Quantization In the foreground is the detailed, head-and-shoulders portrait of an elderly man with a long white beard

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T15:07:39.935042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:770f12c1423907b0c4cbec603a2b692e4ebad5f4994adf44b46a1f6d09b98072

Pith citing papers

Observation 351ca4d4-d0e0-43e3-b8ea-cdfbfb372e8c · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Autoregressive Video Generation without Vector Quantization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:437f261a8c9e26c8ada7cdf35e40ed8205223b288b7db370ecabccbce59a4618

Observation 8ec2b89b-6235-47d5-84c2-3476b3164c24 · inbound

Unified Video Action Model cites this paper.

Unified Video Action Model Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T17:50:29.675358Z digest=sha256:6578bde26101591e4137b85b3a5d1c1db7e5883abcb7bc36d7270eec881acbfa

Observation 9351a4d0-8d54-45d2-af3f-bb19625df75e · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Autoregressive Video Generation without Vector Quantization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:0a7660a87b0fb3d6ad9c102ef1af3491bf7a24b5283ca36407fb3b77836f0139

Observation 37c23299-3855-48a5-ae6f-f3ebb672636f · inbound

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models cites this paper.

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models Autoregressive Video Generation without Vector Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:42.997910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:42.997910Z digest=sha256:26dc4670b09b79e2ba0bfb0d738e2cbe20969882d91f0c9b555c433b8785088a

Observation 5207c38b-a15d-4fc2-9cad-8683c84282c9 · inbound

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval cites this paper.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Autoregressive Video Generation without Vector Quantization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.529993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.529993Z digest=sha256:a8c5d6c039ee0df73fef278e6b69548e56fc863774093b1f75a93f2ebc03e307

Observation b0f32f7b-3dde-4e5c-b581-75845d2181e4 · inbound

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions cites this paper.

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions Autoregressive Video Generation without Vector Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:58.367808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:58.367808Z digest=sha256:118e9ee7f7f62a384e2f213ce189c094eb0306fde86c0e4560f258044502ca53

Observation 055f827a-7efe-48f8-85db-8aedcd009c02 · inbound

VideoMAR: Autoregressive Video Generatio with Continuous Tokens cites this paper.

VideoMAR: Autoregressive Video Generatio with Continuous Tokens Autoregressive Video Generation without Vector Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:26.930658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:26.930658Z digest=sha256:f84a56e91b45c061bf0d1c20762ee23d8f21a4b4480e7fcaa1865e58f7d2bbb3

Observation fa2b8a6e-77da-4bcc-bbff-be3e0c57eb34 · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.797784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:ede823b26b445d207de16be0acbce912f5571de622eec843f8a16fedc8018453

Observation 63c6b614-f7fe-41a4-9e96-bca4458caee4 · inbound

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations cites this paper.

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations Autoregressive Video Generation without Vector Quantization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:12.621023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:12.621023Z digest=sha256:d832aed08ffc7da26c3ac039f6670026211f35118326abd52f5870efe15d9922

Observation 2d1a1932-63a9-4f17-afee-c4073be0e256 · inbound

Geometry-aware 4D Video Generation for Robot Manipulation cites this paper.

Geometry-aware 4D Video Generation for Robot Manipulation Autoregressive Video Generation without Vector Quantization

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:14:27.855764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:12:39.787489Z digest=sha256:1462bb5b12ff114be968f7dcaa72fe9e9d748a951aeb53e70237f243024858fd

Observation 54f45424-b34c-4608-89b0-aff910a20f3f · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset Autoregressive Video Generation without Vector Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:54.436324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:54.436324Z digest=sha256:948285820bb8e7fc46f474fcd78d8b2636a793f7b6d818edf2b0fd76aca5b4b0

Observation 44129b92-9419-4a66-8bf1-88f408c28577 · inbound

Energy-Based Transformers are Scalable Learners and Thinkers cites this paper.

Energy-Based Transformers are Scalable Learners and Thinkers Autoregressive Video Generation without Vector Quantization

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:37.166568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:37.166568Z digest=sha256:7d2d5cb3b0f9955d32b910fa750ed37cda98ab8b50ce0ff799647d552f3455df

Observation 6fc87888-0d09-4247-820f-409d8cb5825b · inbound

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation cites this paper.

SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation Autoregressive Video Generation without Vector Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:43.278789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:43.278789Z digest=sha256:e12d374465754a09771ed2efd33a56283a4ce096be20fdf61327ab447acce5ff

Observation 6df5ccd9-a5a0-429a-8c36-59d1980debda · inbound

EndoGen: Conditional Autoregressive Endoscopic Video Generation cites this paper.

EndoGen: Conditional Autoregressive Endoscopic Video Generation Autoregressive Video Generation without Vector Quantization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:34.850039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:34.850039Z digest=sha256:9d15078b83d8633abf6e762e02e53dfe3c7b987284bf07616738481cffe28fdd

Observation 0037e690-690b-45a8-97cf-52e456ca73c9 · inbound

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation cites this paper.

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation Autoregressive Video Generation without Vector Quantization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:35.528142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:35.528142Z digest=sha256:714649710d0bfef0438e9bbe22f9c0c621e3a7bc38425925fe2a048bba4795bc

Observation eb7614c7-ce69-41a6-b982-5ce792142903 · inbound

Time delay as the origin of oscillations in anodic Si electrodissolution cites this paper.

Time delay as the origin of oscillations in anodic Si electrodissolution Autoregressive Video Generation without Vector Quantization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:28.990588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:28.990588Z digest=sha256:b759dfe614f6a5f1616e300f5269155894b98f970bde29683470161a79240fa9

Observation 2495bef3-5bb7-47dd-8736-453fa1f105bd · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Autoregressive Video Generation without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.606456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.606456Z digest=sha256:bd20f577df7ee02bd7d4552c31c8382dd01fffffcf0fd829d45b9e98b9fb0d13

Observation 4b549c20-3f91-4f77-8651-04a44c2cc22b · inbound

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation cites this paper.

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:19:10.290358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:19:10.290358Z digest=sha256:104d0255b943c191f1be68feafed1f3214e9b6b9d8733b30e24332a6f2be2c85

Observation 94c79c94-816a-42fb-a940-c0c35e37a77b · inbound

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking cites this paper.

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T16:43:50.095402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:43:50.095402Z digest=sha256:8a88f5d32f73114220c4c8370b6e6200751bc5dfdf189dc1b9529803901e1f1a

Observation 0919221f-f361-4d8f-96f1-cda18cda724c · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation Autoregressive Video Generation without Vector Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:2d365d690bc9d32e75be15f2b33d462e97203141a82eb03466862d2fd45cc52e

Observation 751c4310-f2ca-4380-9db0-e277bdeec8e1 · inbound

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation cites this paper.

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation Autoregressive Video Generation without Vector Quantization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T18:17:54.943863Z digest=sha256:b9985deeb94eeae6714c542f9adf4aa71b3b378f5db73fc1c98865c3162978ef

Observation 0cff635c-11b3-461d-a07b-1954c6bd875f · inbound

Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation cites this paper.

Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T18:30:02.314605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:30:02.314605Z digest=sha256:3c6a390383aaa8c33662faee5a359e10ae649e6acc4e73666b620fa18f669a47

Observation 42af0742-2a6e-4b11-a7e6-0cc6433a0b69 · inbound

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling cites this paper.

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling Autoregressive Video Generation without Vector Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T15:46:06.827425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:46:06.827425Z digest=sha256:5b88561557e4776f2c26047c6e396b14d071b4ad3df2e8801aa00596203ee4b9

Observation ba79bc34-0b9d-445a-8fa8-f265657c0f59 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Autoregressive Video Generation without Vector Quantization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:32:01.700503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T17:32:01.642256Z digest=sha256:d7510787daa871b2ca45001716d440dad7be041cb555fb4d32e74d5a496a194f

Observation 125742e8-fed2-4658-9a80-29c6c6a1c0c0 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Autoregressive Video Generation without Vector Quantization

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:51:29.724763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:48:00.633421Z digest=sha256:94d4e2f3ef1c7b78c72f9db3151a6ee5b78b2fc3d5acd03d7356c742f7a53a14

Observation c1cc0199-19f8-42eb-ae64-41afddfd3ae6 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Autoregressive Video Generation without Vector Quantization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T05:30:24.520144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:30:24.520144Z digest=sha256:efdb6240ae03ffa321f978b7faebc0fc286597e6800dcb41b6cd3a92102a2149

Observation 1c5145e8-2e33-46cc-bc1d-60a8dfd6993d · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion Autoregressive Video Generation without Vector Quantization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:788c7c01783259d8b3acab226c4cdad0716ffa4ca59dc1e420f231a9d6db9cbe

Observation dbbbec27-db6d-43d8-b957-1c4d35bc4696 · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Autoregressive Video Generation without Vector Quantization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:997aab1c1e74608c47fc9059a4b4da7e66d671b710ff6606660e0365e70326b8

Observation b04dd9ce-9fd5-48ba-8d18-722905a52601 · inbound

GeoWorld: Geometric World Models cites this paper.

GeoWorld: Geometric World Models Autoregressive Video Generation without Vector Quantization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:40:03.340552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T11:39:15.308355Z digest=sha256:dda9583f4d2d1ee709a1123ef02faa2cea85645dbe4c5134535fccb936eb23a9

Observation cfa52c45-6ae9-461b-abd8-29f00bcf1faf · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:896d5c116225e40ca86f63c0becdca2a9d6a83fe86c99d5c4808434145a7e6a1

Observation d95a30d8-84db-4047-ac6b-22ab82f701ae · inbound

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion cites this paper.

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T21:09:36.040879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:09:36.040879Z digest=sha256:1ecb23b263c2c5b162a1af19efa300c4a0024274f46182c456c3643be01c5b89

Observation 87d909ab-6efa-441a-b843-23996e9273ea · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Autoregressive Video Generation without Vector Quantization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:8fb1840d9f1821a024ebae9916c5938d3c5d14dd8d938aa74627b7a6b1d74557

Observation 961383c7-fd5e-4b28-89fa-2655eeaa25c5 · inbound

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation cites this paper.

Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation Autoregressive Video Generation without Vector Quantization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:10:54.363560Z digest=sha256:2f92c04b5aa2395850e5970fdd389f71b9b93efe9a1a68ff0f62d678f99c791e

Observation cd1d0b0a-49a1-47a9-b1f5-0a1f1c2bdfc5 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:41:23.057949Z digest=sha256:4096ac3d21d0c240570e213af0bd523668c2b085efc308d61399a708de909561

Observation 258d27ad-1aca-4546-a70e-80d5ffdfd2ec · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T21:00:35.123495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:00:35.123495Z digest=sha256:11fe937f2c81f9cca00b3babb0745cbf70763e3048e3c0d5a7c2227d6d5d9da1

Observation 754a4ad8-18e5-4294-87d1-b219f0a138f9 · inbound

Efficient Video Diffusion Models: Advancements and Challenges cites this paper.

Efficient Video Diffusion Models: Advancements and Challenges Autoregressive Video Generation without Vector Quantization

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:28:29.706249Z digest=sha256:812b197c9303fac1f4a4aa3e24d9fa7b543fa6d7cd5875b593a539f51aac3351

Observation 7bf1de98-dbd6-482b-bcfa-a72773fce9b2 · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Autoregressive Video Generation without Vector Quantization

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:32:54.235578Z digest=sha256:618be29e788e8aeb494c71e6811a8458bafca6dae0fa0c9b64a9627b410e981e

Observation 00f966d5-e693-450b-955b-7caf90ae087f · inbound

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models cites this paper.

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Autoregressive Video Generation without Vector Quantization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.773900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-05T11:32:38.636335Z digest=sha256:7e8c86fd584f610428fa09d1dccc091740c2db545bd42a1541aa5c3ec1fee33e

Observation 42985d31-453a-4623-b7a7-562518962d7a · inbound

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation cites this paper.

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation Autoregressive Video Generation without Vector Quantization

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:38:08.199955Z digest=sha256:8394915bd164514e914adcf101b39d633bf8d0a9cf97b8495193057a658990b3

Observation ad65406b-a55c-40de-8ae7-716a00eba102 · inbound

Stream-T1: Test-Time Scaling for Streaming Video Generation cites this paper.

Stream-T1: Test-Time Scaling for Streaming Video Generation Autoregressive Video Generation without Vector Quantization

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T18:16:14.985693Z digest=sha256:eeb07f9003316bc0820dbae3f395b6b6afc6c92d2c891a435b4593645aeaa120

Observation 6233f6b5-b622-43ca-a850-0f324940e4ef · inbound

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation cites this paper.

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation Autoregressive Video Generation without Vector Quantization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T13:28:05.540719Z digest=sha256:ab46fa878dbd50f38324fdded396c84f59218345f4a1b7813b1d5099c2a698db

Observation f26a2bc8-6088-477d-87f4-51b47084a11e · inbound

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation cites this paper.

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation Autoregressive Video Generation without Vector Quantization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:07:39.939813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:50:30.588220Z digest=sha256:a0858dd2a082d9d9826ca57eaef9fff7e346c28f9ed06d03a4af112f7825c21b

Observation cbace986-40fa-46a0-af03-ccc006540f66 · inbound

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation cites this paper.

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation Autoregressive Video Generation without Vector Quantization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:04.401478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T20:59:34.847496Z digest=sha256:603273506e270394185956eb3a4baeb08d40f65acd6de54e7169d89993730fb8

Observation 232d82d5-f957-450c-8fad-e9b2a5c9c16b · inbound

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos cites this paper.

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Autoregressive Video Generation without Vector Quantization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:13.907649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:56:16.054496Z digest=sha256:d7d0e2f5502f2ed21265bfcf804d64e5974eb015c7f56b4b7041177c589ae074

Observation 17874482-b131-4eae-9f49-992b5b751083 · inbound

Tempered Self-Similarity Alignment for Physically Plausible Video Generation cites this paper.

Tempered Self-Similarity Alignment for Physically Plausible Video Generation Autoregressive Video Generation without Vector Quantization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T11:44:38.368550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T11:39:06.597513Z digest=sha256:1bb7e1a51ea7a4fd8459a7f26bdf1af4a729a0df81c9f52944a547a02f050ebb

Observation 72131ca2-d689-4239-85f5-040d1357d81a · inbound

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation cites this paper.

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation Autoregressive Video Generation without Vector Quantization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:33:15.185456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T08:30:00.438334Z digest=sha256:4fa2e8a7243bdb46a94aa2c1137edefb234166ca1bb11309001d47d3c3701ad9

Observation 80cdf9a3-9b43-4ca3-aa06-69f76d952d08 · inbound

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion cites this paper.

VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion Autoregressive Video Generation without Vector Quantization

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:23:15.470700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:17:30.660009Z digest=sha256:a1bb4b5a221b80b56626d21d697682fa9837abec35db637cdd2c349f61d18881

Observation 4ef15f71-4345-417a-920f-9c3e1eaf58ac · inbound

OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation cites this paper.

OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation Autoregressive Video Generation without Vector Quantization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.442846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:56:40.001739Z digest=sha256:b8c5f6889c7a0dac74be8e7b058e7a0723d1a734fd7245c182bac418be282f63

Observation 1cea5fc4-4361-44ea-b3f7-9724a7fc8d2a · inbound

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models cites this paper.

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models Autoregressive Video Generation without Vector Quantization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.548835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:56:21.783415Z digest=sha256:10c0b54cd2895b97550d1d53419ef789df028193588555b6a1c31a243800ce3b

Observation 4a955319-5665-4808-9d7b-2ea813439bed · inbound

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation cites this paper.

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation Autoregressive Video Generation without Vector Quantization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:36:54.891991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T03:27:02.232039Z digest=sha256:c024c40299a19f0a9c492704d2f199821ff9d9ecc8ace67e682eafaba8264894

Observation c0f77ff5-ed0e-4565-8463-48316811bc7b · inbound

RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling cites this paper.

RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling Autoregressive Video Generation without Vector Quantization

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:26:56.753440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:08:44.257127Z digest=sha256:1b82a6a95dbec1be9e36dcb7a95331d5adc2d11d3cd3540c529e2d8599b88b38

Observation c03865b5-3d37-4c5e-a2c1-3ca446ac045a · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey Autoregressive Video Generation without Vector Quantization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:09:35.537094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:b944d73e9489644e8c20de2883356fd9990d3c714249b0f3c0abd031cb16bc1a

Observation e8c8ac01-1033-4e6d-b31e-652dbbf30444 · inbound

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control cites this paper.

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control Autoregressive Video Generation without Vector Quantization

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:43:51.126020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T05:03:56.624536Z digest=sha256:6fd6b23a89325f1395467efab03f08e9d2a12a89650c3b5c5bc98fc532d3e9ce

Observation b372428d-11ca-40b9-b4f2-e6110e65fb3b · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Autoregressive Video Generation without Vector Quantization

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.227463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:05698307134aa2b6ba48257190e6d20536ded0500c4b3edf24a58032a789c8a5

Observation c999d9ed-8cc8-44a0-9c0c-b14696254935 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Autoregressive Video Generation without Vector Quantization

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:25:41.898990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:6f9ad0e1e936b080430e60e78f5cd13c6cb036ddb00b7c743b03001510e51657

Observation 47d66d47-afbf-432e-b25a-c946ecd9fb70 · inbound

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model cites this paper.

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model Autoregressive Video Generation without Vector Quantization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T01:59:48.820112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:59:48.820112Z digest=sha256:a9f884ff744712dc8efc582a4d25733d215c155757156828cf92f8b117943adc

Observation 1c05b8ef-2299-482d-bd1b-e20af8af78c4 · inbound

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising cites this paper.

Enhancing Video Physical Consistency via Role-aware Joint Training and Modality-decoupled Denoising Autoregressive Video Generation without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T15:44:22.210969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:44:22.210969Z digest=sha256:6bee27d96a02382695ed5eadeb3c56c1489445f36fd6104229f2d73e2d0b3f1c

Observation 493494b3-5002-4a39-a219-2cc78267f160 · inbound

SGA: Plug&Play Geometric Verification for Educational Video Synthesis cites this paper.

SGA: Plug&Play Geometric Verification for Educational Video Synthesis Autoregressive Video Generation without Vector Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:07:49.805555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:07:49.805555Z digest=sha256:b0b651bde4e61f1597ad17880ff2baa1cff5b74684085f3277817c934be26ee4