Pith. sign in

Paper Citation Record · LEDGER

Taming Teacher Forcing for Masked Autoregressive Video Generation

As of 14 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 7 inbound Pith citation observations for arXiv:2501.12389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12389 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:16:47.640589Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:54:58.686040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:33:15.180082Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8217c3c-1ec2-4900-92fd-e6926197cce8 · outbound

This paper cites Video generation models as world simulators.

Taming Teacher Forcing for Masked Autoregressive Video Generation Video generation models as world simulators

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.399672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.399672Z digest=sha256:40602f06b5081ac56e4a262c1acc3a4ae95a6f7b7dc11bb93a3315a75de96448

Observation ffde491b-2e00-4ec5-8920-a7b15e856a4b · outbound

This paper cites Ge- nie: Generative interactive environments.

Taming Teacher Forcing for Masked Autoregressive Video Generation Ge- nie: Generative interactive environments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.220054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.404985Z digest=sha256:3d3c80801fda533dfd8233c0ecb334fba6990a4279163de7649968b4e92a1a94

Observation a0ad7479-a4f6-4d31-92cd-85adb495846b · outbound

This paper cites A short note about kinetics-.

Taming Teacher Forcing for Masked Autoregressive Video Generation A short note about kinetics-

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.410911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.410911Z digest=sha256:def2afd51a84243f2021956fe1116fac955bda7e71e6ed2678c08d3652941c2a

Observation 0f6a07e0-4a77-47f0-90b2-a47632a79e8b · outbound

This paper cites Maskgit: Masked generative image transformer.

Taming Teacher Forcing for Masked Autoregressive Video Generation Maskgit: Masked generative image transformer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.201480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.423531Z digest=sha256:5a4aba5909086c0009fa255e2d7ad9ee9d1074ea73442d9eab64681b423775a3

Observation 732eac7f-53c7-4c1b-a645-30a724bc8542 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Taming Teacher Forcing for Masked Autoregressive Video Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.428417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.428417Z digest=sha256:c3adefd89f4d4eb51553dcb87a42491d2f18db705bb3b07eaadd0236c1468fee

Observation c65c365c-b92e-45bf-b664-3df86884fa53 · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Taming Teacher Forcing for Masked Autoregressive Video Generation Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.433670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.433670Z digest=sha256:b1f0c669188ba00d4e8547febc45f58593a39f6799ed63a53528430247fa5081

Observation 4938ebbd-83b1-41e1-94ed-0f664cc0ba32 · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

Taming Teacher Forcing for Masked Autoregressive Video Generation Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.438930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.438930Z digest=sha256:c6fdd94e7c98971783842fe070cf0ac84abbc1551a7d27c8e47cff0299b45999

Observation cbcca705-c57c-42c6-801d-1a7716c2711d · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

Taming Teacher Forcing for Masked Autoregressive Video Generation Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.443669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.443669Z digest=sha256:149fa796d1df5aaff5495b14aa990f4d5124c77f00361abe1395b52478ca72ea

Observation 1445b4f2-170d-4b06-8734-2ab442a3606c · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for llms.

Taming Teacher Forcing for Masked Autoregressive Video Generation Model tells you what to discard: Adaptive KV cache compression for llms

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.183696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.448666Z digest=sha256:4d6bfb9f4683e10b2b4314aabb8435e38d34d288cf5a5775394f8a12fd2d2bcb

Observation 3747f35a-5d14-43f0-91d7-ffaa29e449cc · outbound

This paper cites Courville, and Yoshua Bengio.

Taming Teacher Forcing for Masked Autoregressive Video Generation Courville, and Yoshua Bengio

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.173588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.453799Z digest=sha256:341e4dd3b6ffde9a44ee5c9ff7af736745f49f871760a4eafbb3363427ab463a

Observation 5c2b762b-3225-4277-8b34-04e9748f73ae · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Taming Teacher Forcing for Masked Autoregressive Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.459726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.459726Z digest=sha256:5e02788e26f4e3c60c26d3c80c8a112e8dabf615b9887501bad98f876e13c4fb

Observation 7d6aa9f6-43dd-4825-9437-05471e6be217 · outbound

This paper cites Girshick.

Taming Teacher Forcing for Masked Autoregressive Video Generation Girshick

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.162173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.466231Z digest=sha256:a0717a624507de4d600c0d044d6767083067ce2df5118911ae39895168a08df0

Observation 9f9aade9-01a3-42b4-b197-10715474df14 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.471762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.471762Z digest=sha256:c8160e4c5123bfcd3a585f1ac8a47a354836277f1ca2addade7d7e3a75ea81c7

Observation 1fef39fe-5590-46f6-86dd-06d1877d2ddd · outbound

This paper cites Denoising dif- fusion probabilistic models.

Taming Teacher Forcing for Masked Autoregressive Video Generation Denoising dif- fusion probabilistic models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.150461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.476470Z digest=sha256:305362d1e965d6322e44a8124a81d782fc62780f8284cb1a0286d8cd46b38a16

Observation 20159b2a-dec4-49b3-a901-7709bb2635ef · outbound

This paper cites Video dif- fusion models.

Taming Teacher Forcing for Masked Autoregressive Video Generation Video dif- fusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.138967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.480978Z digest=sha256:1462e1e5a77a27434922340e3b881b758ec64d83da38e4c33f443f6c80722dea

Observation b9eed922-6652-4e54-b265-d2e2379095d5 · outbound

This paper cites Scalable Adaptive Computation for Iterative Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Scalable Adaptive Computation for Iterative Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.485858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.485858Z digest=sha256:542021f0edcc5d5444528253252b7822374362ed8808858427615ae7d03a9ef9

Observation 360b747a-18fe-4dfd-b74e-0848a829e321 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.490673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.490673Z digest=sha256:f3c1272560d9a1a92e73dd967cd42ec664f6998fe89d0e962e2447446e753437

Observation 81b6d41c-4923-40d1-8397-0282b6c91ce6 · outbound

This paper cites Dart: Noise injection for robust imitation learning.

Taming Teacher Forcing for Masked Autoregressive Video Generation Dart: Noise injection for robust imitation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.125866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.495655Z digest=sha256:40b0e4ee6c47722f14a090c857275508cb5ec2301732b35ce6bb79b1c56577ca

Observation fcb101ee-7e9c-4071-bc4e-82a394b57b61 · outbound

This paper cites Autoregressive image generation without vec- tor quantization.

Taming Teacher Forcing for Masked Autoregressive Video Generation Autoregressive image generation without vec- tor quantization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.112751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.500178Z digest=sha256:8ef74c7a24ba5d920213f61949493ef159af7e75ba7f9611b68af396471398eb

Observation 4bc11b39-4bc1-4e78-b94c-91c8ebb49620 · outbound

This paper cites Keep the cost down: A review on methods to opti- mize LLM’s KV-cache consumption.

Taming Teacher Forcing for Masked Autoregressive Video Generation Keep the cost down: A review on methods to opti- mize LLM’s KV-cache consumption

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.098726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.505391Z digest=sha256:9fee0ef8a848857e8fa6f88315e3910fd8fe8074d25783e180d2a943717acc4b

Observation c565ebda-6104-4ab8-8f56-29cb9a5cc815 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.509903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.509903Z digest=sha256:3289d86206fc422246cf417e74eb63280d906f373e4affa18d20994e1518b6ea

Observation c9bd6572-827e-4f82-9b6a-c9a64e551df3 · outbound

This paper cites Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.515030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.515030Z digest=sha256:072bd3def2e4c5bd4dc666aa9ffe2c7d66bdeea00e2211c0138af425d258bf41

Observation 8ca12290-9639-49e7-931b-b5d69692e94d · outbound

This paper cites Cosmos tokenizer.

Taming Teacher Forcing for Masked Autoregressive Video Generation Cosmos tokenizer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.085171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.519984Z digest=sha256:a49613c671559adbb7a89d009bd953c3fed3b9269849390ee37556d9639fe4f5

Observation 5f30d9b2-3407-4071-95f8-236f20103467 · outbound

This paper cites Improving language understanding by gener- ative pre-training.

Taming Teacher Forcing for Masked Autoregressive Video Generation Improving language understanding by gener- ative pre-training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.071303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.524356Z digest=sha256:55b07e1c0a770cd34153054d03f6255c742d525699c695a3d177d7751707e3da

Observation 4841542e-b71a-45a4-ba93-583127cb2c90 · outbound

This paper cites Language models are unsu- pervised multitask learners.

Taming Teacher Forcing for Masked Autoregressive Video Generation Language models are unsu- pervised multitask learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.057648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.528509Z digest=sha256:8f2d4984c1f0d0c1a1843c66a59ebbad64ab3f0232ce1533b0a53759bbb0fbc6

Observation 0fd74360-c00e-47ac-959d-3f07d4277060 · outbound

This paper cites Zero-Shot Text-to-Image Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Zero-Shot Text-to-Image Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.533348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.533348Z digest=sha256:1b567a4539af4f73ac27fcf3e8eee306c279976470ef279f1970333f32bafbab

Observation 4cfb776c-15b3-4afa-b194-77d3ae117dce · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Taming Teacher Forcing for Masked Autoregressive Video Generation High-resolution image synthesis with latent diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.044619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.538415Z digest=sha256:0328d2d18546ac0df8278d3e31c48c3014046a4d4b897086f9493e4e0d992c6d

Observation 1b4086cd-b949-400e-bb3e-e2f776919639 · outbound

This paper cites Generalization in generation: A closer look at exposure bias.

Taming Teacher Forcing for Masked Autoregressive Video Generation Generalization in generation: A closer look at exposure bias

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.031417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.542539Z digest=sha256:aeb5c0a4eefe56b38dde14d1f867f87ac5f201e4412f9ed54effa2ff00ddee69

Observation 2e1a87e8-e7f9-4084-a594-9a27adf35bb0 · outbound

This paper cites Mostgan-v: Video generation with temporal motion styles.

Taming Teacher Forcing for Masked Autoregressive Video Generation Mostgan-v: Video generation with temporal motion styles

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.016832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.546816Z digest=sha256:9183a3622886f00b03e6d98028bded7711f42217d4683d9b6fa5baf75593e77c

Observation ab58ef13-103c-403c-a0e1-709477ce37cd · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Taming Teacher Forcing for Masked Autoregressive Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:48.003201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.550916Z digest=sha256:6cddd4569584f96e56a22d8a6640aa442be8c3296fb0c8cf5c9db3087e37e650

Observation 712eb7b0-13ba-4a16-9873-7280b52418ca · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Taming Teacher Forcing for Masked Autoregressive Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.555283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.555283Z digest=sha256:5dc3e5f6976ebe95970fcfc432d8fab343752385e73bb34b60b046e7558e82f2

Observation 781189e5-eab2-4e27-972f-02b01933f344 · outbound

This paper cites DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis.

Taming Teacher Forcing for Masked Autoregressive Video Generation DiM: Diffusion Mamba for Efficient High-Resolution Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.559474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.559474Z digest=sha256:93c4ec0fa39ea7fc062301656a47841d1e6f289540f8d43c55f78ea5c9141a2d

Observation b7f2e466-cb66-47fe-b237-27131a2c7756 · outbound

This paper cites A Good Image Generator Is What You Need for High-Resolution Video Synthesis.

Taming Teacher Forcing for Masked Autoregressive Video Generation A Good Image Generator Is What You Need for High-Resolution Video Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.563491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.563491Z digest=sha256:89797befcad7122e3e71b3d93857943aad15fe083642bbee498d747209c7afab

Observation 620a12e9-8525-42e0-833a-bdddd904b986 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Mocogan: Decomposing motion and content for video generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.568171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.568171Z digest=sha256:2b3a1fcc83b9de06ec3bd9ba6796413d97aa49c1196280dcc54a6b14cabc744a

Observation eb982e26-a697-4d87-bdb8-d9b38f1bf0e4 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Taming Teacher Forcing for Masked Autoregressive Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.572241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.572241Z digest=sha256:c925430f3871f6f3647bcdc48d7ba845f58bbdde8ba95aac6a823934e35f6572

Observation 55f9b719-ed24-401f-88fb-51c87abc9ecd · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Taming Teacher Forcing for Masked Autoregressive Video Generation Diffusion Models Are Real-Time Game Engines

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.576958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.576958Z digest=sha256:5ff5ac4ff1f6400f7b621e219a4512ff4c466113307f0a102fe3f1a9cbebcb3f

Observation 6ccbc150-f911-4b07-8925-2c523cbc6894 · outbound

This paper cites Neural discrete representation learning.

Taming Teacher Forcing for Masked Autoregressive Video Generation Neural discrete representation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.983020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.581515Z digest=sha256:27250bd9e4a9716d88762c543cc4cda0f67f0dde1e59ceffc9d57ce2f9e3eac7

Observation d09973aa-30d3-4d5e-a283-40586f223a62 · outbound

This paper cites Attention is all you need.

Taming Teacher Forcing for Masked Autoregressive Video Generation Attention is all you need

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.969438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.585602Z digest=sha256:dc3b42ae657a9c6fe2420403139ff3cb11646e695b42401f6089c5f4187e9494

Observation 05191d2d-92ee-447f-bcbf-bb57554b3d3a · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Taming Teacher Forcing for Masked Autoregressive Video Generation Phenaki: Variable length video generation from open domain textual descriptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.957586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.590142Z digest=sha256:abd11617a6501dac2c4b748e7c6540ba4c51882805e9138cca4e285e4e9a2c8a

Observation 80260792-2f5c-477d-a704-a08243a0995d · outbound

This paper cites Predicting Video with VQVAE.

Taming Teacher Forcing for Masked Autoregressive Video Generation Predicting Video with VQVAE

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.594537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.594537Z digest=sha256:a2243b4a074306314a59208e795ee3897e8181b08c89d08ef95a16e620d6cec2

Observation df6172df-8e15-4382-aa81-bcc7abf63b89 · outbound

This paper cites OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.599371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.599371Z digest=sha256:ccf6e0de590b24955a79dba7784cd2338337c628788e961b7985a7338b3a03c4

Observation 7f39279b-c28e-4617-bd95-c2194542e5bf · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Taming Teacher Forcing for Masked Autoregressive Video Generation Emu3: Next-Token Prediction is All You Need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.604165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.604165Z digest=sha256:3064fa2b6711a85f6c187bc71950b023a948eb7da92e4131e796955e9529bcab

Observation a85941ed-ffd9-4442-a632-1245d13c93db · outbound

This paper cites Williams and David Zipser.

Taming Teacher Forcing for Masked Autoregressive Video Generation Williams and David Zipser

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.945800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.608815Z digest=sha256:caf8d9bd7c9ee222d43ca20331a418c0ca56d68e2653bf8e74dbfb5eda69e23a

Observation 8e5f15cc-13c9-4714-9b9c-a9ba42b1bce1 · outbound

This paper cites N ¨uwa: Visual synthesis pre- training for neural visual world creation.

Taming Teacher Forcing for Masked Autoregressive Video Generation N ¨uwa: Visual synthesis pre- training for neural visual world creation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.932679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.613408Z digest=sha256:e85990c3c88fac0b62d1a7294830ddc601323af664a353fb7f5a5ac28f5049b1

Observation 1a778c2c-eb55-42ff-a069-2a61d8e47074 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Taming Teacher Forcing for Masked Autoregressive Video Generation VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.617853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.617853Z digest=sha256:e464192827539b4c9d4f2c37b5b1680b961d6c96541991d7437de141b88f1a11

Observation f9606f3f-391e-49cf-a136-982f5ac0cdef · outbound

This paper cites Magvit: Masked generative video transformer.

Taming Teacher Forcing for Masked Autoregressive Video Generation Magvit: Masked generative video transformer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:16:47.919331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T17:16:47.622561Z digest=sha256:003e05f2e773941b1daca0a627a10bd5404ef927224c6227ce0b03974bfe175b

Observation 02f907dd-8d63-440c-8fb2-970dcbbf1b5e · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.626847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.626847Z digest=sha256:d87106377a9b160f32ad3f117cef9fca2b7e26e4d8344f0e7f2e2bb528a1e2cb

Observation 38ab387c-7f8e-4f81-901f-9c640634f9eb · outbound

This paper cites Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks.

Taming Teacher Forcing for Masked Autoregressive Video Generation Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.631553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.631553Z digest=sha256:dd566a6c5123a295c019345b0dfe1de806205f370ba2b40d328574049490d57b

Observation adb9fa19-0050-4e3e-9ea8-a9e00091d17f · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Taming Teacher Forcing for Masked Autoregressive Video Generation Video probabilistic diffusion models in projected latent space

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.636085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.636085Z digest=sha256:03e618292d1647dc04a29dd4df2eab24b7f588fe65313e70a312754d2b904486

Observation b6517ac2-4092-4ea4-b88b-0c4730de4ddd · outbound

This paper cites Bridging the Gap between Training and Inference for Neural Machine Translation.

Taming Teacher Forcing for Masked Autoregressive Video Generation Bridging the Gap between Training and Inference for Neural Machine Translation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.640589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.640589Z digest=sha256:24babc5a5d943ec8187ecafc98048532dcae2d85306492d46dc873d7d1ec1a37

Observation e6e7aa42-9bc0-4694-a878-c8052c99a8b8 · outbound

This paper cites A Short Note about Kinetics-600.

Taming Teacher Forcing for Masked Autoregressive Video Generation A Short Note about Kinetics-600

Reference 600

Resolution
unresolved
no resolver link, observed 2026-08-10T17:16:47.415805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:16:47.415805Z digest=sha256:51eeea23c918dc650e4402101d07c0e3df43e5bc03a1b0546f994ed28bcd5d55

Pith citing papers

Observation 345c7c81-7ca0-4f60-9dcd-030355ebec85 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.524350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:33b99bf0aad5fae6b228bf8e0f2df4ebfa607fa94fc9210498935b4a9fdf5027

Observation 55028210-12d4-41f0-a85b-e1e6bffa0950 · inbound

Long-Context Autoregressive Video Modeling with Next-Frame Prediction cites this paper.

Long-Context Autoregressive Video Modeling with Next-Frame Prediction Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:05:17.429070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T23:05:17.201790Z digest=sha256:3a97de01b4d19d2fcb532d162e8f399f9c795d6f8d0cf2a67e81ddea90dfb966

Observation 3c4f6c91-da64-4f9a-bd15-5b1934e03dc2 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.751003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:beaa105593cf5a5dd25cf2131fbd6695d71922ed465ee18ec68083d723389638

Observation b302fce3-bb57-4116-9db3-a28e75632476 · inbound

VideoMAR: Autoregressive Video Generatio with Continuous Tokens cites this paper.

VideoMAR: Autoregressive Video Generatio with Continuous Tokens Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 37

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:30:29.656888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:29.656888Z digest=sha256:63affe52ef1038b25d4ecc6186bba6af0b028b72633459ba1f9135cd1810554a

Observation 6fa80ef5-e219-4eeb-9890-eb40db8a9af9 · inbound

LongLive: Real-time Interactive Long Video Generation cites this paper.

LongLive: Real-time Interactive Long Video Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:52:59.443581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T03:52:59.287555Z digest=sha256:2df1bb720284e3bf7a82fc3c6a91670f1c2fd3fcc027728726a7e76a2a478918

Observation 802ff53e-027c-462f-90df-0204cb5e436c · inbound

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation cites this paper.

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.181604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T08:30:00.438334Z digest=sha256:cf5b605f45d9ee6bbf1898caec8e459af7f817ba7a25fc90213bd9892d8bd20a

Observation ea3bcdc1-ed91-4ee5-aa03-278c26125255 · inbound

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation cites this paper.

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation Taming Teacher Forcing for Masked Autoregressive Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:58.686040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:58.686040Z digest=sha256:a857db39691b350efee29b0a4552321dfac212cae3d6d282ec02330164e700a5