Pith. sign in

Paper Citation Record · LEDGER

RefTok: Reference-Based Tokenization for Video Generation

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2507.02862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02862 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:14.295583Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:19:17.796377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:10:50.188969Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 098bb933-9d9b-4ebb-aada-84f2a6363a13 · outbound

This paper cites Towards high resolution video generation with progressive growing of sliced wasserstein gans, 2018.

RefTok: Reference-Based Tokenization for Video Generation Towards high resolution video generation with progressive growing of sliced wasserstein gans, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:22.914818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.129043Z digest=sha256:4823c0c42ef0a1d943a86e36f55872bad05b2c20edd703a04b598966fbd6f57d

Observation d3ab17fe-85bf-4262-9272-a154f76a4fa5 · outbound

This paper cites Fitvid: Overfitting in pixel-level video prediction, 2021.

RefTok: Reference-Based Tokenization for Video Generation Fitvid: Overfitting in pixel-level video prediction, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:22.782558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.188041Z digest=sha256:1d29b4c4e9ba7383a384ea0649ff8da57439b44d429f52e2bc5b84d85f5e701e

Observation 03ca2c76-54a2-48a9-ab82-576afa890293 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

RefTok: Reference-Based Tokenization for Video Generation Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:22.622548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.230753Z digest=sha256:05b425aa7c2107dc1b5d3f3d1b6cdfbe8b94e09e000daa30dbba541c7815cdad

Observation 378dbeaa-440a-48c6-bef9-fc7d44d73fc2 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023.

RefTok: Reference-Based Tokenization for Video Generation Align your latents: High-resolution video synthesis with la- tent diffusion models, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:22.455985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.294395Z digest=sha256:2c6f01c5c42e1ff502041094172c7a4e317a10247dfa1868549fb258a05b9552

Observation 7b97bbcb-b6fc-44c8-a34d-0ff7184ee671 · outbound

This paper cites Efros, and Tero Karras.

RefTok: Reference-Based Tokenization for Video Generation Efros, and Tero Karras

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:22.279970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.368764Z digest=sha256:21ef1abd66e3fac6a74b165426e2c67353dc4ab21544305dad022cde7ff95598

Observation da83308a-3143-401d-b3f5-d2a632de1663 · outbound

This paper cites A short note about kinetics- 600, 2018.

RefTok: Reference-Based Tokenization for Video Generation A short note about kinetics- 600, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:22.112017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.461138Z digest=sha256:7f913505196c8561010ce139214cb025a4e82929113b246f7da3d9ea10a071c7

Observation 429cc97a-4f2d-45e0-911f-edd4032faef0 · outbound

This paper cites an unresolved cited work.

RefTok: Reference-Based Tokenization for Video Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:24:21.979392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.545679Z digest=sha256:e1bef0edb2af6dff1d7cdb387f04f97cd5e4c6623eb28fc71e0e129016731876

Observation f66c88fd-2c5f-47df-8da9-93cc55eb4f82 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.

RefTok: Reference-Based Tokenization for Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:21.792845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.634361Z digest=sha256:550e86f6813c870b5e4c36c35ab1a7cd2865b49ae2cf2b20f76f8c93588ddc68

Observation dea1d19c-7803-4338-86a8-d3830c2b118c · outbound

This paper cites Adversar- ial video generation on complex datasets, 2019.

RefTok: Reference-Based Tokenization for Video Generation Adversar- ial video generation on complex datasets, 2019

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:21.595627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.722749Z digest=sha256:570f56c8d29492605301d924ad7666201753692b87e1ab58a9c089c3d0400700

Observation f4fa4c16-a335-4d62-a831-9bfd4f057a23 · outbound

This paper cites Av1 bitstream & decod- ing process specification.

RefTok: Reference-Based Tokenization for Video Generation Av1 bitstream & decod- ing process specification

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:21.356506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.790282Z digest=sha256:c69c99565117bdd654295b65f677e67481786b1977e4f30cab7403212d2f88d4

Observation d92bc3ee-ee49-4718-bf52-2e48c643b6f5 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale, 2021.

RefTok: Reference-Based Tokenization for Video Generation An image is worth 16x16 words: Transformers for image recognition at scale, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:10.863497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:10.863497Z digest=sha256:d4b5d621061f4b4cea0ad1accf157ce80a96eb5af76b52e2faf5fd179c00f272

Observation 5ff7664d-0455-4b15-ac70-557ba3977fa6 · outbound

This paper cites Lee, and Sergey Levine.

RefTok: Reference-Based Tokenization for Video Generation Lee, and Sergey Levine

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:21.150885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:10.946076Z digest=sha256:ac3256bb09c24dfd6a712bb2072bac54cc161c1a2fc1753daa260c2cc2e6b072

Observation 391dd248-6495-418b-8fd1-ae3e744a7f44 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

RefTok: Reference-Based Tokenization for Video Generation Taming transformers for high-resolution image synthesis, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:11.015793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:11.015793Z digest=sha256:fd328e3c19e6f947ccc6344623734bb4e0ecd9058dafcab60d38f74565398ef1

Observation 8090a8b8-c661-4472-819e-8426b31d3e38 · outbound

This paper cites Videoshop: Localized semantic video editing with noise-extrapolated diffusion inversion, 2024.

RefTok: Reference-Based Tokenization for Video Generation Videoshop: Localized semantic video editing with noise-extrapolated diffusion inversion, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:20.872740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.084282Z digest=sha256:1a1a89bf73b99d8fda75d10a2d48fe4d2f04a3bdaa0fe680a880ea54b44128f1

Observation e6c393a8-10c5-4827-ab10-b7df5784897e · outbound

This paper cites Rv-gan: Re- current gan for unconditional video generation.

RefTok: Reference-Based Tokenization for Video Generation Rv-gan: Re- current gan for unconditional video generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:20.624651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.175488Z digest=sha256:bb7624f7d75fac45a989c67ce8b1ac841900a0dd775eb5d35f469ffda88cc44c

Observation 4fa49dc7-472c-41da-9239-80804c63ffaa · outbound

This paper cites Masked autoencoders are scalable vision learners, 2021.

RefTok: Reference-Based Tokenization for Video Generation Masked autoencoders are scalable vision learners, 2021

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:20.258138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.239482Z digest=sha256:288893212745fcad3550f078f99bc7f8b89843c0a5d6ab1cff5c8afc33dd866f

Observation 29266b77-6317-4306-b2b4-44c4b5358082 · outbound

This paper cites Advanced video coding for generic audiovisual services.

RefTok: Reference-Based Tokenization for Video Generation Advanced video coding for generic audiovisual services

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:19.986902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.335555Z digest=sha256:68fb3195a15d776dd535c5380e6cff37cbab471fed917df6e66eb3b59c32b1c2

Observation 2d78a30f-5714-47e8-a60f-2e01219a423f · outbound

This paper cites Rehg, and Pinar Yanardag.

RefTok: Reference-Based Tokenization for Video Generation Rehg, and Pinar Yanardag

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:19.563711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.396147Z digest=sha256:21c074c17fa06e977a7fe73655cebb961d965479d0a93724f151473592d4763b

Observation 3a39cbb7-f6f6-441e-b937-aa0ab04a5e72 · outbound

This paper cites Lay- ered neural atlases for consistent video editing, 2021.

RefTok: Reference-Based Tokenization for Video Generation Lay- ered neural atlases for consistent video editing, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:19.195599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.531815Z digest=sha256:efa62052b6b49c84fd19f17de1206f43dfb5bf39216be8f80f87f2bd75e9c0ed

Observation 15439237-de57-4a82-96c3-9d568e36ce36 · outbound

This paper cites Auto-encoding varia- tional bayes, 2022.

RefTok: Reference-Based Tokenization for Video Generation Auto-encoding varia- tional bayes, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:11.605620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:11.605620Z digest=sha256:d4c225f5c4e425965d09fafccb6812de0cce438175e8da6f25de609a83f20c26

Observation c802dbd8-8428-48e5-b64f-651c3fdb54ae · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

RefTok: Reference-Based Tokenization for Video Generation Ross, Bryan Seybold, and Lu Jiang

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:18.908936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.688637Z digest=sha256:f52ebed2a4704fa1abc3e794757c2f99bbd689e2f39893ec149aaa5b4c7fbe4a

Observation 459f9262-0c5a-45c5-bb54-178ed5399bcf · outbound

This paper cites Autoregressive image generation without vec- tor quantization, 2024.

RefTok: Reference-Based Tokenization for Video Generation Autoregressive image generation without vec- tor quantization, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:18.586615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.767157Z digest=sha256:1baca15b8c72b99763e18d3acf57681d36d061e2ed825f6e5763ec83e3848504

Observation e582abfe-66f5-4ee8-b0a2-10da7bfe9053 · outbound

This paper cites An algorithm for vector quantizer design.

RefTok: Reference-Based Tokenization for Video Generation An algorithm for vector quantizer design

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:18.267472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.846367Z digest=sha256:2b1f2b19b7b37b3712bc602fc09b7aae3422572419159520ed93cb891041c717

Observation 10c39e5e-cf17-4646-8f6e-58613dc71705 · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis, 2024.

RefTok: Reference-Based Tokenization for Video Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthesis, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:18.006224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:11.919606Z digest=sha256:597572f1a2b3faabce33acd480dd7a8394848e00299a1e38d59aedbed94c2422

Observation 75802674-e965-4252-a370-78facf87798a · outbound

This paper cites Finite scalar quantization: Vq-vae made simple, 2023.

RefTok: Reference-Based Tokenization for Video Generation Finite scalar quantization: Vq-vae made simple, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:17.811559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.005213Z digest=sha256:cf7dcb2ed94955fc739e61d5d793ff280fb4d1cfc3d14f1c86d8d7ab4fe399c7

Observation 2cacb4f6-ecff-402e-98e1-525911416ead · outbound

This paper cites Hotshot-XL, 2023.

RefTok: Reference-Based Tokenization for Video Generation Hotshot-XL, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:17.634533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.100635Z digest=sha256:6792b27916be191ad580ee8309f139f917cc6416de973551749ec4bfd790b17e

Observation 46272d8d-d82e-462f-9b66-2b8f74e0b9a9 · outbound

This paper cites Perazzi, J.

RefTok: Reference-Based Tokenization for Video Generation Perazzi, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:17.490697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.207440Z digest=sha256:6ceac464931874f4a2be5bdbd0b56e8d5c4d6a8c211423baef9cc7e0d3b9a324

Observation 18c400e5-bb8e-44f9-93c7-4feb6f3a7e47 · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing, 2023.

RefTok: Reference-Based Tokenization for Video Generation Fatezero: Fus- ing attentions for zero-shot text-based video editing, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:17.238715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.329799Z digest=sha256:5fb792c2fe7792d20287edde1d4c5a4b3500379fd06c3a9352e3733149e9206e

Observation abcc7be8-7d51-48c1-bd26-ee86b2e7172b · outbound

This paper cites Cosmos tok- enizer: A suite of image and video neural tokenizers, 2024.

RefTok: Reference-Based Tokenization for Video Generation Cosmos tok- enizer: A suite of image and video neural tokenizers, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:17.183544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.405040Z digest=sha256:5fb97d4e97d96221d6236997ad60e03de787815750986bcfd74c3b03f8340b44

Observation 1ce1e44a-336b-4a7d-bffc-d5ce3043461d · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2022.

RefTok: Reference-Based Tokenization for Video Generation High-resolution image syn- thesis with latent diffusion models, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:12.491032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:12.491032Z digest=sha256:c15aa4965a5eb1ff7b09f8b20370e1e13b1bc45ac4b7b55dd7775120b0688eab

Observation dce21332-4963-47e6-9471-9d22e0bb8ca9 · outbound

This paper cites Tempo- ral generative adversarial nets with singular value clipping,.

RefTok: Reference-Based Tokenization for Video Generation Tempo- ral generative adversarial nets with singular value clipping,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:17.065883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.563254Z digest=sha256:330ac2cec1201f7b3831810367f33a7eb3a8fc618336b1197bc99b13a098b833

Observation 4c7a4c6f-68e7-42af-85f4-3364cef47b76 · outbound

This paper cites 9 Make-a-video: Text-to-video generation without text-video data, 2022.

RefTok: Reference-Based Tokenization for Video Generation 9 Make-a-video: Text-to-video generation without text-video data, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.966310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.677468Z digest=sha256:81f513ea39b575ec33b12633e5a76ff6d1883a6b2095204e16267bef2a1c8e29

Observation ad0ac60c-f428-4d9c-b3d3-673545683eaa · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022.

RefTok: Reference-Based Tokenization for Video Generation Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.847244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.770484Z digest=sha256:c3b301c32ccde8f56ff0a2061d9314117c78f8ea7969a744e7eebe5e80dec74a

Observation e7a16f97-85b1-49a8-90f2-b926babf4d60 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

RefTok: Reference-Based Tokenization for Video Generation Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.731515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.864846Z digest=sha256:f68ce02fd965d9ab1e5aa21961bb1446aeb4547834f30479994cd5d49d8b0a57

Observation 707402df-f354-4d8b-9d5f-58091b26c7c0 · outbound

This paper cites Diffusion model-based video editing: A survey, 2024.

RefTok: Reference-Based Tokenization for Video Generation Diffusion model-based video editing: A survey, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.641361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:12.962309Z digest=sha256:c9d381a1f27ac1a199859759cebca114e0c2afad0da34775b50c68834fe712d9

Observation 86103298-8d39-473a-8632-5ae4c19a2248 · outbound

This paper cites Mocogan: Decomposing motion and content for video generation, 2017.

RefTok: Reference-Based Tokenization for Video Generation Mocogan: Decomposing motion and content for video generation, 2017

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.512459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.030915Z digest=sha256:ddcb962aeda33232da51f2637b29256c7d4371614d7ba830ef5ba98514410042

Observation eff2cab0-b20a-4a56-a354-06c9418b3f57 · outbound

This paper cites Neural discrete representation learning,.

RefTok: Reference-Based Tokenization for Video Generation Neural discrete representation learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:13.117421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:13.117421Z digest=sha256:4215359fa612eecc07123ba775f121ef6db5f08ae1a1a3dfc43249f25ead4e29

Observation 7af46399-94c8-48a2-895d-964fb07ffab7 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

RefTok: Reference-Based Tokenization for Video Generation Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.373401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.202432Z digest=sha256:83bfb9ed23d44a67b9a6eac62b492a621c5c08d5b4543446939d00908ef2dd4f

Observation a137870a-f311-4ae1-93f0-64e413b6ae15 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual description, 2022.

RefTok: Reference-Based Tokenization for Video Generation Phenaki: Variable length video generation from open domain textual description, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.251582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.280020Z digest=sha256:912ed7ee15e2022809f7e6cc25b1990527ae0161c52c405892980490d651e3f8

Observation 30f6710d-b73a-4271-b52c-cff11ce17506 · outbound

This paper cites Larp: Tokenizing videos with a learned autoregressive generative prior, 2024.

RefTok: Reference-Based Tokenization for Video Generation Larp: Tokenizing videos with a learned autoregressive generative prior, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.138266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.380424Z digest=sha256:5ccbd84e106a510925cb117c4677216ad4d169a399fff20e0fb6e2809172f20f

Observation d252e1f3-187a-44cb-96b1-a4b544af83cb · outbound

This paper cites Modelscope text-to-video technical report, 2023.

RefTok: Reference-Based Tokenization for Video Generation Modelscope text-to-video technical report, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:16.008344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.470907Z digest=sha256:ee9ce532ab334409da948c6a24a7bea759f01b87c8c42eb5d2533cd84e932705

Observation 9993de7e-b252-49dc-b6ba-b405ba03ec51 · outbound

This paper cites Omnitokenizer: A joint image- video tokenizer for visual generation, 2024.

RefTok: Reference-Based Tokenization for Video Generation Omnitokenizer: A joint image- video tokenizer for visual generation, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:15.882355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.566023Z digest=sha256:92820a382f437da7b66e245628c07b0fd3b5e840a5a4b6a2473d9cfa30beb751

Observation 08039bdb-8277-442c-95a2-5ad17f8446e9 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability, 2023.

RefTok: Reference-Based Tokenization for Video Generation Videocomposer: Compositional video synthesis with motion controllability, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:15.732986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.664212Z digest=sha256:55f16b0aa501fbb5cfbd34315d1baca0b31aec5f50f1f5d2d664fb51f20c5773

Observation 86401808-d66c-42e3-9865-70fd6dddf1c7 · outbound

This paper cites Gmflow: Learning optical flow via global matching, 2022.

RefTok: Reference-Based Tokenization for Video Generation Gmflow: Learning optical flow via global matching, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:15.599143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.732253Z digest=sha256:a0fdb4857fc8c62428e3fff15cf035fd69689e42ec8a31b78b1a73e8b8b6485b

Observation f45e7495-dfba-422e-94da-422213911099 · outbound

This paper cites Videogpt: Video generation using vq-vae and trans- formers, 2021.

RefTok: Reference-Based Tokenization for Video Generation Videogpt: Video generation using vq-vae and trans- formers, 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:15.442988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.828286Z digest=sha256:08dcfbcc71c7624939ed1ac8c426fea2e98bc551d1ff05478ecfa8296303d1af

Observation 7cd3416c-9cd8-4c73-bc30-f9eea3bc6be8 · outbound

This paper cites Elastictok: Adaptive tok- enization for image and video, 2024.

RefTok: Reference-Based Tokenization for Video Generation Elastictok: Adaptive tok- enization for image and video, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:15.294319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:13.928272Z digest=sha256:753358919a180b421121a3f10d39a54d2f82ddb10f347118b52dfa15c8483e23

Observation aeccd932-5da8-4af9-a2c1-474728fbd783 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer, 2024.

RefTok: Reference-Based Tokenization for Video Generation Cogvideox: Text-to-video diffusion models with an expert transformer, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:15.120155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:14.008910Z digest=sha256:2676d454a324f144219429973c6725b8740bf8b3b71b9ce7ece991984f61fafd

Observation b30759f3-ad34-4f3e-80f9-951a66dcbd31 · outbound

This paper cites Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang.

RefTok: Reference-Based Tokenization for Video Generation Hauptmann, Ming- Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:14.946319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:14.085648Z digest=sha256:a99bc96aba21efa5aa816733db8309eef74d4a5bcb8f5586148d4c68419e3fdd

Observation 8f2ebbf9-c330-4846-9e72-af151173160c · outbound

This paper cites Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G.

RefTok: Reference-Based Tokenization for Video Generation Gundavarapu, Luca Ver- sari, Kihyuk Sohn, David Minnen, Yong Cheng, Vigh- nesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:14.803192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:14.155539Z digest=sha256:1de2ff06a2a18b7e2d2b648e039be49b38a4ae4d45c892662c4b1ff13ba71607

Observation 1ccddc64-e647-4587-aa8d-25ef654d341c · outbound

This paper cites Generating videos with dynamics-aware implicit generative adversarial net- works, 2022.

RefTok: Reference-Based Tokenization for Video Generation Generating videos with dynamics-aware implicit generative adversarial net- works, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:14.624074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:14.229205Z digest=sha256:63610dc7c4e73e608992cc02669b2ddc67c0f14ebd014b66c69b7ce559677c8e

Observation 2295b644-e812-46ca-8b0f-35a575a912e8 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation, 2023.

RefTok: Reference-Based Tokenization for Video Generation Show-1: Marrying pixel and latent diffusion models for text-to-video generation, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:14.480920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:24:14.295583Z digest=sha256:34d3ea98e49c8502c55bb9193980a76cd444b4eaba2b30ac47b18c209748bde1

Pith citing papers

Observation 422e5ca2-b01c-4224-9dde-92dbcd19c932 · inbound

A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens cites this paper.

A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens RefTok: Reference-Based Tokenization for Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:10:50.196887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:19:17.796377Z digest=sha256:a28026cd2a69da04b34682e0d363940aaf4001a863c46561fe2e2375621ce582