Pith. sign in

Paper Citation Record · LEDGER

VidTok: A Versatile and Open-Source Video Tokenizer

As of 15 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 9 inbound Pith citation observations for arXiv:2412.13061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13061 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:30:23.560193Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:20:14.972158Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:28:55.542980Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77a296c0-a3c4-4f0f-877a-366b7ade12d6 · outbound

This paper cites UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing.

VidTok: A Versatile and Open-Source Video Tokenizer UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.456291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.456291Z digest=sha256:ac29c8e9e64dcd2eef61216253c79b6e44f4e9957ba0e9491b1fd393ff632da1

Observation 5af6f8cb-127d-4553-9916-701a0d77dd2c · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

VidTok: A Versatile and Open-Source Video Tokenizer OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.476606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.476606Z digest=sha256:cf4b7507c404053aeb43d7085b708f41e4ccdeddb2f5f7ee444382c9a61d4fac

Observation d78e62f2-88b6-4212-affb-e5460ddfd3e7 · outbound

This paper cites Layer Normalization.

VidTok: A Versatile and Open-Source Video Tokenizer Layer Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.496120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.496120Z digest=sha256:43080efa1275b0995244862ef38a92f54e1c4a90bd79c92b5d1517b81f1dfc5c

Observation e239a439-bd92-4ea4-bb36-fb3388c6a8a2 · outbound

This paper cites OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation.

VidTok: A Versatile and Open-Source Video Tokenizer OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.537169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.537169Z digest=sha256:b860df712541d49744d030cd098876fc3556e951e6eb3be85fddd57cfe811c53

Observation 97ab6a6c-c23e-4a05-83f9-00e8e492a65d · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

VidTok: A Versatile and Open-Source Video Tokenizer VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.548962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.548962Z digest=sha256:ea2eb089fa53435c423d76c2c0c85452540728ba2c0e44b28e2a93b5cb5ed93e

Observation 19f470f6-33c2-4dc8-a6fd-297f4b00601b · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

VidTok: A Versatile and Open-Source Video Tokenizer Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.543239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.543239Z digest=sha256:df42ced3d479b9fbc4a6dc9f381cea7b23dbc9b5cd6376a86e8ec792012aa60a

Observation b007f36e-c663-488b-9de0-2b7091447943 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

VidTok: A Versatile and Open-Source Video Tokenizer Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.521476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.521476Z digest=sha256:8ce96fcff37395e22b29383bac30d2c316d0ffbe1c2113cf991173dee82a74d9

Observation 93ccae42-6308-4ddb-b643-350e976d7735 · outbound

This paper cites Kingma and Max Welling.

VidTok: A Versatile and Open-Source Video Tokenizer Kingma and Max Welling

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:23.900088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T13:30:23.489151Z digest=sha256:d55cd1324f8fadb858c650e3a81b311a1cb266343edc132c221f13f39e5bb8ff

Observation f4b6b843-8591-45de-9fc5-9e874022bcf7 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

VidTok: A Versatile and Open-Source Video Tokenizer VideoChat: Chat-Centric Video Understanding

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.503825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.503825Z digest=sha256:0bf985533415478cf9f9afcf6e84551920084c924f94200fd19be4db4db0f3d8

Observation 15c788fe-4e63-44af-8e05-70643bec5acb · outbound

This paper cites Mcl-jcv: a jnd-based h.

VidTok: A Versatile and Open-Source Video Tokenizer Mcl-jcv: a jnd-based h

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:23.879226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T13:30:23.529604Z digest=sha256:ca187695c2c5ea651dd4e80ef4e09619f4f2a6c6ed78ea2760ebe50f62df65a7

Observation 26e6b10e-eba2-4e9f-b75a-467e885db8f3 · outbound

This paper cites Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators.

VidTok: A Versatile and Open-Source Video Tokenizer Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.560193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.560193Z digest=sha256:86195392dd657b42c3927023c2a27579e7243884b96d0ab25d6ca6df3cef58ee

Observation 29f3b717-890d-4ee9-83b6-8e1dc5890046 · outbound

This paper cites aMUSEd: An Open MUSE Reproduction.

VidTok: A Versatile and Open-Source Video Tokenizer aMUSEd: An Open MUSE Reproduction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.511895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.511895Z digest=sha256:3bbaa51940ce9c98c7860ce14801fbc6ee2aeecabd82e197ea3022f4b9fae83f

Observation 8b702859-d4c0-4d7c-a5eb-d9652f86e853 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VidTok: A Versatile and Open-Source Video Tokenizer Imagen Video: High Definition Video Generation with Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.482600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.482600Z digest=sha256:3912f2b0162db1917c8a1856aa6bdfd72316bea3d38a44e938d05ed57ec29233

Observation cea75237-4c0a-42a1-85e9-6cc0e579c6f5 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VidTok: A Versatile and Open-Source Video Tokenizer CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.554765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.554765Z digest=sha256:39a78373bc963db9f9eeb05d7a5035b76f8bc7f541d7c892e6b25b2f7665cb77

Observation c8dc4f06-7ca2-4600-a146-0675eee0062f · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

VidTok: A Versatile and Open-Source Video Tokenizer Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.470488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.470488Z digest=sha256:5718dcff7822925bc3e6b44a23d508ee89e932a53095799aa645fa11ce9c6dda

Observation 7a67941b-4140-4d96-8879-d984c906f2eb · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VidTok: A Versatile and Open-Source Video Tokenizer Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:23.464444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:23.464444Z digest=sha256:36e904404b2a9ece5675a66ff412514f1fe7fb692ddb176862bc768d28fe3b29

Pith citing papers

Observation 85fa5572-3bdd-41d3-ba03-3cfb56585fa4 · inbound

VidTwin: Video VAE with Decoupled Structure and Dynamics cites this paper.

VidTwin: Video VAE with Decoupled Structure and Dynamics VidTok: A Versatile and Open-Source Video Tokenizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:14.972158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:14.972158Z digest=sha256:0d67ed064952197072d01f1009f3ac6010edde4c625c0acb416fa6bc5d138bda

Observation 4427e8f9-248f-42a0-b985-b4244c81698b · inbound

Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces cites this paper.

Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces VidTok: A Versatile and Open-Source Video Tokenizer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:16.114855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:16.114855Z digest=sha256:3594a0a3241deb2ca2033b8ef776837641311cb8b71746f4d715467f5ffc412c

Observation bb307e0d-63c7-494c-a2ec-9efa73f29d14 · inbound

Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis cites this paper.

Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis VidTok: A Versatile and Open-Source Video Tokenizer

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:02:16.754970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T12:00:37.025335Z digest=sha256:250d274f469ebcf7627eeff49e520e1532935bf5351d03032172f744e5743f2b

Observation ef69d201-f105-4960-bf9c-554ac4c7e766 · inbound

Playing with Transformer at 30+ FPS via Next-Frame Diffusion cites this paper.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion VidTok: A Versatile and Open-Source Video Tokenizer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:45.159313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:45.159313Z digest=sha256:f7755bf0bd82ec6a49c8c4fe3e7f9698a1b5bd8a85b0ab41aac03d1ed1fe4643

Observation 33048bd2-577b-4e5a-97b2-1f0dc8a1447e · inbound

Video World Models with Long-term Spatial Memory cites this paper.

Video World Models with Long-term Spatial Memory VidTok: A Versatile and Open-Source Video Tokenizer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:40.611770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:40.611770Z digest=sha256:55c1cc0127780a51fb7bd6d01124b0f253af3cbdf10837c8926e0f81524eb097

Observation 6e61c7d4-8052-44c3-8a74-56aecbe1d57e · inbound

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics cites this paper.

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics VidTok: A Versatile and Open-Source Video Tokenizer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:49:52.048245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:49:52.048245Z digest=sha256:a97537aa0953b39bc2af97e5c3f05e86a99bbcedd0d0d757a5f539dd8dabac9f

Observation 05f6dd02-a832-45f5-9bea-877cd866512f · inbound

TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization cites this paper.

TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization VidTok: A Versatile and Open-Source Video Tokenizer

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:28:55.544391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T01:20:32.508409Z digest=sha256:e565706d952148335ffaf52c05bf1774592634955c0510cf77bcbe08e595a5e6

Observation c365d70a-3fee-412a-8dc8-0e461a32b98b · inbound

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation cites this paper.

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation VidTok: A Versatile and Open-Source Video Tokenizer

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.991649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T06:16:13.420481Z digest=sha256:62283154a016e19485cb82240f3c5b9148f7078bf72925fc1e55ef87ed4d40d8

Observation 817a5b73-23f6-472f-8b79-7989bc4ca8c3 · inbound

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers cites this paper.

FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers VidTok: A Versatile and Open-Source Video Tokenizer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:50:14.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:50:14.926631Z digest=sha256:decfb1bd3130dac4eb33564a30c0f1345a3005fe196caa7d49ea7f8f6ce321ab