Pith. sign in

Paper Citation Record · LEDGER

Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2501.09755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09755 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:54.714314Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T15:43:53.781661Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 09dadb60-9087-4483-9174-0060d7155eec · inbound

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization cites this paper.

MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:54.714314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:54.714314Z digest=sha256:7c8788d5a36ccc62c67aa27caf046a92e54e1e424b7ffea2df9fcbad8d66acc5

Observation a76ca4bf-c885-4701-ab4f-dcb157ef389b · inbound

Flow marching for a generative PDE foundation model cites this paper.

Flow marching for a generative PDE foundation model Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:51:25.457933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T13:48:14.532529Z digest=sha256:6f4c51779868a7f86f551f5bc83448b946d29fb31adce418e48bc5f732cb3fc3

Observation 2aa396cf-8da4-4bbc-b54a-8e497bd7139a · inbound

TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders cites this paper.

TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:57.138473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:44:14.636654Z digest=sha256:8e8e0b17a078ea046d713011de1711d17b3e2cde1a34c401fc57516f27048703

Observation ac0c1624-78af-40ce-9719-35b212adf4b2 · inbound

Latent-Compressed Variational Autoencoder for Video Diffusion Models cites this paper.

Latent-Compressed Variational Autoencoder for Video Diffusion Models Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:03.077326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:23:19.583713Z digest=sha256:b8208dab69ce56990fb0efa7ee250c862d14fc25cf580368fc242ee27e4704d4

Observation b0ace6f3-d26d-4699-b281-320374b491fa · inbound

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters cites this paper.

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:51:06.445578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:06:21.438738Z digest=sha256:9dea2a6eafb3a938c8b5ce5a6f98b1448afcb6845208908bcb7a15d6d48c8682

Observation d0055b1f-9d88-459b-bc31-c64cb3afe703 · inbound

LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation cites this paper.

LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:30.898606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T03:42:56.202960Z digest=sha256:135391dbdc922e457b02cdafd0a99a54919fba93b5a94826d12bf84189fd800b

Observation 7e29255c-bf86-4f97-9053-82a549e05319 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:57.333354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:c7b23303d3001cf72253afb4dcda04e1673b1a974cd712612d15e7570d42c74e

Observation 3dbd517c-af80-47d7-b9be-5335072a6892 · inbound

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation cites this paper.

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:29.699689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:01:05.492969Z digest=sha256:3886867749c0930d6f97aa5a628e59d71563b630efb2e27526baf91aa12d1485

Observation a0bbab4b-173f-4e72-9dca-99de3929f7dc · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.439682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:5b92bbc0f1ff923edbb524ad6cba54a92a0a77db1c8e954bad86e6bcfc48becf

Observation 51798922-a3bf-4a4f-9b45-9dd38f20c13a · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.456464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:30fe37f03f67dbc23785f4b728b5ed2d1b38e1e972700dceae17b51d7ef04e13

Observation c02e9aa8-1cbf-473b-a748-2a9cb2180d83 · inbound

Multiplayer Interactive World Models with Representation Autoencoders cites this paper.

Multiplayer Interactive World Models with Representation Autoencoders Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T15:43:53.783116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T15:38:36.897705Z digest=sha256:f73c71294271a7e48628f9cbaedc14c57397aa7ffbc0b78c21ded6b5498000a2