Pith. sign in

Paper Citation Record · LEDGER

Image and Video Tokenization with Binary Spherical Quantization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.07548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07548 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:04:02.451468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.369173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a329c8a-3a3e-4058-ae43-f914da5e935c · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Image and Video Tokenization with Binary Spherical Quantization

Reference 248

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:46.176065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:13ad0d3bde159b89d8ae3725ee4747cde1651df1dbbd0a5f5ff96438b75ddeb5

Observation 0d66fc1e-a3a2-4698-87f2-d800bc7c8802 · inbound

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation cites this paper.

QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T20:04:02.451468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:04:02.451468Z digest=sha256:b54e0a58bacfafad66f00aad99c0d44390833eed77a877c1f135a3b1f68bf149

Observation 53a839f8-3586-4ffc-af44-54354bb9b2c8 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.913350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.913350Z digest=sha256:9aac1a0cebbf5d8bae5c66287ec4d5198f983e74e047bdb6a64d7f0c46316995

Observation 042198ce-a5bd-4b17-8d00-45b9a34297a5 · inbound

TokBench: Evaluating Your Visual Tokenizer before Visual Generation cites this paper.

TokBench: Evaluating Your Visual Tokenizer before Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:44.839423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:44.839423Z digest=sha256:7ec189cb32ba1500a84692d4140cbcfab173f5813763a3e68c3039fdc2e5879d

Observation 38e3a0e4-6030-48eb-9e23-e98ba94771b3 · inbound

multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data cites this paper.

multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data Image and Video Tokenization with Binary Spherical Quantization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:09.611149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:32:09.611149Z digest=sha256:45a5e5df081911dbee3241cfa746c9a4194e8202a14b645603ec22fff9a9963a

Observation 82af5880-746c-42df-aed7-16c7075b59c0 · inbound

PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models cites this paper.

PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models Image and Video Tokenization with Binary Spherical Quantization

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:27:18.938661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:26:17.991938Z digest=sha256:6970165828de074c97424d3bb47b976e8efeb76d2e06c0343eb78193870c9e85

Observation 0eab98de-a119-4e18-92bd-603a1c5d7329 · inbound

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization cites this paper.

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization Image and Video Tokenization with Binary Spherical Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:52:28.135246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:52:28.135246Z digest=sha256:28c6b524d14d4735a72f518e83375b1b22bdd494cec7119777df777569400e88

Observation 6be68bb3-d653-4d4c-832b-eaf9de91d993 · inbound

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive cites this paper.

Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive Image and Video Tokenization with Binary Spherical Quantization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:44.578141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:44.578141Z digest=sha256:b20571a366c91fc9206bdf60a22b566982d444c5248f54a8ae06d456b90fc697

Observation e0c3e197-b638-4c99-8321-2b6944591867 · inbound

WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting cites this paper.

WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting Image and Video Tokenization with Binary Spherical Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:57:19.687253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:57:19.687253Z digest=sha256:ac57fc14297f7085e8de895e3cbc3cdd5bd19f5d0d0d1f0fbfe442e089692431

Observation 4db5a033-e809-4e52-969a-df2aef4bdaf1 · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Image and Video Tokenization with Binary Spherical Quantization

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.263008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:a989c7bb1f79d5a65013b9cfb16978527f433910bac040fa2b77c4d93d3a6a28

Observation 53889dda-c3f6-4c01-a36d-a41535389e37 · inbound

Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications cites this paper.

Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications Image and Video Tokenization with Binary Spherical Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:56:42.310530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:56:42.310530Z digest=sha256:3a67e354ae3be6c7cbf942706a0835acb297341029892f3caa54d3ca1aed9e4d

Observation 107f7c47-d0e9-4b38-915b-5cb972d24187 · inbound

LGQ: Learnable Geometric Quantization for Image Tokenization cites this paper.

LGQ: Learnable Geometric Quantization for Image Tokenization Image and Video Tokenization with Binary Spherical Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T22:43:28.615040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:43:28.615040Z digest=sha256:62b09df52e6c52a3652e190ea2ce3333db7c83333a79793730b1b7a31875c8b3

Observation 0740cf01-3056-475e-8ac8-d875c6eb6c39 · inbound

Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI cites this paper.

Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI Image and Video Tokenization with Binary Spherical Quantization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:49.442294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:43:41.036191Z digest=sha256:311e70556eb96dc582f8e6bbe23f9ce0455f5ddbf4062aecc47af27b743edf03

Observation d500312b-a983-406d-8833-9560ec6231cd · inbound

ELT: Elastic Looped Transformers for Visual Generation cites this paper.

ELT: Elastic Looped Transformers for Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:06:00.090311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:19:22.543462Z digest=sha256:88f983d428572ceac92f87b3515cab5998ef501e4c27b68cbecc11e97d3569a2

Observation b830cbcd-5aa4-4060-93e3-146c61d73648 · inbound

ELT: Elastic Looped Transformers for Visual Generation cites this paper.

ELT: Elastic Looped Transformers for Visual Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T16:35:03.432151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:35:03.432151Z digest=sha256:7a6fa14f7aacc594e94f038440ed0540b8fb0ec2514d800b40a99006098d6d3f

Observation aff94852-2587-417d-bcd5-818e8d7b99e0 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Image and Video Tokenization with Binary Spherical Quantization

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:03.361529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:41:23.057949Z digest=sha256:6c3f9f354a6f21e1b68db9b6dce7042a8ba3da5ef1dd7d9493d27d48e51b62c1

Observation 917f8982-2f24-4a38-96a7-f7a155f52e53 · inbound

Generative Refinement Networks for Visual Synthesis cites this paper.

Generative Refinement Networks for Visual Synthesis Image and Video Tokenization with Binary Spherical Quantization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T21:00:35.123495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:00:35.123495Z digest=sha256:2474a9b0ba0b5f2036e90b25620435037476345142262bfed6400146ec4d758f

Observation 2eb6b7d5-2334-4925-adce-1565de30279d · inbound

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning cites this paper.

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning Image and Video Tokenization with Binary Spherical Quantization

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:08.120351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:05:21.069489Z digest=sha256:cc757c4ebc07ce548e3f896139425544f5a9bfd3d573c2744e0ffa3ecd218b7f

Observation 82c4b2ac-23be-4cbd-bfb7-afe1354ffeb3 · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion Image and Video Tokenization with Binary Spherical Quantization

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:57.135970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:59c18c0b7167a73a9df9abe691f48091df29329468fd4d10fedf2cf78f47dfe9

Observation f670e45b-f998-4c30-8c9d-d7f4fa84749e · inbound

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation cites this paper.

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T06:27:24.788773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:24:01.820975Z digest=sha256:d4d49fcbc2ac108d5a0265d772fcf85145cb83b27f1e4b26a04024890832c272

Observation 4d54be22-07df-45b3-85a1-13f8abf699a5 · inbound

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation cites this paper.

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-08-02T14:19:04.486699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:04.486699Z digest=sha256:fe869e5477e0d17a0f90a9a27b538740d3812c86d85d506f60949fd76a4735b0

Observation b3587e9b-bcc7-417b-b6d6-618fc9e189b8 · inbound

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models cites this paper.

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models Image and Video Tokenization with Binary Spherical Quantization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.307706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:01:01.427273Z digest=sha256:d9c3200a99b160748ed2d139acfd496940e55b4b63a10a6a2e304d3e50613f9f

Observation 5195f449-27e6-47ae-aa2a-7076b077e4c8 · inbound

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation cites this paper.

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.008832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:00:16.005187Z digest=sha256:3569597086b61af27f0ef3c00a1465404a3cf082717c2d53c0618c9afa988b5b

Observation fc5e4f06-5d5c-43dd-bbfb-e88e0fa74da9 · inbound

ChannelTok: Efficient Flexible-Length Vision Tokenization cites this paper.

ChannelTok: Efficient Flexible-Length Vision Tokenization Image and Video Tokenization with Binary Spherical Quantization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.334531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:09:25.049534Z digest=sha256:495e6063f6cb3c15985f40a443277df4e2a8d0bc97af852c9ef0c830d84e4aa6

Observation 7d5c36a2-c73d-4a12-ac10-fa0a315ef9de · inbound

Concept Removal for Frontier Image Generative Models cites this paper.

Concept Removal for Frontier Image Generative Models Image and Video Tokenization with Binary Spherical Quantization

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.370684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:18:32.620951Z digest=sha256:8491e55b0e5ad43261f40b3f8ed7c4113b00a8ead406bdb58b07e5aec3551e59

Observation 67e9d513-4836-49da-98ce-eb506e862f5c · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Image and Video Tokenization with Binary Spherical Quantization

Reference 204

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.152440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:5b725ba57b95886c90b5a3cd3ec25c23c07749b16e52b24dd0afed142c1a3ef0

Observation 388dd666-a4a8-4db3-81b3-187d5cb46766 · inbound

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units cites this paper.

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units Image and Video Tokenization with Binary Spherical Quantization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T00:58:38.644366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:58:38.644366Z digest=sha256:617bb0d97cca85a42824b7b5cc50bd2a3b7af21a6f0771fbcf0b2d750e326450