Pith. sign in

Paper Citation Record · LEDGER

An Image is Worth 32 Tokens for Reconstruction and Generation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.07550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07550 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:23.840987Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:13:49.010757Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6567b4a-4243-43cb-a42c-2000efc017e8 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:46.153147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:10c16c1c070c93f5e8da71b12898df4003a41a8901edfcb387bc2da71ba73461

Observation b2baccb6-563b-4823-bd9b-d3ce09306f82 · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:21:45.136776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:642b7857ef700e5cdce900277ed7d58319f492543f9ea22b616f67fcf5812a96

Observation 1406d1ef-0d27-4c18-a42b-e0873baf7cbf · inbound

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation cites this paper.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.840987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.840987Z digest=sha256:63ae3f5503b96d7e261f222079511b699501ceafb3f9ab3492b133dfd2419dbb

Observation 97c6b236-800c-42a5-8c4d-0f0cb0fda1c2 · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:04.678813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:04.678813Z digest=sha256:d336a6b178c2a0a89df8bdf035da7085fc9b847ea18d00f4fd7203fb94e96586

Observation e40974e0-86b1-4647-ab9f-ab8079e07911 · inbound

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? cites this paper.

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:06.140146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:06.140146Z digest=sha256:f7e7afe381f92974ea7281ab2b56fc3bd466fec41c719c34237a3ff6f4c092ec

Observation 4b2fd8ff-3c60-4015-bbfc-aa3e9bcfd79f · inbound

Hita: Holistic Tokenizer for Autoregressive Image Generation cites this paper.

Hita: Holistic Tokenizer for Autoregressive Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:58.541916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:38:58.541916Z digest=sha256:ea7feb032a5714e883b786abf512a499f65f5521edfb9faf06fa4d59e5ad41e4

Observation 5e6fd36e-aeb7-483f-8b76-514cd63103d4 · inbound

Single-pass Adaptive Image Tokenization for Minimum Program Search cites this paper.

Single-pass Adaptive Image Tokenization for Minimum Program Search An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:32.847351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:35:32.847351Z digest=sha256:88edc1def3a7dbfcbbfac746d624c7c4b23d3b03fbe143166ef8edf5ce75cad2

Observation e043c121-f9b0-4b56-942b-a2b6b4a95145 · inbound

CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio cites this paper.

CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:41:39.636626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:41:39.636626Z digest=sha256:8b7455e1cd5d8e012d2d75fe7e10d1661124be1ab1910a6f3b21d4e9be9c60e3

Observation 282f7c6d-80b6-4c6c-9a70-54a0225dd3b0 · inbound

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization cites this paper.

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T18:12:48.006661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:12:48.006661Z digest=sha256:f6d786cf80f7e3a8dc860e2942fd0092292abe3b260ad6ecbd09c385880bac33

Observation a0028c5e-e42b-4cdf-ab2c-7b66998eca75 · inbound

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters cites this paper.

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:17.503383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:06:21.438738Z digest=sha256:d35fc2d937aea134d6c6ea906af3b65f2eb61fcacc16ccd5e660204009b30357

Observation 8efbf0fa-5750-4b9e-b108-0f04cf73663c · inbound

Does Engram Do Memory Retrieval in Autoregressive Image Generation? cites this paper.

Does Engram Do Memory Retrieval in Autoregressive Image Generation? An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:28.137756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:33:17.978924Z digest=sha256:7bc9cd1a089cffb21ac207d701d4a8ee1e33ea6b770a6ebf5fff359730dfce2f

Observation 4c9f453f-ea96-4bc9-a3d5-e1797b412881 · inbound

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice cites this paper.

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.039023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T22:41:44.510546Z digest=sha256:558cad0f1b8574bf4908a74f81d6551f2236ef2aa4a213ed27c90f8155433f86

Observation df817617-febd-48ed-9a38-f9ca51fe0792 · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:03:13.487184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:183189db10ab0a4bed3fc4e2975f29bb4da55043ebeb612b09551ded33b0182f

Observation abde206c-91da-495f-980a-3aa274828bc2 · inbound

Structure over Pixels: Learning Variable-Length Visual Programs cites this paper.

Structure over Pixels: Learning Variable-Length Visual Programs An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:13:49.012240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:06:01.684713Z digest=sha256:062ea19eebfdf62f591fe50e2c3534f2c407771b1439d9e0ab82991570b4346e