Pith. sign in

Paper Citation Record · LEDGER

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2412.03069.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03069 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:29:16.883675Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5b9e76d7-d684-4499-8d69-a1284bc0d394 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.158653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:e6e5aa1857e200ca90d658adcce84fe4a1abf4f1c3a8df0ff948ae63119c22da

Observation f55f4173-db02-4d93-a5e4-96030d315cea · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.721256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:d6cce708495c25c3a320d63ec1432df498c48ba60d8ad9ebea4be849f2a51080

Observation a9e53c24-575a-4fc3-b371-025ed7011778 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.603862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:f442de87dfcd79843f4085cdff527a7935c5ec997ada3dfe8bce3c6c0a0a9d54

Observation 6869a934-ada0-4a3d-a060-0bf823273652 · inbound

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset cites this paper.

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:34:27.042984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T23:34:26.878354Z digest=sha256:5106292d0d29b4f3c403df51f1e915847c9c11b943bce9ed528f5177dba41adc

Observation 98fed8a3-a44b-49fa-8069-5a61c92b004b · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.211623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:605d58ababf592f1aa88c00ce53b61c03f45f0686bd45d9902936218bbfddec7

Observation fb7f0dab-7b70-4b83-b7d9-fbfcbae00a30 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.542475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:423017a24a729f9326539d31dc3bd70289d62211732d72dc33bac91a42bfb118

Observation b88afb4d-7579-4fad-ae74-0661592f29e0 · inbound

A Unified and Controllable Framework for Layered Image Generation with Visual Effects cites this paper.

A Unified and Controllable Framework for Layered Image Generation with Visual Effects TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:57:50.252224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T11:54:46.989748Z digest=sha256:894a6b7c1d0b0dbe39da5d661d4da762b6d8692eb7501bb3d0eca15161640318

Observation 609a9b36-99f1-4287-a95c-8be0c11ae5cb · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:52:59.470222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:359d4ed59c239a7106dd46367a3c9fd58345139eb57f5780b0d92ba95232c651

Observation f7bec7b5-3764-4502-9b7e-e1511547bb45 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:d65c61591316df2d63e9875ab944251a49736ef31f2e081928224c16e4070639

Observation 211c994b-f3f4-45f2-a9ff-72e6f011aac3 · inbound

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning cites this paper.

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:10:53.714152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:16:58.323955Z digest=sha256:1335bfffa13fd0bf370f938e95ea8a75cbc58ebf95d12e7a7dd8d76ddb7572b2

Observation 85b178fd-4ad0-47f0-974e-080507de5924 · inbound

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding cites this paper.

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:15:28.979944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:15:02.723774Z digest=sha256:5d00a941ddea8cfd2c888208b9a1d1ebc9a24495c4adb9574bc5156cb2fb72cc

Observation fec4d6af-7265-4a0e-a833-957939712b01 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.367534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:c141a4017760c61801217ce67745aaac31c32e0761f5fbe591514d22b7e3680d

Observation 226722f2-0983-4fef-883d-c2374909187d · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:29.041921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:c194c64132418e640bdce9b66930d56425b107391eab11b9e6e47c406910899b

Observation b7e1ded3-b160-45dd-bb73-6d9cfdeb56cc · inbound

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion cites this paper.

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:05:57.477353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:57:24.033068Z digest=sha256:8fad9dc5ae08e27c9e14323ebb917e1484bbd0972d50aca730db8af7dabba9d0

Observation ab1bc623-546d-41bc-9581-ddde19cc7ab6 · inbound

Vision Foundation Models as Generalist Tokenizers for Image Generation cites this paper.

Vision Foundation Models as Generalist Tokenizers for Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.455436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:01:24.738195Z digest=sha256:4c4e953921b7d97ee5cc5cb2542230f2669d4cbc5946f7ddd5de7e471b4d503b

Observation 179e42ab-e52d-411d-96ca-81da6fff4ca9 · inbound

Histogram-constrained Image Generation cites this paper.

Histogram-constrained Image Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:41.186922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T05:51:19.756153Z digest=sha256:8b01a6ea7db3f6fa280dcbe9a898eff3e788e8c10f9fc168b124b4a144f6cc60

Observation c07c4d84-88bb-4bb0-bf4c-91baa3e94539 · inbound

OSVE: One Step Video Editing with One Step Diffusion Models cites this paper.

OSVE: One Step Video Editing with One Step Diffusion Models TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:29:16.883675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:29:16.883675Z digest=sha256:48d7774e9f44ae77fa91c6eec620f5479ae1df3e4e30bb40c25a2878dee25752