Pith. sign in

Paper Citation Record · LEDGER

UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.04423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04423 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:12.950240Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:15.732563Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be42375d-140f-439b-80ec-a9391c13516a · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:12.950240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:12.950240Z digest=sha256:c94f734527e9812ed31048b9fcb8590684049b7d265ace8a8ef5f2a9952eed15

Observation 31234611-dab3-4a1d-a3d0-a54352ddb84b · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601708Z digest=sha256:24f4b66097c3db3a12206218c87df87935d74e83bbca26128b2666890622f55a

Observation 4fb701a8-a7fa-4fd5-adfd-b33b36b63ab6 · inbound

R-Genie: Reasoning-Guided Generative Image Editing cites this paper.

R-Genie: Reasoning-Guided Generative Image Editing UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:35.925099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:35.925099Z digest=sha256:904ec707eb21f9035a2f2f285caea2f620538e91566b8c4fc884e9b8e4608d55

Observation 1ed1e2e4-6fd9-447c-89d2-e5d3ad74bb88 · inbound

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks cites this paper.

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:32.219643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:32.219643Z digest=sha256:1e5892f1f21e2c9374befba31abce09fa5bfa4403acde62c67d7867799f31fc2

Observation b9e78131-bb3b-4edd-94c8-b901daf3346d · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.736533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:494936e3e2ee1e8b9b76b859364ba8110ed0e80061a5e649941f48f180cb93f4

Observation 185a68ec-e17a-4264-a362-0c073d654f6d · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.667271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.667271Z digest=sha256:9047993d6013902c6f204dd9e9008e3cc4f24c9dbff4d3bac1149f42e9665976

Observation 1b1501ce-e4d0-4e53-970f-12d20944505e · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.932896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.932896Z digest=sha256:75501e2fbee4c4c312d4fcb83c8c6757df4e04039905cdbe47871773b303b4ef

Observation 5d1e0680-aea5-4e43-802b-bef60542ce56 · inbound

2D Gaussian Splatting with Semantic Alignment for Image Inpainting cites this paper.

2D Gaussian Splatting with Semantic Alignment for Image Inpainting UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:06:36.052885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:06:36.052885Z digest=sha256:a432160c651b2d2bbf46aaca2072c6c55d89d3276bc29015189ffafb30826888