Pith. sign in

Paper Citation Record · LEDGER

UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2504.04423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04423 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:49.667271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:15.732563Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b9e78131-bb3b-4edd-94c8-b901daf3346d · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.736533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:2c7bf21557d5d9d5381aecbb1e6d7945b1a72569fa8e31b4a3b229f84ad6f30e

Observation 185a68ec-e17a-4264-a362-0c073d654f6d · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.667271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.667271Z digest=sha256:59c55343f3d7bed1bfe329124e8e028eb6dee49b32d95be0d54222223e825247

Observation 1b1501ce-e4d0-4e53-970f-12d20944505e · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.932896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.932896Z digest=sha256:e77858fe0db5b820278294e84c31a68512279f746a925933c1a8404137c31846

Observation 5d1e0680-aea5-4e43-802b-bef60542ce56 · inbound

2D Gaussian Splatting with Semantic Alignment for Image Inpainting cites this paper.

2D Gaussian Splatting with Semantic Alignment for Image Inpainting UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:06:36.052885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:06:36.052885Z digest=sha256:f0cca005118c17b27ec3058dee3ef407a9292e7e7c5b4d5cdfb6c279414188a5