Pith. sign in

Paper Citation Record · LEDGER

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2503.21979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21979 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:18:31.895518Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2b3bbee4-5ae8-4865-83ec-1108112e45b0 · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.976154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.976154Z digest=sha256:0fde528652f8b8e0b40b3e59ba9c3574ad09283640f339df19633cc0212e66c0

Observation a80dd079-0f4a-4332-859b-5b1d5949f4c9 · inbound

DiSA: Diffusion Step Annealing in Autoregressive Image Generation cites this paper.

DiSA: Diffusion Step Annealing in Autoregressive Image Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.231098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.231098Z digest=sha256:d3d5f8e5fc6ec9d55edb53f12d3a0482f7b96b35013416b763c9e6610724400b

Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.245982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.245982Z digest=sha256:e15c7fbe083d1ded8fa41b32df7ec3f484fd27d88d32e6b54d6b108a1d8ab8fe

Observation 58a09def-c607-40c0-b40f-9ea28861d1b1 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.859294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.859294Z digest=sha256:0df60edcbc09645e397e1a0c03ac7c2855cb183b6ff699ac84c3b367840f3bdc

Observation e75b459d-78b3-443f-a894-9270dc10d264 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.248358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.248358Z digest=sha256:576cb81d1057c7b5e945c2931da0c7a37788e31ce466f704e68af8a7c487fdc8

Observation 9f7a7664-068b-480c-a038-23757b08a911 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.481610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:3beb1008c2e0c08e5adaeffa85d4cf560cd1c2133d1c3569aff3cb42eabb460a

Observation 1149a9a4-a285-4b5d-b66c-f5aceed07d9c · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:640ace5f5785900cc207de8af17251fc43d2caf4d3a90b15b6fdf2b3aacc9f62

Observation d00a1655-ebce-4abc-ba91-ce65d35f17d2 · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.081884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:35c8725038e8418893f02f48b1d9f6874ed4dad60fd16c37306d3ae307bee34d

Observation f4871b2c-0371-4777-8a3d-eedb6c0ba895 · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.277057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:87c1b2fb304e40e4a72e2d78c6a9f3573017a2fd5d872eb5e6db29bc87853dfc

Observation 02848f04-40c5-4cc7-be1e-6f6c2432fe1b · inbound

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens cites this paper.

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:24.458005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:58:03.974053Z digest=sha256:9b90d5ed7d909c2ca672cfd90c32093fbcbdf98ac934aa13b50a92e50c7e2e1e

Observation 3af64175-1c8e-4cb9-9727-eba43ed6d6fd · inbound

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection cites this paper.

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:15.287320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T22:21:10.133655Z digest=sha256:b27c9dcb50b4238f5b71f1e14759571ae059b75e781f260a2eb7c8fa24af5537

Observation b6af3b54-1388-4e76-b3c0-4b9ffa81aff5 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.714030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:17c3a5ebebb96166fe9155536527817e004af032017d4d8f46e3ec3896fdb75f

Observation 40a9a53b-5dbf-41d4-bd34-b9cde05af524 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.793520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:e272dd625180330fd6ab8edd9b98c00a7e3ba92cfde3ade1aa64b34cf0556122

Observation 879ca5e2-fb28-429b-9c95-a44b0ce6226d · inbound

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models cites this paper.

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.437930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T05:35:46.236860Z digest=sha256:4d43ab3c07738ec4d0124d8f012b6585ba88331db8fae4aa979027609755dc11

Observation f5913998-4216-4ce2-8925-878a650dcab1 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.881233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:7c83193223e0b993b91d4a3b5356b99f1a4831452c56ad9d8208b0ac2f26d942

Observation f80ece7d-474a-4a7f-849e-0cf729c50f3b · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.677719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:bf1a0170f84a1f741ea19bfe023875fe7a74557a1cd7987aa9c4c52ee7ddf918

Observation 48829565-1a03-459b-a7f2-600152fa830a · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.477984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:5550861063e944d345253a0b79edd03e17c7025581091f04d672d7fad2008099

Observation 02f63038-6940-4cd3-af38-483d52477528 · inbound

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation cites this paper.

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.559037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T00:50:22.839005Z digest=sha256:0404a1fec27d249437d198bde1b67f265a949a5b8f2cd135ba0164b9f95f4404

Observation 4784d571-1a55-4630-99cc-5d6dea84cfba · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:03.046647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:03.046647Z digest=sha256:47171a370d59332e1c9c84b80e13f70846d1a6b3b461e69dceb549fe5d110195

Observation e794fb98-3746-4095-a919-f93c121b53c6 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.791253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.791253Z digest=sha256:57296b1fa593adb640af1b863124fef02a18407bd81bb75b6b0f98fedc228c96

Observation 40aeaab2-5f0f-4ce1-8d4a-ac5288f0a93c · inbound

XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling cites this paper.

XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-16T00:18:31.895518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:18:31.895518Z digest=sha256:2d0272228b45e1e63474b731ddd4a9b5f50a41b0621cb2e6516672635bb43414