Pith. sign in

Paper Citation Record · LEDGER

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2503.21979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21979 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:04.231098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a80dd079-0f4a-4332-859b-5b1d5949f4c9 · inbound

DiSA: Diffusion Step Annealing in Autoregressive Image Generation cites this paper.

DiSA: Diffusion Step Annealing in Autoregressive Image Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.231098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.231098Z digest=sha256:0568aaa9b9a6cd51d527f0442b3e6b103d375288049e019930eb67ffd90dd608

Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.245982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.245982Z digest=sha256:eeb912082937ec2136bac193624b7af452152192b354fea366461fd341197b43

Observation 58a09def-c607-40c0-b40f-9ea28861d1b1 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.859294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.859294Z digest=sha256:296288d7df932c88017a5af6dff8502ba0283e53f1e3b6ad58af242bf9bfe674

Observation e75b459d-78b3-443f-a894-9270dc10d264 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.248358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.248358Z digest=sha256:e62171a0b3fbb8546163eb41a2dffe00f408db25ffaaa6de855de62b5f9e3ce8

Observation 9f7a7664-068b-480c-a038-23757b08a911 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.481610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:96b734eee947285dbadde14f0bcbec4e3548e76325dd96f7ab96a3ba02360756

Observation 1149a9a4-a285-4b5d-b66c-f5aceed07d9c · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:9d79a763c072efb83c0874842f20b8818fc909a566c67b8a71deb97202c178b4

Observation d00a1655-ebce-4abc-ba91-ce65d35f17d2 · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.081884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:cc6ce8cf13adf7dc4034550b0acfc261e03f629fac12aeabe849037ff7196d32

Observation f4871b2c-0371-4777-8a3d-eedb6c0ba895 · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.277057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:d0a94d715a45e344d75d93893f238b7d1f93e3d31542c2d4f51f8ee06c9597d6

Observation 02848f04-40c5-4cc7-be1e-6f6c2432fe1b · inbound

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens cites this paper.

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:24.458005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:58:03.974053Z digest=sha256:fdfc1c5c1faa4670a5f6268d067a5ed8a10958ffaa01e60561ff2ee9232df885

Observation 3af64175-1c8e-4cb9-9727-eba43ed6d6fd · inbound

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection cites this paper.

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:15.287320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T22:21:10.133655Z digest=sha256:d5a31be0dadf5fd35c7799280e04f2a44f6e99b2d91804aa76b6d81305a38f55

Observation b6af3b54-1388-4e76-b3c0-4b9ffa81aff5 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.714030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:b2937426bbd4fb120ba1518388853e1902c99bc6b123f7097917e4172b986a0b

Observation 40a9a53b-5dbf-41d4-bd34-b9cde05af524 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.793520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:4bbe3d70c8013b9f29d4f98caaf2db13ade5fe3c7e1f3b07a8062a955359e6fd

Observation 879ca5e2-fb28-429b-9c95-a44b0ce6226d · inbound

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models cites this paper.

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.437930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:35:46.236860Z digest=sha256:816da36ba7b7d1e0e5d6459fee5dccb27d217c61da0ec8e270e6e7ac98e6843d

Observation f5913998-4216-4ce2-8925-878a650dcab1 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.881233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:77bfa7b3b0bf1930d03869a0bdfd444f3be17efa202cdd22ed01e9527b8575b8

Observation f80ece7d-474a-4a7f-849e-0cf729c50f3b · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.677719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:610ee6bc446dcf862a6d7241263d3824cd0be46fc4de2d04e0f12891e62ae5e7

Observation 48829565-1a03-459b-a7f2-600152fa830a · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.477984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:10513523f7832961280c156e1217a8c62c2e48c4bd9f9c186d6b55851e3834d6

Observation 02f63038-6940-4cd3-af38-483d52477528 · inbound

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation cites this paper.

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.559037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:50:22.839005Z digest=sha256:d3d880b16d44acbcfe34381f91d0b86319c3d6dff3d8ee0e4f0961a572afef38

Observation 4784d571-1a55-4630-99cc-5d6dea84cfba · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:03.046647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:03.046647Z digest=sha256:ed23bc6b36609a7164ea76b7b0fbb90dcf1f215dcd48b672db71d9cf6965d6de

Observation e794fb98-3746-4095-a919-f93c121b53c6 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.791253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.791253Z digest=sha256:2eaa7b1deb51da6d62bf7dc51800995836dea09fb99726bdb6e494956e030de9