Pith. sign in

Paper Citation Record · LEDGER

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2505.23043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23043 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:06:58.520524Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.359705Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 844b3e6a-ea89-4fc1-b5cf-a92b9f6d2639 · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.086711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:798e0337048479f4d39a9c7d59194e984adbf3d6882823946986048d671c62df

Observation 46b1327c-347f-44a8-ab5b-a159e9c18129 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:58.520524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:58.520524Z digest=sha256:8130bec6771c27f89f192a4736aabe222327bfe8eb1391b6834cd0b2a13a0d7f

Observation 90b13ec3-a465-48dd-ba92-993348a6ce22 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:9e5a7bf232a5ce81454c9a7c1ad9152866eea13f2f2364c2faf06780dc66a198

Observation 3739057f-d01c-45e9-a4a0-4a849778039b · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 261

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.360990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:79e81ee98a84a0dd4cdf2b755efb51c32826f9595b10ef4eb066c3bda18b9415

Observation e2ee9f89-6e09-4582-958a-ab9830eeb666 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 261

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:f342e31548ff99380864ab7e2f7d23000a4411fffc7ff9be719e93c549b052a3