Pith. sign in

Paper Citation Record · LEDGER

The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.10462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10462 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:05.867758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.015082Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97348b64-4114-4889-9a2a-b871dc941f75 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.804522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:b6a7b07ae0be5e4c944ff34a905d0091938d17ef9ef51b9ff6944192d2aa363d

Observation cb0a40c4-3a26-4ea0-bf95-5f2fffdf4283 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.474028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:c159f6ae5a632bcfed8a01ecc931f2dd00dda04871b94a63ae718d66f4c3e464

Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.867758Z digest=sha256:ae5b81dad8234db2fcaf6267f286e0e77f480bff0812dfaf01ba344fa20d48ac

Observation 42056ed7-d647-444e-86bf-c2139b686188 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:28.552429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:57d07ae879c778a646f5ff5c4ade5f285c23daf764a0e268c88917a0c9178026

Observation 9299fd8b-3ac9-4be1-b289-3704a39184a0 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:18:05.174455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:539fc44cfbe6bfdcbaa57c27eb14cdcbf7ebdac46c33d70d15ba6b50524b886c

Observation d3dcd580-27df-4fee-9729-3b3de8ca34c9 · inbound

From Pixels to Words -- Towards Native One-Vision Models at Scale cites this paper.

From Pixels to Words -- Towards Native One-Vision Models at Scale The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.016525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:c345c0565a87df126d76b1c4560a4142a257d2bcc118077d31affed667de2994

Observation 458afbc7-108a-4f0e-9017-244ed6609077 · inbound

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization cites this paper.

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:21.339189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:21.339189Z digest=sha256:75f1507f31c6fb621f372427d2c58a3c9358173fb9712301f9e9a8a0efbfbcbd