Pith. sign in

Paper Citation Record · LEDGER

The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.10462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10462 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:53:05.867758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.015082Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97348b64-4114-4889-9a2a-b871dc941f75 · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.804522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:3a9c7bdadb8577a39db9ccd848d5c7316b2b3680e4540add9ddcf8c48660eafc

Observation cb0a40c4-3a26-4ea0-bf95-5f2fffdf4283 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.474028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:29fd4535be059455a4a4a5decec6f3f320ae6e886d3ebbfebe9e43f0b7254cf7

Observation 8e3d7bf0-8038-4b4e-9817-32ed004715e8 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:05.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:05.867758Z digest=sha256:ae5b81dad8234db2fcaf6267f286e0e77f480bff0812dfaf01ba344fa20d48ac

Observation 42056ed7-d647-444e-86bf-c2139b686188 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:28.552429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:c5ec16e5e6f213ae28ccd5a84e1d492375d33dd01baca50a8e7adac9601092cc

Observation 9299fd8b-3ac9-4be1-b289-3704a39184a0 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:18:05.174455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:28303cfb4a15aca14ced5c4cf97407073d463a026874627a1b6d1b88c6193ab3

Observation d3dcd580-27df-4fee-9729-3b3de8ca34c9 · inbound

From Pixels to Words -- Towards Native One-Vision Models at Scale cites this paper.

From Pixels to Words -- Towards Native One-Vision Models at Scale The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:28.016525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:29:53.006064Z digest=sha256:89765795128e96987584667001e006ac1012b14cd8294aa0a93b68080a8e73c8

Observation 458afbc7-108a-4f0e-9017-244ed6609077 · inbound

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization cites this paper.

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:21.339189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:21.339189Z digest=sha256:857e35bd1f1a80dc074bbf49060c706505980dbbe187198986a4c11521ee3ea0