Pith. sign in

Paper Citation Record · LEDGER

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

As of 16 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2505.19406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19406 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:04.977166Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:11:57.228949Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:29:53.301476Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2015ac20-33ef-4a2b-983f-c547a9ca023d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.915052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.915052Z digest=sha256:c265564c1be875964a22cce2367750e4555fe797e2627ee36a83d3871ede33d3

Observation b1e26433-bd32-401a-a42f-42851105daa1 · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.318580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.318580Z digest=sha256:c7a2183ab6c7a721f75812bc836e72283fb3069a660e4f5fa65638d8e5772f1a

Observation 18acaf29-64ab-4962-bfb0-4fb940ccd2e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.504324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.504324Z digest=sha256:97deebbbc0f34f4122d3223b702801e247ea0477c3cf9c50967cd566974e388e

Observation 82ab3d3d-13f9-4382-8cf0-d5cfdb58dcb1 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.650097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.650097Z digest=sha256:418a211d9424137ca3afd21765a6ec296e038f1ca333753204aff6620e59dced

Observation 4e063a37-5cf6-4970-b35f-34de18d07322 · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.805674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.805674Z digest=sha256:6c7d249af8a8ea9e646100aa42b3aa8d19b73be357f45c27e6b51c8212640b19

Observation bdc4a137-b272-4730-8576-cbc00ffd0c81 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.964530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.964530Z digest=sha256:83b4e7ce5a4185e02bb0df01f407d26e6905a517667f034971aa6888b7551545

Observation 0a2542e1-2b09-4150-a480-e600b0aea9d6 · outbound

This paper cites OpenAI o1 System Card.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.067077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.067077Z digest=sha256:f044f8c7e44cffe2cf77f4e2fa4cc92ce0748b5f1bb6648c128e4bf4a5649313

Observation c20c6894-9fcc-4cb0-a18d-f9177110cd8f · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.201102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.201102Z digest=sha256:ab487be58c1a50c9b046a8e7bc88a0705bd29b61a4abd2e8f9c2d41a63461709

Observation 5bfc92f5-42a7-4f04-9283-7df32d795d5b · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.442626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.442626Z digest=sha256:0d559efc08f1c2ef61a628e2321af77e34983577795e2b5e3eedc62411509238

Observation a380a69c-12d4-4f35-a5f8-1577f7eb76b2 · outbound

This paper cites Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.584308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.584308Z digest=sha256:ae78bbef69ca268c44fa4ea9ac091e76b4f1d0bcc447cd2dc6a31780a3ae81da

Observation b9799e86-71b6-418d-afac-5e6ab7f019b4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.674719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.674719Z digest=sha256:26e1600caccaebed4f2216ca6dcb714d03e6193e4a11756a6725c2509874c578

Observation 6a944832-4ef3-4194-9ba9-a99067b7f3c5 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.823999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.823999Z digest=sha256:c8dc1f460d3de591792ee2c84705a50c40cbb674d0b6b41cfe7bc952db9a58da

Observation 9de96a95-7ef6-4650-b085-1905bfe6dfd2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.941349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.941349Z digest=sha256:5023a20fb6f9e45050c19dd981163c9eeca71e7a012e9d7755597284913d8885

Observation 2c5e2e8c-d1f1-45c9-808b-d94d2edc3773 · outbound

This paper cites Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.095763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.095763Z digest=sha256:c9907df97609e4c30b1844a8797d58cb87b212dc54394a7e03e9c7599b3834a8

Observation 05057945-d3bb-4106-8c93-394e0f6d9656 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.218387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.218387Z digest=sha256:1cb54bea6419c1d181a72712dfd853d00d13fd30180c7ebcd052446686391b8d

Observation 47598c7c-5742-4d49-b975-a4a009bdb9c0 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.371188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.371188Z digest=sha256:6430b289133126231f74469901bcf2f4a13b18b0e7b5a6f41697f864a5e08238

Observation 9c7d8425-f4b5-4bbf-9fc8-2eb76722dfad · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.527707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.527707Z digest=sha256:220d35481789a70ad15e63090d071232f9e5775c4ca3f910b4bf7dc1f9cc402f

Observation 2f88aef6-8697-47bb-a8ff-b2997eefa401 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.657935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.657935Z digest=sha256:139df6e665bb86fadbb42327cb67927894621914e2184d2de2192e84a6f2ef35

Observation cd2034a0-62da-402e-93e3-4bcac76dd10f · outbound

This paper cites Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.798512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.798512Z digest=sha256:b0eb316558d089731da40680aa4c6a31180a20b7b571c8dc5d21dc115f448d98

Observation 8b14d4b8-a2bd-44ab-8590-7b87811c4b6e · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.977166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.977166Z digest=sha256:b85421a05b01b89a3c9b35efa612cef490623dc7db72209635012acb5ae594e7

Observation 71bf01ed-c9d8-45de-a256-2f291490d0f3 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.335586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.335586Z digest=sha256:2f217de3f797684808017b2e8d79093d3ceb584fb08014ea42a0c4d2a2a44ef3

Observation 86d89c6b-f0d1-4f19-b05b-551e409a7a4b · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.165225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.165225Z digest=sha256:bf9b7053284cbf6a465d2d704a660b174b4581709dedd66c360a596d9e0611df

Observation f295fbe9-90b6-48b8-ad51-5d90ac34222a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.977150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.977150Z digest=sha256:094023ab317b3069b7b79cfa326192daeff692a3756b339d9be7e76d70226ac0

Pith citing papers

Observation f2b08c53-5e61-4ef5-888e-dc17d57686b4 · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:29:53.304356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:ff87405bdd6266f5b5b875ebfb28be6522bea4ee2310c58527249fbce995b8cc

Observation 031bfc5f-2dd0-4126-9ef8-6cc40ce19550 · inbound

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval cites this paper.

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.304743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T15:11:57.228949Z digest=sha256:6fec1fac8a966f6f3b707bd6d288134cbfc5259bc23075e37e9d06ece21f4588