Pith. sign in

Paper Citation Record · LEDGER

Diving into Self-Evolving Training for Multimodal Reasoning

As of 15 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 10 inbound Pith citation observations for arXiv:2412.17451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17451 v3

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:32:39.075363Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:58:29.367654Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f711b940-b588-4b8e-a534-e579099ee3d2 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Diving into Self-Evolving Training for Multimodal Reasoning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.774438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.774438Z digest=sha256:76102777e21ed728ff0f79e81544036f9532131f5be5bdd890122efad8c03595

Observation 34c0833f-b3d2-4e6c-940d-429b87bd4460 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Diving into Self-Evolving Training for Multimodal Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.780886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.780886Z digest=sha256:a4f3969796e6c071626b03b64556452c81605dc29f0fb104e59df0c8e309c1bd

Observation 6d00c2a2-1d33-446d-b02a-02850e139d04 · outbound

This paper cites InternLM2 Technical Report.

Diving into Self-Evolving Training for Multimodal Reasoning InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.788722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.788722Z digest=sha256:257a9e4e6e6caeb123de0351801357ed63057905903e090251c31e0da87b8ce1

Observation 48607c01-32b4-403b-8dfd-6591e1ada89f · outbound

This paper cites H., Kwon, T., Kim, M., woo Kwak, B., Kang, D., and Yeo, J.

Diving into Self-Evolving Training for Multimodal Reasoning H., Kwon, T., Kim, M., woo Kwak, B., Kang, D., and Yeo, J

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.795835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.795835Z digest=sha256:c0643b2f3cc9269fb332c0c304adf70749dc8ef7925e466f0230569e742e9f5d

Observation 811197b6-32ab-4737-97df-ad071e627c3c · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Diving into Self-Evolving Training for Multimodal Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.802488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.802488Z digest=sha256:265e2c692b5af4ee3618e27b645962de0d4745c3aef8c97650ede10eacc8d872

Observation b2973cdf-adc0-4b74-bc35-b417c65ed775 · outbound

This paper cites M ^3 C o T : A novel benchmark for multi-domain multi-step multi-modal chain-of-thought.

Diving into Self-Evolving Training for Multimodal Reasoning M ^3 C o T : A novel benchmark for multi-domain multi-step multi-modal chain-of-thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.808839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.808839Z digest=sha256:3607547a23a2d924bb579bde69e85e766eafe1965ea24c6c1716f74a7e5bb308

Observation 77a42cc4-b89a-4c6d-aae1-52cb8455ced7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Diving into Self-Evolving Training for Multimodal Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.821457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.821457Z digest=sha256:15a7992a98457177a177bdb45068b6e48a0fd5f17bd0fdc6f0c815519bb66a5e

Observation b800a505-5f2a-4f43-8fda-25933e118368 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Diving into Self-Evolving Training for Multimodal Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.827443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.827443Z digest=sha256:54350852156ea9395d2b72f1255b71076ba66afbb90b3b6ffabf91362732ea6f

Observation 6a5ffa05-c1b0-4db3-91fe-d36ed9ce6e90 · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

Diving into Self-Evolving Training for Multimodal Reasoning Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.834249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.834249Z digest=sha256:693ccade2922e4c29dd1f83239772824c4240851699bd7a15f6b004571147c35

Observation 2587cb1e-cb40-49d6-ac8a-b0ebdf8b05c9 · outbound

This paper cites RAFT : Reward ranked finetuning for generative foundation model alignment.

Diving into Self-Evolving Training for Multimodal Reasoning RAFT : Reward ranked finetuning for generative foundation model alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.840905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.840905Z digest=sha256:c70c8f2d1e59630e2c8db3a7cf6fab8e9c97c4cc88409d6de228a729b69a0292

Observation d989aa54-570b-4353-a095-ae05d791c607 · outbound

This paper cites The Llama 3 Herd of Models.

Diving into Self-Evolving Training for Multimodal Reasoning The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.846788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.846788Z digest=sha256:ab0b87aebc7b013abcd5e44c068ae0f9274010c6e9673a151d21edb835952af8

Observation ffbc9fa3-e027-46dd-ba36-5934b04eadc8 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Diving into Self-Evolving Training for Multimodal Reasoning VILA$^2$: VILA Augmented VILA

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.851989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.851989Z digest=sha256:23ff4dd168832bfc0fde42ac9677c79514032cefe245b235e3f2b5793ff3c6cd

Observation 826696bf-c942-4d62-ae1c-4d635417b65a · outbound

This paper cites Incoder: A generative model for code infilling and synthesis.

Diving into Self-Evolving Training for Multimodal Reasoning Incoder: A generative model for code infilling and synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:32:40.297445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:32:38.857303Z digest=sha256:ac948685429d2336fac9d2f5c71f50d911d2e48c50a09cf5e2b4fd30d6ad6276

Observation 5c278405-21c9-4421-bcf6-27aff4bbb879 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Diving into Self-Evolving Training for Multimodal Reasoning G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.862368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.862368Z digest=sha256:b7f272248a3a2acfdd0fc90a4242539ef007785bfd0d9492d6c6729df8811b28

Observation 969ad025-6c86-4cf4-b28a-e3082c24e5c0 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Diving into Self-Evolving Training for Multimodal Reasoning Reinforced Self-Training (ReST) for Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.867648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.867648Z digest=sha256:514a8b92f7650afe4369bd6d984f5f7bcfb276f1cff76ed442817ba0545ccb3f

Observation e348a9a1-f9ca-4d47-8d74-ed40ae29640f · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Diving into Self-Evolving Training for Multimodal Reasoning V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.872790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.872790Z digest=sha256:5b237edde29a21a5c4fb9a3a4064d4abf6e8a15783c415112a7d1864a9db149f

Observation 0e45c6ed-6f4d-4a22-866e-93d0621f8740 · outbound

This paper cites Large Language Models Can Self-Improve.

Diving into Self-Evolving Training for Multimodal Reasoning Large Language Models Can Self-Improve

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.878759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.878759Z digest=sha256:40b5b2fb6128ffbdf9d539c56b0c96388bd2bad0753533347e9ff510d861b224

Observation f69d33ff-0a87-4061-93eb-208e75244e8d · outbound

This paper cites A diagram is worth a dozen images.

Diving into Self-Evolving Training for Multimodal Reasoning A diagram is worth a dozen images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.883351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.883351Z digest=sha256:e0faf392d091f6f2c321da792369ecd8b99dee488996e598e320132b11798a92

Observation c4633b1a-269c-466f-9b52-d77707814ff3 · outbound

This paper cites ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation.

Diving into Self-Evolving Training for Multimodal Reasoning ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.888290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.888290Z digest=sha256:01ac5d32ba0d3cb30d86ee3b7f8efffef6dda5681194527824438cbdee2eecdb

Observation 2946bda4-133e-4ea6-a6eb-6b3882d44d95 · outbound

This paper cites Let's Verify Step by Step.

Diving into Self-Evolving Training for Multimodal Reasoning Let's Verify Step by Step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.893520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.893520Z digest=sha256:6b60ed3ee86c16e373aed300aca46e4251ccb104e6f552c90b0497f26e8f99ae

Observation aad732cd-7290-4c63-995a-1b85766ec56f · outbound

This paper cites an unresolved cited work.

Diving into Self-Evolving Training for Multimodal Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:32:40.270810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:32:38.898704Z digest=sha256:28342c96c88e868b33af4bd09f0b75a28892a3155c1719af3c65fb696ccf99aa

Observation d1470872-d6d6-4705-beab-8cb192fa6bfe · outbound

This paper cites an unresolved cited work.

Diving into Self-Evolving Training for Multimodal Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.904473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.904473Z digest=sha256:b77e829e6db03561292fefeaa429a3e72e9153d12fe534c42494e9fa49d3f606

Observation 821f263e-fb7c-4416-bf35-602f12aefb39 · outbound

This paper cites A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation.

Diving into Self-Evolving Training for Multimodal Reasoning A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.909320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.909320Z digest=sha256:30cfaa1694ac3523437f016d502b2115eddda1dbb29ce0c4cc6a5a3f73f4f165

Observation 8600d892-1138-4d4e-a81b-c5c9164fc385 · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

Diving into Self-Evolving Training for Multimodal Reasoning VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.915513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.915513Z digest=sha256:6bffdfe090a2f8f47a7c7924c4c88f15b2e4c0767bdc3d8bf60115e77687df19

Observation 55e769cf-df5b-4dc0-b08d-9d274dcf0e8f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pp.\ 216--233.

Diving into Self-Evolving Training for Multimodal Reasoning Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pp.\ 216--233

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.921385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.921385Z digest=sha256:68efaea59e2c64530b6c6004a0846a6886cff5f01b21ecbf7e252b38d2f751f3

Observation fd22d068-7a4d-40f7-b0b7-e141f9275fca · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Diving into Self-Evolving Training for Multimodal Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.926897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.926897Z digest=sha256:194cf7a469f678be8b12450cbac502783b02a246084518eb1ba24eca2f977e1d

Observation 6f0d2b5c-08d9-43cd-b7ed-c501a75db81b · outbound

This paper cites L., Tan, J.

Diving into Self-Evolving Training for Multimodal Reasoning L., Tan, J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.934182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.934182Z digest=sha256:6015841f6352b80d6543d9a69ea67b0048055c98a4916b240021d6edd9b425c8

Observation b192fef6-e7ae-4733-b147-49850334b8cb · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, 2024.

Diving into Self-Evolving Training for Multimodal Reasoning Introducing meta llama 3: The most capable openly available llm to date, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.940538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.940538Z digest=sha256:212a5b2aa44878d640e2fecc9828625851a7d4dea4ecc1e4edb6885f50dc7979

Observation 93151882-fac8-45b9-98ca-245f5337c5ad · outbound

This paper cites Iterative Reasoning Preference Optimization.

Diving into Self-Evolving Training for Multimodal Reasoning Iterative Reasoning Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.945911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.945911Z digest=sha256:3dc2c8753ddc6bcb7c1847eb16ab791c7fb2854ac64e0ed8a8ad4066021016c5

Observation 43a5cce4-cb50-45cf-ac97-1b6fabdb673f · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Diving into Self-Evolving Training for Multimodal Reasoning W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.951753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.951753Z digest=sha256:1915921b678d6a0b9ebe4df8d1a50f3be88f677f3e97e7a6e374b6952871335a

Observation b5bb887b-d140-4278-b32f-ffc07ee6f29f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Diving into Self-Evolving Training for Multimodal Reasoning Proximal Policy Optimization Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.958511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.958511Z digest=sha256:5b886c897037cf38ce0ff0e7a5b43a0bd7ee7f8b275ada172c03a4a17cb2efb4

Observation 6cf46a2d-ccf4-4153-b15e-830695889de3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Diving into Self-Evolving Training for Multimodal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.964101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.964101Z digest=sha256:e4280b1e9ee9f2c2aed04c0a708f3109759a9b714559a3c4697f5464a1dfbf64

Observation 8ec73ec3-d9ad-484e-9584-50f3fe4e18ff · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Diving into Self-Evolving Training for Multimodal Reasoning Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.969733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.969733Z digest=sha256:3e00fae5456bf3bc24b1da9e5a6ec83b5241e558d8e164286b1f44754855bb46

Observation dff38718-995e-490a-abe9-9cc6bbafe7eb · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Diving into Self-Evolving Training for Multimodal Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.975103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.975103Z digest=sha256:1899a0e7c1e73e6008ef45b57daa92ab6867e2ef144c57a1a9d48864717362e7

Observation 493225b0-a140-448a-a552-4826cae769ad · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Diving into Self-Evolving Training for Multimodal Reasoning Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.981633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.981633Z digest=sha256:a1c88ce31fc65c3d995102edc07382319cdc9e122319a7a66546fe6cf3a01f21

Observation 5cf9947e-8342-4a8d-84d5-709b272f4580 · outbound

This paper cites Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations.

Diving into Self-Evolving Training for Multimodal Reasoning Math-shepherd: Verify and reinforce LLM s step-by-step without human annotations

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:32:40.068328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:32:38.987376Z digest=sha256:72dd8f13b44ce128f9f422d45ee21161e8ebbd9125d4b2f4bb40e4268c506680

Observation d6acb88e-0c2e-47ab-92aa-d05159bdc3a0 · outbound

This paper cites Progress or Regress? Self-Improvement Reversal in Post-training.

Diving into Self-Evolving Training for Multimodal Reasoning Progress or Regress? Self-Improvement Reversal in Post-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.992525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.992525Z digest=sha256:bcda34640a2ca22d3fd1ffa0e026f48c97f7d3447ef5607b4d96b335073f6ce4

Observation 41e40ab7-0702-4e79-9cf8-63cdc68ae477 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Diving into Self-Evolving Training for Multimodal Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:38.998633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:38.998633Z digest=sha256:aa41332f420111e3e918d32b2709fa4b181eb9f3a64f81e9a796fa2e526f5c3a

Observation 8711f226-4e21-4c89-851c-71bb01cdbaa9 · outbound

This paper cites ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline.

Diving into Self-Evolving Training for Multimodal Reasoning ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.004817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.004817Z digest=sha256:05425e8bfbf0f2cbf8e5cbf2be7e5ae78292908b1c1cb6dafa4b2161674ff7c9

Observation 25e57039-bf7c-4185-9448-6f9feaf9f5cb · outbound

This paper cites LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding.

Diving into Self-Evolving Training for Multimodal Reasoning LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.011578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.011578Z digest=sha256:e2c170a7a5c55615bfb22b1db2ed8ae97ccea203e9e1992a4998a9a230a6838a

Observation 2bdbebda-d4f6-475e-a1b9-1b4c876c8e13 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Diving into Self-Evolving Training for Multimodal Reasoning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.016623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.016623Z digest=sha256:fedf23c937314b3107d9edb4a9ce983e4e20eb928722acc1ae1075144af4ddac

Observation 2c808077-1d9b-44f7-91e2-b63224468d19 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Diving into Self-Evolving Training for Multimodal Reasoning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.021608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.021608Z digest=sha256:f80098896b0c1d4d5a4e6e9ebd59b694e9ad66791eef74189514eba1b5dc32fc

Observation 346921c3-282a-439c-b4bf-1f547f5ea9a2 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Diving into Self-Evolving Training for Multimodal Reasoning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.027142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.027142Z digest=sha256:18fea26626d340ff01fc054bba24c2ab81f0fab451cb18d3e9f42d72a6a75953

Observation 444f5f3e-4ae1-4f76-8e2b-4644c154d6e8 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Diving into Self-Evolving Training for Multimodal Reasoning MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.032356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.032356Z digest=sha256:dbb0ca8a7ce03006fdb1bc1dcea899ad7f4969f64e45663741a20935c22fca6e

Observation bdb4cfa7-e5b0-4dce-be7a-f11e06b53167 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

Diving into Self-Evolving Training for Multimodal Reasoning Star: Bootstrapping reasoning with reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.037946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.037946Z digest=sha256:d1b2e217a80cca0b03f5b2db3b7f950618bd1a3d675a2fe1805dbb35d425c65d

Observation 9528d355-cb91-468a-88e7-7668b1f77f8b · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Diving into Self-Evolving Training for Multimodal Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.046843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.046843Z digest=sha256:035a086612c05cef02ec9f40682da8cbd88646e0e35725966ce2285f1c3ac306

Observation 2b3774b4-3156-4f8b-b5bc-62f8c1cecd4a · outbound

This paper cites B- ST ar: Monitoring and balancing exploration and exploitation in self-taught reasoners.

Diving into Self-Evolving Training for Multimodal Reasoning B- ST ar: Monitoring and balancing exploration and exploitation in self-taught reasoners

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:32:40.038164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:32:39.052894Z digest=sha256:d5bb8e909cebb10053a1ab140119b50e4f49b08ef37db25c7a28a2d786e82475

Observation 7bb639e3-64aa-43b3-89a1-7bfe5d2d98ba · outbound

This paper cites Sigmoid loss for language image pre-training.

Diving into Self-Evolving Training for Multimodal Reasoning Sigmoid loss for language image pre-training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.058274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.058274Z digest=sha256:489fcdae9f1f76a4ebf130064695bc94ba8f70c1fd67f2f7ad203488142bcc9d

Observation d7dc7ea2-d931-4bfb-8a99-abcde8ad81a0 · outbound

This paper cites Re ST - MCTS *: LLM self-training via process reward guided tree search.

Diving into Self-Evolving Training for Multimodal Reasoning Re ST - MCTS *: LLM self-training via process reward guided tree search

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:32:40.019254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T05:32:39.063984Z digest=sha256:c13d7eaf33faa61c2f26fa1c62f12d973c6867594d7d82fc6bca859c4ca72737

Observation c6ba98ab-5eb9-4534-b122-ca04f7d1ad34 · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Diving into Self-Evolving Training for Multimodal Reasoning MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.069771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.069771Z digest=sha256:7713e7e3d0ed80feb6e894c7ee2918041ac78a1cc0205203fcaa877c8fa66ee6

Observation fd60f60c-468f-4ed7-9a24-ce632c310fc1 · outbound

This paper cites write newline.

Diving into Self-Evolving Training for Multimodal Reasoning write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:39.075363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:39.075363Z digest=sha256:07c3c16a12748ff16ef6b1a001a35191f45334aae3cd94200d6531f05a0bfbd9

Pith citing papers

Observation f1986e7c-ec77-42a7-98a7-b29c2dbfc94f · inbound

CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions cites this paper.

CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions Diving into Self-Evolving Training for Multimodal Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T10:58:29.367654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:58:29.367654Z digest=sha256:413036e585e8efbca329cbed3337d5ee75b48c46c9bc513540d68a046cd1ff23

Observation 8b622981-74cc-499c-9762-b210c1de0bbc · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Diving into Self-Evolving Training for Multimodal Reasoning

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.226918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:d5fa7d97338f82c0e50c8c0fc626000fac9921dd43827e199002ab28376ac3f7

Observation 8745bd44-0a92-4216-839b-d168c0e96028 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Diving into Self-Evolving Training for Multimodal Reasoning

Reference 241

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.551812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:c80f0dace54b349b798577029ea6201c9b0a8e53d684bf9bdb799cb68ca7f243

Observation 0ad5efed-1cde-4734-b7ec-3e0100d528ac · inbound

Realistic Evaluation of TabPFN v2 in Open Environments cites this paper.

Realistic Evaluation of TabPFN v2 in Open Environments Diving into Self-Evolving Training for Multimodal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:16.756120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:16.756120Z digest=sha256:087972a087607b84b33e5c20a745ebc6ece0ba0cbad7e8ef5586b475b56d5a51

Observation 1208e7d4-dfad-4aa9-ba0d-592b73ae6df9 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Diving into Self-Evolving Training for Multimodal Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.938364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.938364Z digest=sha256:18e93786935efe21ad7431bc46c4d67c201dbee97389ddb0d4ea4ba4e9be8848

Observation c7716c05-1979-4b7e-9fbe-c8f56a9c90d3 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning Diving into Self-Evolving Training for Multimodal Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.403063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.403063Z digest=sha256:cafccb83dc2024f045641036664845c3af300ab7225d24835dc9916425b74333

Observation 682d4100-6566-402a-92b2-bc73aefc2ca9 · inbound

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations cites this paper.

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations Diving into Self-Evolving Training for Multimodal Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.902622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:23:45.814042Z digest=sha256:515ca0fcfcbfd3db050f13c26ee63f3f0b834396a7ab933faea7bb08953476c8

Observation 5aa54b37-2009-42e7-9242-1e423e50f7bc · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Diving into Self-Evolving Training for Multimodal Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.663287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:17a6ed19f6ae7cf920dff13ffc04e4e6877303f8c459eebf4924951afe99e559

Observation aad56041-7083-42d2-bfe0-3fb4067aa380 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Diving into Self-Evolving Training for Multimodal Reasoning

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.314324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:db043e44d949d4749fed789462585655bf216a19648a04a56c31ae455e416f42

Observation 7ce451c5-430f-4eee-9485-b72b089c4062 · inbound

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models cites this paper.

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models Diving into Self-Evolving Training for Multimodal Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.974194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:01:47.207296Z digest=sha256:eb8c498886710b33bd661cd4b63aa6636b396d2ba315cc26f495c5cc74f1e0d2