Pith. sign in

Paper Citation Record · LEDGER

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 13 inbound Pith citation observations for arXiv:2505.05464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05464 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:08:16.658495Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:28:18.936793Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ccf3fd5c-4381-4b9b-aae4-790e0695fb5b · outbound

This paper cites Evolutionary Optimization of Model Merging Recipes.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Evolutionary Optimization of Model Merging Recipes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.422983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.422983Z digest=sha256:23edbc267f71e5bd4470f308b0bdb1d775751155770229db6af2ec548ec64054

Observation dd8d76f8-0905-4178-a219-a85d901271b1 · outbound

This paper cites REMEDY : Recipe merging dynamics in large vision-language models.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging REMEDY : Recipe merging dynamics in large vision-language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.486690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.430069Z digest=sha256:049965f3802a00dd6bae1c3236dca96cfe4723baae23f85fd81f173bf8deb1dd

Observation cde49d14-bd8e-459d-b01e-00aa23c5b8d8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.435598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.435598Z digest=sha256:74c555b503fe5aaa0868755810e31dcc0709301aecc86c01e1fa418dfe82c88b

Observation 8c402e13-1642-48cb-bd30-73d4dff25892 · outbound

This paper cites Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.441638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.441638Z digest=sha256:99ef7e543aa60a2eb12a3c7dee0f2a7720a9e3e436ba70e3dd7208fb1aa60137

Observation 6f8f40c3-4add-4f31-baac-6d8aa8f00aa1 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.446722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.446722Z digest=sha256:645ae4701ea2e7e0edb471ccb3bb6f1e661e3dd40821a9bf05690e819cf85807

Observation 0997f6d3-00c6-4dfa-bcbc-9925e639c325 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.466108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.451648Z digest=sha256:2c5e542af4ed4ed783fd18be37482ef87fba9959672a7cb5a13a68525b4c1356

Observation 5a06c253-00f0-4158-82f2-fa1b4e4bec81 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.456750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.456750Z digest=sha256:28cafd0d9fad70bce49fb1705e26a8e9d416ba8e58c7f71533b84efb4f050e75

Observation 4189c3d9-3df3-41ff-9220-22b1edf407c4 · outbound

This paper cites The Llama 3 Herd of Models.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.461410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.461410Z digest=sha256:834bea1a45a8b72d601f6c44d8826fa5f2cafd13cfbed8a7c22a64c16c8d9e34

Observation f74812ff-a28e-45b4-af5a-dda916c8d9ff · outbound

This paper cites T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.436551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.466252Z digest=sha256:db6c7c8b2fd95422b318992c1a43b0131cf4d8e628d23e3c644702ab43beda37

Observation 57462b7b-c950-43e9-a8f9-df24f4bc77ce · outbound

This paper cites Mistral 7B.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.472192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.472192Z digest=sha256:1da099e5de8741ae4244df059022ab363d501601e530acd68e7080d5fea6cb71

Observation 5f08ca0f-ae16-49a4-ac77-ec5b5db4bd5a · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Dataless knowledge fusion by merging weights of language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.412229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.477058Z digest=sha256:6c5890d132977063f1540b10d5c2834b2b8f4a4339fe52f013ec1468c8d9a258

Observation 56f4915b-39f6-4608-842d-2efc3cd56b50 · outbound

This paper cites What matters when building vision-language models? Advances in Neural Information Processing Systems, 37: 0 87874--87907, 2024.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging What matters when building vision-language models? Advances in Neural Information Processing Systems, 37: 0 87874--87907, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.393565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.482450Z digest=sha256:66603c0a0e1b4d95da65685883f05275ae67e4a34e557083a60f87069090a7b6

Observation 8cf25caf-15e2-4715-802b-c9528d975a69 · outbound

This paper cites What matters when building vision-language models?, 2024.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging What matters when building vision-language models?, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.376191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.489065Z digest=sha256:b24986ef88aa6ea48c51ef255118a05cb3a3d8e213dcc95c37fe084966bf82ef

Observation ad09fd08-c2ae-437d-9346-76d49900fc2c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging LLaVA-OneVision: Easy Visual Task Transfer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.494882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.494882Z digest=sha256:5851b66ca0b5ad442601fe63d19e7c26eb02e17a946e2b01259144d8114451f8

Observation 0d3e6ac4-67f8-4605-b331-2de84277439d · outbound

This paper cites L ogi C o T : Logical chain-of-thought instruction tuning.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging L ogi C o T : Logical chain-of-thought instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.499089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.499089Z digest=sha256:d12c49f6c0678c28f59c272a2f57b13d77773ba987e25327030e653921e21e9a

Observation f8a3714b-448c-4f30-a914-7fa1af3a5733 · outbound

This paper cites an unresolved cited work.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:08:17.360854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.503466Z digest=sha256:e09c762601ab52e6d8dcd17760e70cede6d1a4c746efe54ab7601a32f733bbd9

Observation 84667f08-28f2-4abd-947c-7ceb7ddea65f · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.509070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.509070Z digest=sha256:baa56bd1143336939d3804d933b5afdeed4c7bc43b337c09057a1154c7998fd8

Observation bde0602a-d2a7-4702-ae27-d545e40c5aa9 · outbound

This paper cites and Raffel, C.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging and Raffel, C

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.335353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.516428Z digest=sha256:d427e5012c78132a4d7b5c78b94514b4bdc781168e48e651fd0f7ba0bb63ee9b

Observation 2a2e1e3e-efff-4c76-99f1-2dc0aa9ff052 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.521800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.521800Z digest=sha256:01ab0b29b121dbf27cb814fc94076c0a3840bf2e9dd078b2ca9f1c1cb47ac05a

Observation dc922c32-eb26-4a40-983d-c8eaccbbac5d · outbound

This paper cites MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.532473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.532473Z digest=sha256:84aa7fe52ca7bf1c8a6bd28a52e558bc097dcd013bd2d652566ec0a4d853c2a9

Observation 99f00c08-1220-4658-9078-45431ee3810a · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.537631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.537631Z digest=sha256:8d72092d4e9832b4364ecea63a7c35c6322f78769f2fde912827fd0254d8d4db

Observation 0aea0911-47b1-4f82-8f5c-8be4f8bcccee · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset, 2024 a.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Measuring multimodal mathematical reasoning with math-vision dataset, 2024 a

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.297225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.542750Z digest=sha256:d5a97af0770db17842311b5b647f068341826309650c66f4168dbf770281b4d4

Observation e5453415-e96a-4976-9142-894bf7d0efdf · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.548458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.548458Z digest=sha256:fb1f7df068b366131ab87f569039e8f26ede55b5f148df8544b495a90d014d49

Observation 8a31c3ae-5984-44c7-bd30-5150a696efe6 · outbound

This paper cites Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.554851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.554851Z digest=sha256:c7a2b320cfe8c17d6a17d5365bdb295ca2573d6246e5f27b954c7eeb0e0cdf0e

Observation fff6c16f-60e2-4b84-ba46-6d30d5f1d899 · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.560353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.560353Z digest=sha256:c5540ba8bf1a1416784a64f3ff58432e8f9df1479515a52296a3939b121ab3c4

Observation 722bc919-0364-46c2-902a-1d514f241cb0 · outbound

This paper cites A., and Bansal, M.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging A., and Bansal, M

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.269231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.565100Z digest=sha256:fcabbdd3163e0271936445421d22d706ef4d2c2c8bb94b8db4a76627c24cbbaf

Observation e2c0f643-a7c3-4af1-8155-95926d035679 · outbound

This paper cites What Matters for Model Merging at Scale?.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging What Matters for Model Merging at Scale?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.571966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.571966Z digest=sha256:a6eb81dc5dcd01eb61a43bf2906ac379436fd03dc241306f4dc2b1ac4c864220

Observation 6146401c-192f-4e73-ae18-d9e70734471f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.577402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.577402Z digest=sha256:1f63920f032209825568bb8b77595f0de40a6e4b8f104c01fdb0296e7702db15

Observation cb39ff89-8176-4e48-9d2e-bca122b63720 · outbound

This paper cites AdaMerging: Adaptive Model Merging for Multi-Task Learning.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging AdaMerging: Adaptive Model Merging for Multi-Task Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.584752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.584752Z digest=sha256:a534e2e331ab11a1c671430430e92a02034e38b25d1e787598af9d78d1b855db

Observation e8ceb6a0-ae80-4260-8409-c6dcd1092dca · outbound

This paper cites Metamath: Bootstrap your own mathematical questions for large language models.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Metamath: Bootstrap your own mathematical questions for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.246881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.592047Z digest=sha256:8cea31b223162dd6ea8daaf6cd31fc9666a763bfded52b8aead832d8f7bf14ac

Observation 82340a4c-b78f-4d68-8bb8-a14dac31a9c4 · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.231811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.600090Z digest=sha256:3dcfde129478e0a5f85f95584beae2b12292e49f29dcc9b5a6b259a60c0032d5

Observation 4809fd7e-7b06-4fd8-aded-56e37ab6aba7 · outbound

This paper cites MA mmo TH : Building math generalist models through hybrid instruction tuning.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging MA mmo TH : Building math generalist models through hybrid instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.211154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.605856Z digest=sha256:533f431e1366e2565e15972454fca75dfbe7d69aef4dca7c67001a97d99b09cd

Observation e229af11-847d-4058-8984-2c2be6a21326 · outbound

This paper cites Mammoth2: Scaling instructions from the web.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Mammoth2: Scaling instructions from the web

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.184490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.614523Z digest=sha256:36ee05571570adc1e9295c1d9973f80ea51d72d535162e7dcb4e6c8e7b2eb2fd

Observation ee765e21-7378-4868-85f8-c2af11ecf95e · outbound

This paper cites Sigmoid loss for language image pre-training.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Sigmoid loss for language image pre-training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.621755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.621755Z digest=sha256:705a0b9ab5fe39606673a1efa3a69e2e4285826c5452c8803157acd4019bdf41

Observation 909c2e30-124d-4367-a7ef-4d2e7b45ce31 · outbound

This paper cites Composing parameter-efficient modules with arithmetic operation.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Composing parameter-efficient modules with arithmetic operation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.165996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.629317Z digest=sha256:01c058be3c99ebba02de8133deadac461a37f5136b5476d6864939abccda6416

Observation c9b4285c-884f-40ae-a2ca-b8ca5f4e0a31 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Improve Vision Language Model Chain-of-thought Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.635085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.635085Z digest=sha256:f71ee2e8272a9cc7f3134ab428e53f916ba831d97223c616d67495b848e904ac

Observation fcd7e9a2-a7a9-4a32-a579-8ec3abc0227b · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pp.\ 169--186.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pp.\ 169--186

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.642768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.642768Z digest=sha256:e85e4b3e47bf2cbdf35074ad0144692bb5ddc91b42c9aaa633c5ca4fa5c2dd27

Observation 11312815-6f71-4dda-a9a1-b46fdb279f7b · outbound

This paper cites H., ZHANG, R., Gu, J., Zhai, S., Susskind, J.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging H., ZHANG, R., Gu, J., Zhai, S., Susskind, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:08:17.137961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T23:08:16.648334Z digest=sha256:843a19a50b5946d66a749c81e9c1ad73c1ca8901a16fc7a4f1f6bfe983763504

Observation 9b5ba975-18cb-48b2-b248-6d6ab60b6e18 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.654069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.654069Z digest=sha256:9082702620135d38c774c05d40dc38e0659e19fe4691cf5c918a2a85101fe24d

Observation 77a79bef-b527-43e7-b51d-c03f2b4a677b · outbound

This paper cites write newline.

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging write newline

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:08:16.658495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:08:16.658495Z digest=sha256:26f898a157d5273c0af0df49ddd0b1c83dcd5e131d52fe4276064a4d294c272f

Pith citing papers

Observation 00e8cf36-898b-4728-a796-78f4cd00c571 · inbound

Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic cites this paper.

Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T17:01:02.057343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:01:02.057343Z digest=sha256:5263a11bbca1fd4e97b0df5a51507d452a43129ad3fd5994bfcb8f24d63858d1

Observation df9585c7-f145-4952-9ca8-d51d77c80e69 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.876272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:f2ac9517787bab3378c71854450618a04a5787c248807bd1f2c7a4f0242226be

Observation 1e1bae59-1e0b-4397-9f9f-566475f688ee · inbound

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective cites this paper.

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:18.936793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:28:18.936793Z digest=sha256:534a86006d709c6d5487ef824acce6e93435fcf06d2db8e1f003cd9f1810173f

Observation 52166a51-173b-4065-baed-1b8849038595 · inbound

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models cites this paper.

PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:50.323121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:50.323121Z digest=sha256:4f3895ed1e3a28ffb577fa7f911f36ac2f3adb84083be504a347ea055a670e86

Observation c9568ce4-e1e1-418e-97cd-37470082b171 · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.819405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.819405Z digest=sha256:f97e9dc93583a8d9e29e440698e2f52e404ec047a8e89685a09510cdde12529b

Observation 18bcdaae-1d8b-42aa-9472-b489a4c8ac21 · inbound

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models cites this paper.

VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:35.625880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:35.625880Z digest=sha256:b6bc0bbfd356661c06a61aab36f6c90d16937d62787f73fed80cdd296636a473

Observation 3c15fbc6-41d7-49d8-9fe3-6b4c95398690 · inbound

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging cites this paper.

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.779148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:42:43.948462Z digest=sha256:13e3ce752f1d4b48fc716aa38df5b604db206435993f14fe5ec363822453f8cd

Observation 5e991e75-8394-4b31-8ba6-d69525255154 · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:00.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:33715aad513b1d5ff9246fa6bdf563a45922d0bf48dc0f78719703ffad26d0af

Observation dd4b5312-f3f0-48b6-aeb6-f3645440beb6 · inbound

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails cites this paper.

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.487917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T23:03:01.151745Z digest=sha256:54a600ee4541e33f9e24f91d9ec329668aff5772683b02e0e36fa112fea05bb2

Observation f9a56e1a-b4b3-4e7f-852a-9ff13eb396d9 · inbound

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning cites this paper.

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.623754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T15:45:26.891621Z digest=sha256:964f438cf5e554efd9ee9a1bd11bf68df429b4d90ef78697ee2df41fb1db75ad

Observation 613f2f3a-19a1-4ec1-a37f-867336c1a40e · inbound

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models cites this paper.

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:56.348287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T02:10:12.974244Z digest=sha256:8f5955a43b48ec6e76598e04e72c3bb43ed7f0d70e275130579c9bf5ff97cffa

Observation e4417778-bbfb-4451-8ae9-3e4b983a85e3 · inbound

Recursive Vision Language Models for General Symbolic Reasoning cites this paper.

Recursive Vision Language Models for General Symbolic Reasoning Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:11:24.638371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:11:24.638371Z digest=sha256:741a8cda4ba56a47232c76655be332f4b1f89f76290f9b674f318dd7a894c303

Observation 97c0e90a-481d-4511-b135-8327fc93146a · inbound

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging cites this paper.

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:06.746182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:15:06.746182Z digest=sha256:e90b5e20284a9935b08b185fa8ad890469c0076809dddae6a3455308b84df6de