Pith. sign in

Paper Citation Record · LEDGER

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2507.11114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11114 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.476211Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T03:00:51.318412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92470eb8-9efe-42d2-9980-248036356263 · outbound

This paper cites Zhang, J.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Zhang, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.905378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.393362Z digest=sha256:76ab28f379c4d1b32212f7f568febf47e01c52bdf2809c2ec57ded30d420fb59

Observation f0ebc0d1-93a4-4554-b8ad-342654e74dde · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.404396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.404396Z digest=sha256:efbe70847f7d0a74191805c91b375647b8a9ff8d337c9dbe4792625575059011

Observation 0c1a05da-e30b-42c1-a013-8bcd0f13e334 · outbound

This paper cites an unresolved cited work.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.408964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.408964Z digest=sha256:cad9b564a74e7e48c64fb9eb9de7a7bd499070526ce5eca5b926e733e74130ec

Observation 3ad7f9ac-8795-4f5d-95a6-fbdb1fba706a · outbound

This paper cites an unresolved cited work.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:31.894357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.414629Z digest=sha256:c55b3fcd268d5b8d44db89d3cb5cb6bf005dad13284a67bef28fdcbce122d728

Observation a6f10935-7194-4020-a331-c6cc9f155762 · outbound

This paper cites Huang, et al., M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models, in: NeurIPS Datasets and Benchmarks Track, 2023.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Huang, et al., M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models, in: NeurIPS Datasets and Benchmarks Track, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.884289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.421210Z digest=sha256:57e02422ee6f4fd39a2eef7bde40938734e45f4f7fbcb93015bf1758b0366ff6

Observation 4ca57af0-7765-4b7d-88df-1aa8fe32a921 · outbound

This paper cites an unresolved cited work.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Unresolved cited work

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-06T17:21:31.710786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.424431Z digest=sha256:c76ec4b20b08325689b68b3432149f35d132c247d3f60f83928a12c878a6cb67

Observation 5c6d92af-a4f8-4e0d-9de5-f11ea9f2a027 · outbound

This paper cites Dimitrov, M.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Dimitrov, M

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.873465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.427692Z digest=sha256:44108278986b9f83b57444655226e72915bc093ce7320069b4be79359d2be1a0

Observation f8bdb6af-5c30-4b2d-862e-2737345e9e68 · outbound

This paper cites Ionescu, H.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Ionescu, H

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.863199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.437470Z digest=sha256:07f0914f1b68278d31bbacb8c5dec8af74645e4b6751512b5cb555e4bd56b8fc

Observation a7672fc2-0c8b-453f-821f-56b28ce2e13a · outbound

This paper cites M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.440457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.440457Z digest=sha256:f480bc2582ab96bebe364b81df5f4f71696ea9fb5d7535ac4fff8e547444f7bc

Observation ac521802-bc06-497e-8b97-05af889726f3 · outbound

This paper cites Zhang, M.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Zhang, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.853022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.444814Z digest=sha256:011127cdc3273bc1b1194091fd7f469b35cd6f02bba71786a8fe5ca0992d97ee

Observation 55962e93-2d13-45da-866e-11bf31b14607 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Language Models are Multilingual Chain-of-Thought Reasoners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.452850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.452850Z digest=sha256:8886d1b79c0f4252780d586a4e99d1ef318ef8341ffcd1e09eef1192b85ee076

Observation 3c67d9fd-629d-4a86-8262-13bf014a3514 · outbound

This paper cites an unresolved cited work.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.455296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.455296Z digest=sha256:e12e04ebc6acd3579cf478af2585e1f944e43665b10a937041ec818d58809de6

Observation 3d562751-8eec-4e9d-b3d0-72c93878a930 · outbound

This paper cites an unresolved cited work.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:31.841294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.458593Z digest=sha256:fffde365c4ac50a37b8387e497acba89b27f0e7058c9e3aa500196aff552ed8c

Observation 43f61f8a-3aa4-4aee-b567-374803ebd6ff · outbound

This paper cites Accessed: 2025-03-15.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Accessed: 2025-03-15

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.831966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.461180Z digest=sha256:2e5aa1fc37f5cbf06310a2714a7e7c23066d24118e6fdb2fb5b7699aef124f0b

Observation 63caafe2-1361-4995-9fad-3dbb5499b695 · outbound

This paper cites Accessed: 2025-01-10.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Accessed: 2025-01-10

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.820829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.463774Z digest=sha256:3b1a55a516be242126848191f75b998c516c052a6e4a36923110d2d5ccd219f9

Observation 0f126019-ecbf-4787-b9d4-364c8322d39f · outbound

This paper cites Accessed: 2025-05-28.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Accessed: 2025-05-28

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.810382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.466903Z digest=sha256:21738fe3732ecd67066e5fbcd3598d8a088afe5c7b98ca2a6c2c25124a594522

Observation c849ec16-283d-49a4-8dfc-a9fb5dc2b533 · outbound

This paper cites Phi-4 Technical Report.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Phi-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.469737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.469737Z digest=sha256:0f5887a0e824b79c27a628194167cea8d64c720c36bd266fdc5d6e35976ca135

Observation 58cbb8ad-5ddf-4a5b-ba60-a8177482ab0d · outbound

This paper cites DeepMind, Gemma Team, Gemma 3: Advancing open language models, https://blog.google/ technology/developers/gemma-3-google-new-open-model/, 2024.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models DeepMind, Gemma Team, Gemma 3: Advancing open language models, https://blog.google/ technology/developers/gemma-3-google-new-open-model/, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.800766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T17:21:31.473297Z digest=sha256:71d476595e316edcd7b61e9d761fed6e06c6a019ddcf7715da5d057eff371358

Observation aea30014-7dc6-4b96-9110-31fe232355ff · outbound

This paper cites Mistral 7B.

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.476211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.476211Z digest=sha256:52451122c04c3ec3f83ab62f644e807847a373a72e02deb7be993b3dbe84bb5c

Pith citing papers

Observation cf2ba68c-894c-4090-8107-87739cb3a732 · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:55ae6db8363afade0b444589c249db4826611410fdb9c14184efa696f88ba38a