Pith. sign in

Paper Citation Record · LEDGER

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests

As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2506.07418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07418 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:39:00.747897Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a484e773-47e5-42ae-b5d5-8fb790abd933 · outbound

This paper cites Zhang, Y.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Zhang, Y

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.503087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.536458Z digest=sha256:ad2a2de8606c0a64b94a57a0b3a460f0835da069cb3cf773b6f3ccf4ffd4c074

Observation 3aae2431-7507-489e-bea7-b08a9baa8cb8 · outbound

This paper cites Achiam, S.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Achiam, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.313361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.650734Z digest=sha256:b60792847318388a1f6ce32cc640a4af4d53a939ee5b43e81ce4b8ea96fd7c6e

Observation fc18f617-188d-4cea-a35f-ca373b0d3f75 · outbound

This paper cites Qwen2.5-vl 7b.GitHub Repository, 2024.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen2.5-vl 7b.GitHub Repository, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.069406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.737402Z digest=sha256:7144f9743cfcae28d9f0457b9faf7f6b121ddb8b683065a6aabae96d2f6ad4e4

Observation 42929a52-a881-4b7b-bc17-e02c0c02a3e4 · outbound

This paper cites Visual Instruction Tuning.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:56.825172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:56.825172Z digest=sha256:fc70d14a02a67fee6a2b5b3d6c3bfcdbabcb34350811d9f81f09183aadd42134

Observation 2630dfda-3d82-47c3-8b56-eef2f69ea43e · outbound

This paper cites Andreescu and R.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Andreescu and R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.898084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.904032Z digest=sha256:d0bc4298e6cb514595fc713f228c42824b03db9f755ce496316d3136e844e276

Observation 93700d1e-7cf0-4a5d-8199-fbe04144cfe1 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests A Survey on Benchmarks of Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.018980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.018980Z digest=sha256:d0c6c51c980e82ad962486ddc04bf9367c933ff09de537217678f4df80f4597c

Observation 50ca7b15-dfdc-46db-8798-0ddf1a0e0329 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.150799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.150799Z digest=sha256:542ee59a5c7397ffad91cbc59d4d873cd7e193bfa37c53d188025360a7b49cfa

Observation 87b1b2c3-68b9-4d65-afbc-738f1b3f6947 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.238958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.238958Z digest=sha256:302c23e1a5b07fa7ab0f3834911ca90f9c5117413ede4348af985926928f0844

Observation cad49236-1d9d-49f3-9a3f-c38b95ef13a2 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:05.722951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.326879Z digest=sha256:ec699847a4747292447754c8a2d7323f29bae0e9d2758083de5680800236c7ba

Observation 304fed1b-577f-4a56-9216-74a69416e9ca · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.399389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.399389Z digest=sha256:9c231369214302ab8ac5daddf2841ce3d1c5ae4fa2b68c50a000b8e9d53fc7da

Observation 5138c6dd-79fd-47d8-89ba-261db7c90624 · outbound

This paper cites Shindo, V.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Shindo, V

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.485076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.478028Z digest=sha256:76af2ef5485b22da538189de8cbe27f7e61ed78a110a17bdb16e310e1162fb1e

Observation caf37e07-1b82-4459-a02c-4927935021c0 · outbound

This paper cites Kangaroo Mathematics Competition, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Kangaroo Mathematics Competition, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.273555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.585502Z digest=sha256:56e1fcd0e0ad80dae64dd653957b283967927aa6283d163f1a5caf4cd8fd2ac3

Observation cb88a8b8-03b8-4266-8d7e-738e2c7f8f5b · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:05.148045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.678482Z digest=sha256:6ec222e0bf25c4da8a9a391acb38c5e1c7f18f97ba3b0aef808305b250cb73d0

Observation 573b588d-a2f9-42f0-9b3f-f827c66c52d0 · outbound

This paper cites Rhomrasi, Y.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Rhomrasi, Y

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.938401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.785035Z digest=sha256:41dffb8b380a2bdf4fb6b5598ecabbe7cc0181ab541f568e352ccbdf616b8584

Observation d745f0ea-b10e-4459-9681-3fd338604861 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:04.752925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.900519Z digest=sha256:1348211ad5e1be1eb4f2d9644fd60c34a48ef58ba683306529bd60275809b1c1

Observation 2d0e39b0-00e1-4238-ba58-8e2830ac619f · outbound

This paper cites Sachan and E.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Sachan and E

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.548848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.004348Z digest=sha256:76dd6dc6bd9b0d00bcebaa4d46c00715df3c82a376495378b9168241a77eb386

Observation c3b793c7-6a0e-4711-ade6-818b4fb0e3c2 · outbound

This paper cites Alayrac, J.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Alayrac, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.364881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.123777Z digest=sha256:359f1a784402974b3a0c579b2f2b9698ff02a3ad37ab671a13529d77ff32a39c

Observation 66b9ed53-75f3-4785-ad3e-e3e9985d2472 · outbound

This paper cites Antol, A.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Antol, A

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.200993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.228933Z digest=sha256:db63080dff5d3b62494a1b2692dd39fbac655e55ec0e124bd204fcbf0aa6fa55

Observation 1b2bac11-e557-41de-90d9-0663e5af949e · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.315986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.315986Z digest=sha256:647a235698f1e8c705552f8bf7f1db5d6ec4d6fe6a4bbde7f2316fd0389bf98c

Observation c7c7c834-d8c3-4486-a128-551a39b058d1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.407354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.407354Z digest=sha256:9a496251cda37f273d30cbcb15d1f8ac50d619484efb5db0409719ae793c1397

Observation 95d62a69-1cec-46fb-91d8-974add5e83dc · outbound

This paper cites Zhang, D.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Zhang, D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.016375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.556096Z digest=sha256:39c15e8e2ae0bb8d592b197efbe0fb4ff8bfc70c1ace48561c47b97674332a61

Observation 97d54cd4-ff90-40b7-93e9-00ce9a5c2f64 · outbound

This paper cites NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.647891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.647891Z digest=sha256:b78015e970670df2bf43482e79b3ce5f62157345c8d22b96dad1e6d772816658

Observation af557481-fa63-43f1-bd55-bc3fbec83de0 · outbound

This paper cites Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.780594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.780594Z digest=sha256:e4f61b8bda936d56e4deec7724facd420c4291dfafdfb9106fea003a8d299028

Observation 5f2f6cf7-204e-4d06-8cb0-e32ee99c223a · outbound

This paper cites Goyal, T.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Goyal, T

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.831158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.862704Z digest=sha256:530dfa12a0ece3e2a9e25d4141fe6d3d6261dd98c315c4b84d6dbca4b97dd1ae

Observation 453f1ba5-a98e-4c99-b932-26767347630a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.990977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.990977Z digest=sha256:59c9089aa96ae00899ed6c886c7ff3b0977314d5e01d27914e8391997d7a16cb

Observation ee210642-10fa-4f89-8037-d4ece739fa93 · outbound

This paper cites Saikh, T.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Saikh, T

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.637710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.074449Z digest=sha256:faebcc4dfcb90193eccceb2a7845a982150f45d02a00a27439a977f1d5d8f3a4

Observation 0c87dee2-7f1b-42df-a252-7f049e8e7e4f · outbound

This paper cites Caffagni, F.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Caffagni, F

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.476997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.168749Z digest=sha256:c4b3277cac3311f49871c1b5b634a2d3c42fc299de9cf656a56b7869af115705

Observation 896556e9-dafe-4f9a-a901-4e5005ad5654 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:03.348832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.258951Z digest=sha256:92adf8431ad4d07654890d6c5ee8b569421e1a36887432431ede70707ad0541f

Observation 5f9d92a7-117f-4199-b4b3-dfe6dd2d72e9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.377109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.377109Z digest=sha256:27f6868bfd9fd66ca666896c96a5d9d3aae265cc69de140479bef870a1f0e711

Observation 83ba023d-cbfb-4439-bba7-77341301036e · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.464859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.464859Z digest=sha256:02808b7d026b6bdcaa6339afba72f43463afe6bb81811673085dc8aa6c3c805b

Observation 0c8dc9d2-db98-46ab-b742-339bc071c773 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.546002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.546002Z digest=sha256:a3228beb5adfb7084eddf087766cec4362cfb0f38746cf1824b9d2c595ef163d

Observation 690bfc09-49da-4b96-893b-3cf4e768ce37 · outbound

This paper cites Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.227171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.623333Z digest=sha256:9271a4eac9f70e6aa03dbb108ff6726c7e4c0146123bff15a1baa97e38fce595

Observation 038865bd-3558-44ee-a4bc-43f4220e087e · outbound

This paper cites Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.034380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.711872Z digest=sha256:ee773620c9556d831beae9a31692b6742e346fd9aac0278684f8b2e433c16a36

Observation 6f62abc5-0f3f-46ef-a12e-76aeb0052716 · outbound

This paper cites Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.886641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.815322Z digest=sha256:7c6cd99bef48c24c2e3b82938dcb2ae676f95007c6a046191246e81bb4be9a69

Observation 3497ccd6-ef16-4153-856e-188d3b84e651 · outbound

This paper cites Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.746861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.925316Z digest=sha256:84d1f6b5488941469526d5be23e81e304503dddb39c95ba530da50395d00ca7f

Observation ab2100e2-8416-4dd2-b4f6-4323ec5a9866 · outbound

This paper cites Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.570734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.023172Z digest=sha256:253ad59a342c7c0a59260f26faa695b4ed579416b920f7e3512e35ffeba4c064

Observation 6a14e7c4-98a9-4cf5-8489-7a1132ca4f4b · outbound

This paper cites Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.409137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.113148Z digest=sha256:a3cff8a1b64d3823bf760345840f91e04c08d41cf71847a899fada078306bc9c

Observation 8d7fae9c-9770-4187-81e9-32ad6a13f7b4 · outbound

This paper cites Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.243023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.238036Z digest=sha256:adcbacc83052279a6a0b2bd97d8a1c9535ad9ca608a744f74ab7f2be7e363441

Observation f00e5a31-7e5b-41ab-855f-3f97e42af644 · outbound

This paper cites Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.121010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.305158Z digest=sha256:d813c7375881be1a52fae9b9c9248d6cb33523381c3c7ff87c65e7267d9584b6

Observation 1b0fbfde-3c37-4e78-9479-eec4f713b624 · outbound

This paper cites Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.924226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.397520Z digest=sha256:09681e992623b85a28dbd24e17078a2ae29495a4cb96fad94b8250712889ab15

Observation d9315d22-26a3-40be-84ad-ba9dff0c6d1d · outbound

This paper cites Kangaroo mathematics competition, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Kangaroo mathematics competition, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.732821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.468727Z digest=sha256:4fbe97fd969a178e5bdf59a63cc3277a91a36cd6a1157470d078723deee38e36

Observation cd56db18-9f26-4c6c-8acb-231b9d04268d · outbound

This paper cites Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.559522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.567258Z digest=sha256:7461abfc60cbc9443c9087466f6c9c7f48e5ff160e43ea9bc491b1191263bd96

Observation 5cd0efb3-23a0-45a0-8ebd-0e81152075a3 · outbound

This paper cites Concours Kangourou de Math´ ematiques, 2025.https://www.aksf.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concours Kangourou de Math´ ematiques, 2025.https://www.aksf

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.377765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.656507Z digest=sha256:bcd4c4673ac6579b13635954d4b318b7b9a035267d6e68ef29aab358df5316a2

Observation dbb4fb08-ac9f-4904-9217-c03ab09a386f · outbound

This paper cites Concurs Cangur de Matem` atiques, 2025.https://scm.iec.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concurs Cangur de Matem` atiques, 2025.https://scm.iec

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.164441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.747897Z digest=sha256:173ea067dd88ea809911b9d38f776602166cc9767d8cb2c210f84e74c8724ac4

Pith citing papers

No inbound Pith citation observations are available.