Pith. sign in

Paper Citation Record · LEDGER

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests

As of 8 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2506.07418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07418 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:39:00.747897Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a484e773-47e5-42ae-b5d5-8fb790abd933 · outbound

This paper cites Zhang, Y.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Zhang, Y

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.503087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.536458Z digest=sha256:3f23e3b0c71f4bb37e278395193fb12cfd2523f7f57820b9d8fbe415ab42304d

Observation 3aae2431-7507-489e-bea7-b08a9baa8cb8 · outbound

This paper cites Achiam, S.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Achiam, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.313361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.650734Z digest=sha256:7b060e22d4e7b2839ee0596b28cb390e074fa9739d393fe9d52df8b551e6b6c7

Observation fc18f617-188d-4cea-a35f-ca373b0d3f75 · outbound

This paper cites Qwen2.5-vl 7b.GitHub Repository, 2024.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen2.5-vl 7b.GitHub Repository, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:06.069406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.737402Z digest=sha256:3aa733a4f8243ce6dd875f0918b5e353fc8bc708b3a86022865752c6d903dffe

Observation 42929a52-a881-4b7b-bc17-e02c0c02a3e4 · outbound

This paper cites Visual Instruction Tuning.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:56.825172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:56.825172Z digest=sha256:7ebae5677ff89d4d761e37227246bf0d1bce1c73ebd571d7533ab928cdca1f11

Observation 2630dfda-3d82-47c3-8b56-eef2f69ea43e · outbound

This paper cites Andreescu and R.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Andreescu and R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.898084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:56.904032Z digest=sha256:ed133d893a7b1cb78d68aba2682e93f4108acc09a1670804083d7082f68d264b

Observation 93700d1e-7cf0-4a5d-8199-fbe04144cfe1 · outbound

This paper cites A Survey on Benchmarks of Multimodal Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests A Survey on Benchmarks of Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.018980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.018980Z digest=sha256:a01a057e42687aec8e2f5b716b2bd2935deb7f649b2e0695788dcd02ed0b3fd8

Observation 50ca7b15-dfdc-46db-8798-0ddf1a0e0329 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.150799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.150799Z digest=sha256:b88eadd2053b242ab63534368919f0fefcd0e6308a43f9aa593a531586183359

Observation 87b1b2c3-68b9-4d65-afbc-738f1b3f6947 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.238958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.238958Z digest=sha256:4b2275277de415d8f9ccd75d810b95ec3521d80537292aac12e2272851acfd26

Observation cad49236-1d9d-49f3-9a3f-c38b95ef13a2 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:05.722951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.326879Z digest=sha256:1de9e431a8771d803159c6dd5842d52e0829d7a5e279652de9a2d93bec728e02

Observation 304fed1b-577f-4a56-9216-74a69416e9ca · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.399389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.399389Z digest=sha256:bd79e41f58b794721e3a493a1189ef3607a2dab3dfd44baea3adc9a86a9ff217

Observation 5138c6dd-79fd-47d8-89ba-261db7c90624 · outbound

This paper cites Shindo, V.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Shindo, V

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.485076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.478028Z digest=sha256:2c90ca0fde5433bec7e2d6c2e7cc416fdceb3e72944b3c8b0b62d9cca2131285

Observation caf37e07-1b82-4459-a02c-4927935021c0 · outbound

This paper cites Kangaroo Mathematics Competition, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Kangaroo Mathematics Competition, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:05.273555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.585502Z digest=sha256:315c8c6141a854ddd970e7ffcd6d417655c2c23a99dde360d67d14e58c276fd7

Observation cb88a8b8-03b8-4266-8d7e-738e2c7f8f5b · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:05.148045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.678482Z digest=sha256:158bb79a60261e82ad61f671b13d5523e7f7b4fe536d5b12183c78c9a4d5280c

Observation 573b588d-a2f9-42f0-9b3f-f827c66c52d0 · outbound

This paper cites Rhomrasi, Y.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Rhomrasi, Y

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.938401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.785035Z digest=sha256:f648768fe61d3ae132a02f08e0a0a1bc026e986ab334dac68a111a8c3427659f

Observation d745f0ea-b10e-4459-9681-3fd338604861 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:04.752925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:57.900519Z digest=sha256:c3e45a4feb607f9f674be1a8f7ff4fa137a9336bc48a4a6b38d7f9dd7b4519b1

Observation 2d0e39b0-00e1-4238-ba58-8e2830ac619f · outbound

This paper cites Sachan and E.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Sachan and E

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.548848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.004348Z digest=sha256:2c2a05ac357368b03493468ecb3c517ce59700e44f15b114880488ad5fdd2c2b

Observation c3b793c7-6a0e-4711-ade6-818b4fb0e3c2 · outbound

This paper cites Alayrac, J.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Alayrac, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.364881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.123777Z digest=sha256:537739ef7015a19f3e1777d092253264ae97489bc4f80d72a417455a6336f125

Observation 66b9ed53-75f3-4785-ad3e-e3e9985d2472 · outbound

This paper cites Antol, A.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Antol, A

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.200993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.228933Z digest=sha256:364666d340afbc6ef6752de8162535eb8778a130f34c37c731c7fb1a3523de2d

Observation 1b2bac11-e557-41de-90d9-0663e5af949e · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.315986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.315986Z digest=sha256:325ff824ee8eee529992af4d9ddabaf610802f4527ab6d6c891dfbf58d1b788a

Observation c7c7c834-d8c3-4486-a128-551a39b058d1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.407354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.407354Z digest=sha256:b658d9f331d073d1cd8664cb6001f133ca0f9822d04ee3ae759072a8057360af

Observation 95d62a69-1cec-46fb-91d8-974add5e83dc · outbound

This paper cites Zhang, D.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Zhang, D

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:04.016375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.556096Z digest=sha256:4ae52271961f7c683589ba4150caaf4aebd29953a3383185348caa8a7444cd6c

Observation 97d54cd4-ff90-40b7-93e9-00ce9a5c2f64 · outbound

This paper cites NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.647891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.647891Z digest=sha256:9dc80e8870a2773594115341701cb23a7ff66629fdbc54d96090f6246a6ed706

Observation af557481-fa63-43f1-bd55-bc3fbec83de0 · outbound

This paper cites Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.780594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.780594Z digest=sha256:7e01a6f3e9b9d36dcc833d6c1d1530d30adc6f29bf607051c74c59ea5aa5644b

Observation 5f2f6cf7-204e-4d06-8cb0-e32ee99c223a · outbound

This paper cites Goyal, T.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Goyal, T

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.831158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:58.862704Z digest=sha256:76d8c81567603be2b9f1f14bd043f1c60771917ced56846455bef3926668fe14

Observation 453f1ba5-a98e-4c99-b932-26767347630a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:58.990977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:58.990977Z digest=sha256:6d9d0b38440e70ba82028b2df39c479e18f33a969b7583dc96733c0231fcac08

Observation ee210642-10fa-4f89-8037-d4ece739fa93 · outbound

This paper cites Saikh, T.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Saikh, T

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.637710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.074449Z digest=sha256:774bac676aadf437420dcfdfc180f762ef552dc5f2aa687f32b3c983f7d95ede

Observation 0c87dee2-7f1b-42df-a252-7f049e8e7e4f · outbound

This paper cites Caffagni, F.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Caffagni, F

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.476997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.168749Z digest=sha256:9cc9174eaa495c71d089a82b15fb24b3b6601fb4eb12b81129fc72a8ed645f38

Observation 896556e9-dafe-4f9a-a901-4e5005ad5654 · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:39:03.348832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.258951Z digest=sha256:099756abdb928d8c786a9e36f277f2bcfcf892377be6e57acec4d4132c9cd7d6

Observation 5f9d92a7-117f-4199-b4b3-dfe6dd2d72e9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.377109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.377109Z digest=sha256:f58a62d848f76f41468f2638ea5443b6751ce8f799f537a59efe9f2842ba7764

Observation 83ba023d-cbfb-4439-bba7-77341301036e · outbound

This paper cites an unresolved cited work.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.464859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.464859Z digest=sha256:e9f73abf2574f1515dd5856d1c4a7a951bf4555e36a2251e30db24970d9974e5

Observation 0c8dc9d2-db98-46ab-b742-339bc071c773 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:59.546002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:59.546002Z digest=sha256:24e26cb37b37f79065b2d0ac767a1d079c1af16d1ebf5f47230058813407f977

Observation 690bfc09-49da-4b96-893b-3cf4e768ce37 · outbound

This paper cites Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Pixtral 12b.Mistral AI News, 2024.https://mistral.ai/news/pixtral-12bLast retrieved, April 25th, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.227171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.623333Z digest=sha256:808914938749a5146b2ecb9813bd2e11f7fff92f5ac453d67764236b5b492ea2

Observation 038865bd-3558-44ee-a4bc-43f4220e087e · outbound

This paper cites Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Llama 3.2 vision 11b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:03.034380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.711872Z digest=sha256:af608ce50977ea042c549ef115646ee3a14f533af619ac91c969a9830c3f7e9b

Observation 6f62abc5-0f3f-46ef-a12e-76aeb0052716 · outbound

This paper cites Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Llama 3.2 vision 90b.Meta AI Blog, 2024.https://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/Last retrieved, April 25th, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.886641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.815322Z digest=sha256:2ae3b22de15b256da7652795e0fef77cc64470e3da130a98afa90e2714cddfb6

Observation 3497ccd6-ef16-4153-856e-188d3b84e651 · outbound

This paper cites Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Pixtral large.Mistral AI News, 2024.https://mistral.ai/news/pixtral-large Last retrieved, April 25th, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.746861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:38:59.925316Z digest=sha256:92bdc2865dbab2d8be077698d0fda4eee4d91e089715edf5c4afed1746bada55

Observation ab2100e2-8416-4dd2-b4f6-4323ec5a9866 · outbound

This paper cites Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gpt-4o.OpenAI Documentation, 2024.https://platform.openai.com/docs/models Last retrieved, April 25th, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.570734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.023172Z digest=sha256:83179e6eeae74fd73a8f866568fec1eebd1048292abcd1d262bb14534c31ae0b

Observation 6a14e7c4-98a9-4cf5-8489-7a1132ca4f4b · outbound

This paper cites Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 2.0 flash.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash/Last retrieved, April 25th, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.409137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.113148Z digest=sha256:bfdd191ef545fdec1331604235f01b518b6370801e29562d0b3b33b0d421c3f1

Observation 8d7fae9c-9770-4187-81e9-32ad6a13f7b4 · outbound

This paper cites Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Gemini 2.0 flash lite.Google AI Blog, 2024.https://blog.google/technology/ai/ gemini-2-0-flash-lite/Last retrieved, April 25th, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.243023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.238036Z digest=sha256:859a0d05cc3cb0da4ec5619565014e307b4c941ed683145aa18d49cb788b5577

Observation f00e5a31-7e5b-41ab-855f-3f97e42af644 · outbound

This paper cites Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Qwen2.5-vl 72b.GitHub Repository, 2024.https://github.com/QwenLM/Qwen2.5-VL Last retrieved, April 25th, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:02.121010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.305158Z digest=sha256:38f553a20b32f7a5f4677f9eb9f69d8ae69383d955ddf5159992dfdeee2f76c3

Observation 1b0fbfde-3c37-4e78-9479-eec4f713b624 · outbound

This paper cites Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Australian Math Trust, 2025.https://www.amt.edu.au/ Last retrieved, April 25th, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.924226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.397520Z digest=sha256:76d5e35cd81a6476cf63b3f350c1ebf39ccb4a48ba0e7bbdc7215f6a8b12d80d

Observation d9315d22-26a3-40be-84ad-ba9dff0c6d1d · outbound

This paper cites Kangaroo mathematics competition, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Kangaroo mathematics competition, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.732821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.468727Z digest=sha256:f6fc56269a148aa1c0136ccb3d1a9f0aab19c0c440da990b4a415e307793e07e

Observation cd56db18-9f26-4c6c-8acb-231b9d04268d · outbound

This paper cites Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.559522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.567258Z digest=sha256:92b6a213fa573bee07c72d90906f04e10c4455ed3438a7dad4614640250adac0

Observation 5cd0efb3-23a0-45a0-8ebd-0e81152075a3 · outbound

This paper cites Concours Kangourou de Math´ ematiques, 2025.https://www.aksf.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concours Kangourou de Math´ ematiques, 2025.https://www.aksf

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.377765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.656507Z digest=sha256:6dfd77c56fdf45716b79d9fbaf6e3fb4857a1bb07c178845e44232ea0432b976

Observation dbb4fb08-ac9f-4904-9217-c03ab09a386f · outbound

This paper cites Concurs Cangur de Matem` atiques, 2025.https://scm.iec.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests Concurs Cangur de Matem` atiques, 2025.https://scm.iec

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:39:01.164441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:39:00.747897Z digest=sha256:620124584cd9bba7c914bd7da24a5b36d6d65d0cb04aca148be045df6daaf2b3

Pith citing papers

No inbound Pith citation observations are available.