Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:06:59.670819Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.13498.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:06:59.670819Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e353f138-250f-4b03-9e29-db2994e5ab16 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Evaluating large vision-and-language models on children’s mathematical olympiads,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3f767a34-2456-4a86-b61b-ddd47f7b6b36 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Towards reasoning in large language models: A survey,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 98326dee-95f0-4cfa-af71-1f5066f24ff4 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The bitter lesson learned from 2,000+ multilingual benchmarks,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b418d6ba-a19e-419f-8bd9-4012ac3c1c4e · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Challenging the boundaries of reasoning: An olympiad-level math benchmark for large language models,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9cc1d127-af62-41ed-bdd0-8bb2927c4619 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Humanity’s last exam,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c8500227-a87f-4036-9a47-50806b63f63d · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The multilingual mind : A survey of multilingual reasoning in language models,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9b1decdb-fb12-4d6c-bf1f-63298839868b · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Scaling test-time compute for low-resource languages: Multilingual reasoning in llms,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7a05bebb-d546-4099-9fa7-09d67df29cd9 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Towards measuring and modeling “culture
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 65badd25-f43d-44ce-b1d1-1282e6b31169 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Memory of Peoples Series, UNESCO, 3 ed., Feb
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3e918326-2133-4940-9df9-6cd4de5ec00a · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation SIB- 200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ff32c060-5313-4954-af16-1751887b485f · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The belebele benchmark: a parallel reading comprehension dataset in 122 language variants,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb4f1d79-23a6-4346-84b6-706864247b03 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Uccix: Irish-excellence large language model,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 94372199-e63e-4ac2-a3e2-1908b4e85d68 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Survey of cultural awareness in language models: Text and beyond,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b2daa91-c408-4fd6-9906-67f924b15966 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Judging LLM-as-a-judge with MT-bench and chatbot arena,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d562f72e-d14c-49d4-9065-e1613f418b99 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Justice or prejudice? quantifying biases in LLM-as-a-judge,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 676b34f5-4ec0-4eb7-96d4-3a333addebfa · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Language models are multilingual chain-of-thought reasoners,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb9bcbe7-0ba7-4f67-9fdd-0ae6e6f326f0 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eca2b40d-4bdf-45c5-a489-f3d63625147f · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9702f91d-7bcc-4094-bdfc-b89415656684 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Training Verifiers to Solve Math Word Problems
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b938eb7b-4fa9-49b8-85fc-9b9ac3ae1ed1 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Measuring massive multitask language understanding,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 74846c90-94cf-4c22-a4ad-ccf8b9018d09 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation CMMLU: Measuring massive multitask language understanding in Chinese,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2ed1a92a-772a-4ea1-9279-315c66d6fad8 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation KMMLU: Measuring massive multitask language understanding in Korean,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e80bcadc-f996-4193-8989-e01a5157dcca · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation ArabicMMLU: Assessing massive multitask language understanding in Arabic,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c98c950c-70c1-479d-a39e-e9003b5ae395 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation None of the others: a general technique to distinguish reasoning from memorization in multiple-choice llm evaluation benchmarks,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ebe5e9b-80b4-4daa-8e32-3a5ab1817e8a · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation SPIQA: A dataset for multimodal ques- tion answering on scientific papers,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a24f8483-3c3b-40c5-b657-688e93adfc29 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Introducing gemini 2.0: our new ai model for the agentic era,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4d7baa99-52d9-4371-ae1e-c4f89d245b0c · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Gemini 2.0: Flash, flash-lite and pro,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cc371727-9f41-49ed-a630-d8f3116adfcb · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Bleu: a method for automatic evaluation of machine translation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0d321d6a-e496-4656-b662-0f001065d8ff · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation ROUGE: A package for automatic evaluation of summaries,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22831e4-866f-41ec-9ef3-4e51fefe0a03 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation LIMA: Less is more for alignment,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2e9c4103-e71e-4883-a8a9-a4efed9b5f14 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Bag of Tricks for Efficient Text Classification
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0fc659-4e43-45cf-8ddb-97d2566cdb01 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation FastText.zip: Compressing text classification models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568b6873-8f3f-49f3-a5d3-90ef9d205ed8 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Introducing openai o3 and o4-mini,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9052b625-9441-44f0-80b0-3dd0ed6fdc19 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Introducing gpt-4.1 in the api,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 820c2407-65b8-493a-bbf0-4fd9e43c1d6d · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation The llama 4 herd: The beginning of a new era of natively multimodal ai innovation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3abb851c-3f92-4017-a412-687421e7c253 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Aya vision: Advancing the frontier of multilingual multimodality,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9aa5105c-3e93-495d-bb01-9b7076063a37 · outbound
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation Measuring short-form factuality in large language models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.