Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:36:35.996643Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2506.19468.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:36:35.996643Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:17:10.649943Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
40 of 40 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 98b03044-3a20-4beb-b15a-d62e21f70207 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://www.anthropic.com/news/claude-3-7- sonnet
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d06bff1-adb1-4c2a-b536-8652fb2e3d64 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://openai.com/index/hello-gpt-4o/
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e170ba35-2370-4d50-a6e5-7ee71e86fa26 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://docs.anthropic.com/en/docs/build-with-claude/multilingual- support
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 43d641cb-caf4-4ffe-a21d-ee96087d40b8 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://qwenlm.github.io/blog/qwen3/
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f2f3377b-0996-43bf-9a82-71b2c8a32d5b · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of- redpajama
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2b33e47f-0555-4ddd-b267-a19c28153346 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa721713-73e3-40e9-8e7f-58e2af6ec58a · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec99d56b-849a-4789-8fa6-08cc4a8293a5 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bd3a764-b1b7-4488-bc5e-ca2cec99fe8d · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models, March 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b85ca9f9-e540-43a0-8c2d-7041d9e25d95 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, March 2018
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a29e9fc7-b056-4f24-b63f-28f04e968567 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Emerging Cross-lingual Structure in Pretrained Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f04a8942-0535-47b0-aa16-f701872a2cd8 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ab1e1922-a584-4e3d-bd47-2d85e3875240 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Measuring Massive Multitask Language Understanding, January 2021
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e9b682f-6662-4a3a-88f3-bc2b4059c93e · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models, February 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 246c5349-1ce6-4dd4-874f-5a5cbe828057 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Evaluating Code- Switching Translation with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fe10d072-5651-429b-9f32-826fae534430 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 99bf655d-ca07-4eac-9bd6-792c4bd608be · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 23f05cb1-4611-48e0-aab3-35eaf884a8f7 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Cross-lingual Language Model Pretraining, January 2019
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 25ef6dd0-ac2e-42af-8422-ef5f55bfed11 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages The Winograd Schema Challenge
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c65a92f0-4b1e-4118-83e8-6240f5cc6671 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages CMMLU: Measuring massive multitask language understanding in Chinese, January 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 874ebfe6-2074-419e-af6b-165b3c29fbca · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 548ab402-ab86-40fe-b10d-3515435425f6 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Few-shot Learning with Multilingual Language Models, November 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c13d83f-24ab-4403-b131-a2a595cd01d1 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories, April 2016
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation de2a2721-bf5d-42af-8d45-bcf48bca2370 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, October 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14ab9f1e-bf65-4baf-88ce-68e172d531e5 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 741cad5f-2805-429d-bf59-a6dbd667b067 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Qwen2.5 Technical Report, January 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0cfd18b6-c0a0-4d3b-8a3e-e94a2ce4cd38 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Farinha, and Alon Lavie
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 069a478c-ba85-46ba-8347-31de6d7836c3 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f09d0d-5100-400d-9409-b8029b4a278b · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a20afbec-4b23-4406-a544-717165232b95 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837d2353-93f2-4ac1-b566-a7bcce497b62 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f149561-8e70-44f8-b7e7-7377b1b9afe6 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemini: A Family of Highly Capable Multimodal Models, June 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ed20dbda-5be2-435e-a163-eb61afd3a5d1 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemma 2: Improving Open Language Models at a Practical Size, July 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ac3a7651-b3d9-4f04-92a1-ff7ae8ee86f8 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemma 3 Technical Report, March 2025
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a2718dd-310a-4da3-9463-c30a86374681 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, November 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 01178922-80db-4fee-a5f3-525d8e439acf · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f531ad6d-2fd4-4cfd-9853-d90bb11031b2 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, March 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a86a4da5-b507-4506-8c1b-6af4e12c5808 · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages GeoM- LAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e89d369c-9f7d-466d-8b69-4a16717cde6b · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d708062e-4b90-4a7e-a6c4-3de2a5f2c57e · outbound
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages imitative falsehoods,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d7d670b6-74af-4249-82d0-ad4de16c1c5c · inbound
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 730fd3d8-a5e0-4074-bf2d-508a7376e0bb · inbound
Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c82c7c9-5200-4e50-86b9-ee767e5aaa4f · inbound
Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7e5a3f42-0bc6-4992-9fde-7ea61cbf2cbc · inbound
MultiHashFormer: Hash-based Generative Language Models MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9720bdf5-582d-4673-9979-f8b6660a096c · inbound
Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.