Pith. sign in

Paper Citation Record · LEDGER

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2506.19468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19468 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:36:35.996643Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:17:10.649943Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98b03044-3a20-4beb-b15a-d62e21f70207 · outbound

This paper cites https://www.anthropic.com/news/claude-3-7- sonnet.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://www.anthropic.com/news/claude-3-7- sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.590000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.798449Z digest=sha256:74d1165ed3c4c633ec179430aae5684f5f251a6fc321a37ad534c00688101898

Observation 1d06bff1-adb1-4c2a-b536-8652fb2e3d64 · outbound

This paper cites https://openai.com/index/hello-gpt-4o/.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://openai.com/index/hello-gpt-4o/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.575647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.805740Z digest=sha256:9694777ddf25ae2b62f9de6d699c93e2eb98efe27818847c3d40631f9910c60d

Observation e170ba35-2370-4d50-a6e5-7ee71e86fa26 · outbound

This paper cites https://docs.anthropic.com/en/docs/build-with-claude/multilingual- support.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://docs.anthropic.com/en/docs/build-with-claude/multilingual- support

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.561102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.810385Z digest=sha256:a4dedafc87dea5cf5262362286890b113c36256df0ce19595d3704fc745d6222

Observation 43d641cb-caf4-4ffe-a21d-ee96087d40b8 · outbound

This paper cites https://qwenlm.github.io/blog/qwen3/.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://qwenlm.github.io/blog/qwen3/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.548739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.815557Z digest=sha256:fe493be14b253a87387fc86296c6d41a3fd593f40ae13cac97891edec881620e

Observation f2f3377b-0996-43bf-9a82-71b2c8a32d5b · outbound

This paper cites https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of- redpajama.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of- redpajama

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.537022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.821291Z digest=sha256:d5e867fbd0ffb058bce241a6b6f76d229ebe54f442b424b677e5c1709f3f46b0

Observation 2b33e47f-0555-4ddd-b267-a19c28153346 · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.524255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.826438Z digest=sha256:4ec60fb91d680ced503f86e20937ee7f5aa59b19e0b07319a86a3e154bae9549

Observation fa721713-73e3-40e9-8e7f-58e2af6ec58a · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.508957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.831683Z digest=sha256:02cdffeb81025d68615c515f9ac2a47c610741bcac36323f479356175ad77655

Observation ec99d56b-849a-4789-8fa6-08cc4a8293a5 · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.836435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.836435Z digest=sha256:e29c00f2392a69188f2510a6a237189f8a4b9c8fa5538704e19f027283d928cc

Observation 9bd3a764-b1b7-4488-bc5e-ca2cec99fe8d · outbound

This paper cites Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models, March 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models, March 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.482994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.840727Z digest=sha256:b5f486ddee43d46d1a5164fc29a2a8b39ce460a8f39e1c240ff06702227c9654

Observation b85ca9f9-e540-43a0-8c2d-7041d9e25d95 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, March 2018.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, March 2018

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.469955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.846889Z digest=sha256:f22a899f5b7dd634c5c614bc506ce28205c626e3e851fd8c5cc134bb1b0dfcd8

Observation a29e9fc7-b056-4f24-b63f-28f04e968567 · outbound

This paper cites Emerging Cross-lingual Structure in Pretrained Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Emerging Cross-lingual Structure in Pretrained Language Models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.456891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.852194Z digest=sha256:3e7745b55220ed29444495e1c5980f44f19a811a44671c908403a5393025e41e

Observation f04a8942-0535-47b0-aa16-f701872a2cd8 · outbound

This paper cites Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.442754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.857111Z digest=sha256:d4291feaf1cd7b51b62b71f07ead5f439e1d7ca2939e8197279d9502ffde1ad7

Observation ab1e1922-a584-4e3d-bd47-2d85e3875240 · outbound

This paper cites Measuring Massive Multitask Language Understanding, January 2021.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Measuring Massive Multitask Language Understanding, January 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.861228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.861228Z digest=sha256:ffe27079beb0de873d04a01c427cc82e5912fbe49526ddcc6cb989fbbf42b961

Observation 7e9b682f-6662-4a3a-88f3-bc2b4059c93e · outbound

This paper cites BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models, February 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models, February 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.419675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.865204Z digest=sha256:16f42c916b50452cf381524542c34556559e4028a844524b02278ab4699c850a

Observation 246c5349-1ce6-4dd4-874f-5a5cbe828057 · outbound

This paper cites Evaluating Code- Switching Translation with Large Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Evaluating Code- Switching Translation with Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.404649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.869258Z digest=sha256:c656e6b0278b0941c169633d56b7dd60fba01621b0ac55cda0ba64ef1d0b9bb1

Observation fe10d072-5651-429b-9f32-826fae534430 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.389923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.875617Z digest=sha256:404235c4432add06bb9fc1612c9e6f1235b367a5b4c1e7c5daf6fba7d86b746c

Observation 99bf655d-ca07-4eac-9bd6-792c4bd608be · outbound

This paper cites Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.377679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.880406Z digest=sha256:53e69e620b3b4001ec3b6ee362ef604e86b85f70fae32b7ca2c86de8c072cbe2

Observation 23f05cb1-4611-48e0-aab3-35eaf884a8f7 · outbound

This paper cites Cross-lingual Language Model Pretraining, January 2019.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Cross-lingual Language Model Pretraining, January 2019

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.365539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.884827Z digest=sha256:c87a1f25626aa977be286422fc2a4c65d5a88a52b4ca58bbde7d053908aa4542

Observation 25ef6dd0-ac2e-42af-8422-ef5f55bfed11 · outbound

This paper cites The Winograd Schema Challenge.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages The Winograd Schema Challenge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.352897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.890993Z digest=sha256:c42847cd4158b9f909c65dbd554c90417f5f2aa29fe19701094cabb2ebe4677c

Observation c65a92f0-4b1e-4118-83e8-6240f5cc6671 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese, January 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages CMMLU: Measuring massive multitask language understanding in Chinese, January 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.339722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.903785Z digest=sha256:9b9c6d77c1f8b21db25fa4da965d9cc47aa2e366c75a86991ded8d1e3917c261

Observation 874ebfe6-2074-419e-af6b-165b3c29fbca · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.327265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.910190Z digest=sha256:57fceff8b724d1fe10929b12640caa9ac8f3150b0fe936bce6307dc9a4e82994

Observation 548ab402-ab86-40fe-b10d-3515435425f6 · outbound

This paper cites Few-shot Learning with Multilingual Language Models, November 2022.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Few-shot Learning with Multilingual Language Models, November 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.309882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.915621Z digest=sha256:95ba53908d0d6e1402d76914c0d10b7d54d675522989bc687f3b12057d8b4f66

Observation 6c13d83f-24ab-4403-b131-a2a595cd01d1 · outbound

This paper cites A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories, April 2016.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories, April 2016

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.294228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.921919Z digest=sha256:fb4112adb629ef7aa20b807369d19f2b0f8f6374c4293c6889fd9d2f02a7a17c

Observation de2a2721-bf5d-42af-8d45-bcf48bca2370 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, October 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, October 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.282353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.926421Z digest=sha256:4e2240516d8fa013c01dcb4c10b4b32b999fb90c5edfda73b0c430d6614687e3

Observation 14ab9f1e-bf65-4baf-88ce-68e172d531e5 · outbound

This paper cites Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.271759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.930581Z digest=sha256:a49b08f1e79111330ee2bdfbea6a3ef336aeafa619cfea25e9f26d7781b57de3

Observation 741cad5f-2805-429d-bf59-a6dbd667b067 · outbound

This paper cites Qwen2.5 Technical Report, January 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Qwen2.5 Technical Report, January 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.259326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.934717Z digest=sha256:ad3411f608f73de28582be56ca5835b76013a298863219f8f3743c020c10ecef

Observation 0cfd18b6-c0a0-4d3b-8a3e-e94a2ce4cd38 · outbound

This paper cites Farinha, and Alon Lavie.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Farinha, and Alon Lavie

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.249062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.938458Z digest=sha256:377ce015ad514bda7c171fe6dfa958f13d5fb2ecfcc24fa117a8820815564c4c

Observation 069a478c-ba85-46ba-8347-31de6d7836c3 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.942564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.942564Z digest=sha256:3825f3d9fc72cd6d2b8101e24a97e18dd820cee56688886934e5989df5275170

Observation 14f09d0d-5100-400d-9409-b8029b4a278b · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.238669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.948255Z digest=sha256:a7f58b416494c7d42958c1fe730e419d2398cf64c11fa87d165771f9f9c36f96

Observation a20afbec-4b23-4406-a544-717165232b95 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:35.953037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:36:35.953037Z digest=sha256:44d06d5be2cefd96cd920a4de1817f7e088ea8f05c3d7aa5a2076a3c8dfc7c8e

Observation 837d2353-93f2-4ac1-b566-a7bcce497b62 · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.226925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.959420Z digest=sha256:0bbab3051d9488dceb3c2fcd7c77a9ee9a48b89b4e58d613af6310c43b4be263

Observation 3f149561-8e70-44f8-b7e7-7377b1b9afe6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models, June 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemini: A Family of Highly Capable Multimodal Models, June 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.213558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.963708Z digest=sha256:199af0e4a3892f3f1ab3d5817266716e6fabb4d86769c4b5b6140ea19e785c14

Observation ed20dbda-5be2-435e-a163-eb61afd3a5d1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size, July 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemma 2: Improving Open Language Models at a Practical Size, July 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.197316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.968257Z digest=sha256:15f45e6a20d24407b40ff957793c855346ff7ea592f10ca5b83a46956834e5a4

Observation ac3a7651-b3d9-4f04-92a1-ff7ae8ee86f8 · outbound

This paper cites Gemma 3 Technical Report, March 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Gemma 3 Technical Report, March 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.182932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.972572Z digest=sha256:7d113666b7fc003e5a0331a6d891d80ac8c640264c3dd8e1406fbbcf4d6d987b

Observation 6a2718dd-310a-4da3-9463-c30a86374681 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, November 2024.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark, November 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.167618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.976323Z digest=sha256:5f43b1933d9eb0dbf080514ef4fdae087342f455e0392a8dc9e7e6ce6532aa45

Observation 01178922-80db-4fee-a5f3-525d8e439acf · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.153180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.980383Z digest=sha256:5d8192aad3822bff05e7d51cf75c384ae2b59bec70ad1a3ba51d8345c5caf6d1

Observation f531ad6d-2fd4-4cfd-9853-d90bb11031b2 · outbound

This paper cites MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, March 2025.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, March 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.140870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.984421Z digest=sha256:83c6e758f3ab933238a70682264fded2acf985f55be4aed2b4c72a6659aa58db

Observation a86a4da5-b507-4506-8c1b-6af4e12c5808 · outbound

This paper cites GeoM- LAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages GeoM- LAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:36:36.124851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.988736Z digest=sha256:c2a487cd1ef547f1abea16e2b92e01f6a02f11b88d25cc9bcf0d113ff7651ef3

Observation e89d369c-9f7d-466d-8b69-4a16717cde6b · outbound

This paper cites an unresolved cited work.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:36:36.105733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.992934Z digest=sha256:c5da3adb9f22964a89c6bbbacecbe8fb7c79edcfc8a5fae3cec12f301b9f2231

Observation d708062e-4b90-4a7e-a6c4-3de2a5f2c57e · outbound

This paper cites imitative falsehoods,.

MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages imitative falsehoods,

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:36:36.083545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:36:35.996643Z digest=sha256:6eb2910377ace25d5a5f6ed9f9261129e77caf100661b1c0115615026ed8503c

Pith citing papers

Observation d7d670b6-74af-4249-82d0-ad4de16c1c5c · inbound

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability cites this paper.

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:00:17.818057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T20:59:59.582440Z digest=sha256:aaffda4d0cda763b28462da5e8a0c2b74badfd8d2b6142177874bf359a210671

Observation 730fd3d8-a5e0-4074-bf2d-508a7376e0bb · inbound

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR cites this paper.

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.958889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:04:01.791491Z digest=sha256:0df4392f3a75d42be3b383270efaf900132265e7e0d286b3cd2e6e4bb22df2f4

Observation 9c82c7c9-5200-4e50-86b9-ee767e5aaa4f · inbound

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language cites this paper.

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T20:29:57.635151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T20:27:41.561992Z digest=sha256:58aec8d8eb20f9505c34cee2b1af5ab919927f9d782feada5be4187e943949ff

Observation 7e5a3f42-0bc6-4992-9fde-7ea61cbf2cbc · inbound

MultiHashFormer: Hash-based Generative Language Models cites this paper.

MultiHashFormer: Hash-based Generative Language Models MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:23:05.535623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T04:13:05.082903Z digest=sha256:f35a56f8139b0eaf56d9cb499c38336397ba53e1f4e496aa98f6544109c9ccb4

Observation 9720bdf5-582d-4673-9979-f8b6660a096c · inbound

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ cites this paper.

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:10.649943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:10.649943Z digest=sha256:7ddf253af9350121ab0401bbdafb338cab54cae797d4fe6ecfe63db6f625ae0d