Pith. sign in

Paper Citation Record · LEDGER

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2503.10497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10497 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:35.335823Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cfa8bad-7bbd-4edb-9bd0-d17271414661 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:35.335823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:35.335823Z digest=sha256:f785e75e2ceab18277a605f8ccaa4a8f5173c50bead3dc1c7832c0c4a3160813

Observation b4a59685-f744-4586-8b85-b098e6b5a1e0 · inbound

MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs cites this paper.

MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:02.432419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:53:02.432419Z digest=sha256:03e6032b3e1d34363da5fce2878ee107dcf225b1d767455f4db60ae97427c5da

Observation b62f5b2a-4507-44ed-9183-e799adadf83b · inbound

Do LLMs exhibit the same commonsense capabilities across languages? cites this paper.

Do LLMs exhibit the same commonsense capabilities across languages? MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:43:02.470776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:43:02.470776Z digest=sha256:7ace8f02ccfc849d303713d806fbfe1bd612848243be70df57660f5b17fc2f62

Observation 0bf650b9-a6a6-4ba8-b0e1-e216529eed56 · inbound

Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models? cites this paper.

Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models? MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:20:49.205830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T03:16:01.561180Z digest=sha256:ea02873975c3c59357592a6a3d21a62e175d6ae8f0e40c0a3b5565118b6afe8e

Observation bec86f7f-e4c7-42ae-b56a-a0c52930a7e2 · inbound

Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus cites this paper.

Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 13

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T00:30:53.087429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:29:09.281263Z digest=sha256:bbe7ea54c18a4acccd1a484547954778113bd5dd1849089583a53bd916ddad99

Observation 6239d49b-95bd-4115-bc44-72ccd4a0f39e · inbound

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling cites this paper.

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:05.547859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T01:14:16.831333Z digest=sha256:377e07159f4429e7e68d205964c5917e091f6c4dbdbb28c8c8e0675f6edf71b2

Observation 91014f22-4b00-4901-a80a-72d5e40d0df2 · inbound

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning cites this paper.

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:39:47.780540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-15T05:35:32.806871Z digest=sha256:c0a49be1ae856f5c1eb0f94e6788cc4d05dc24a9688e767655b957c9e6bf9ee2

Observation c67d6216-ad29-4182-bc2c-ddc9650365db · inbound

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance cites this paper.

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.383716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-22T06:19:44.377733Z digest=sha256:c924d43a2283a205f7be359e7cc7705b82dd4fe91fcbfed5f789cd410819c189

Observation 80e1b972-e75d-400f-a470-081157ee5f03 · inbound

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking cites this paper.

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.610430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T07:55:17.832799Z digest=sha256:db3c7a19bb8e120f6c67db4a6ff4551fd2734d0521987d5977cb72a8c855d1b4

Observation 99c475e1-693c-43a0-9268-58784579774e · inbound

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language cites this paper.

Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-26T20:29:57.595067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-26T20:27:41.561992Z digest=sha256:401407848dfe7ce988fe686d362422cb52966e4d237e5c2bc6560421b99decd6

Observation a5ab1097-0f80-4c81-9db2-e2c56a476596 · inbound

Disentangling Language Modeling and Boundaries cites this paper.

Disentangling Language Modeling and Boundaries MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:01.243283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:21:01.243283Z digest=sha256:81990f6d5a43dce9d528c4c5c1b9fddc509705f14125bf2cd45d86450f77b498