Pith. sign in

Paper Citation Record · LEDGER

Disentangling Language and Culture for Evaluating Multilingual Large Language Models

As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.24635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24635 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:33.032518Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b584db00-d352-45d2-8f3e-d42e8cb31f84 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.310743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.310743Z digest=sha256:88c8cb006004307a9508b6304a9ea8a2de347df20138ada724fb48298e45bb28

Observation fe36fe63-bf95-4b02-aba1-fe168ef7da96 · outbound

This paper cites Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.383031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.383031Z digest=sha256:c2f717f98a23ac197743af3bdcf1d8b0c73b411da14ae4bd961e2987fef1683c

Observation 545853b1-e495-4d69-a88b-671cfb2c1227 · outbound

This paper cites Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.450813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.450813Z digest=sha256:0b212d88b8c7ade254f1c6cff4e37f449b54c9ef9c2aba5bb4d4c5557e93d53f

Observation d6e1a612-e347-49a3-b5e8-66b55159d98f · outbound

This paper cites CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.577676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.577676Z digest=sha256:65605fb3f6e97a47a4faea8a2aea35c3f721319b4553d50c341f2c6af336e528

Observation 0f093d86-4add-44d5-b7bc-6130c22afe9c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.670996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.670996Z digest=sha256:47ecc4c165b956b7168e58a0c6d6d434270e840eb1be599dfbc81b174e9150b0

Observation e5a57296-a27f-417e-84a0-0d78ab4a73cf · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:34.318376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:30.786039Z digest=sha256:f08b751c9a03e441dd49379a10ebc1d4737ee1348146e9da05faedd5192c1bee

Observation ab413b52-840a-4d12-bef7-4b465ede556a · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.868536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.868536Z digest=sha256:ad9fbd34a2b2e7972884410930ab63d0202c1538d967d8e791267ad6a9aa8d7e

Observation 97426cd6-c677-4952-b9da-1efedfbd4f00 · outbound

This paper cites The Llama 3 Herd of Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:30.975371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:30.975371Z digest=sha256:9bf1804d6fdf730e7acdb23178d7c46b405cacd8b96328502c18c958f7e2b5a6

Observation 5cc99fa0-6ce5-4cf3-bf32-eaa45ada50ad · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.063098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.063098Z digest=sha256:4b9fcf00e7e9b22bee954b92c300c10429fbdaec66c65f391abf2686c46a2a6e

Observation 4c34f7a2-6192-4306-b09f-bd0a425a6909 · outbound

This paper cites Intrinsic Test of Unlearning Using Parametric Knowledge Traces.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Intrinsic Test of Unlearning Using Parametric Knowledge Traces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.131951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.131951Z digest=sha256:6c735ebfa0b99615169297fc86577e7b350ea7c370eef950d96fac04337ba003

Observation a4467368-c0c4-45b9-9801-02a593adf7b1 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:34.087343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.215879Z digest=sha256:fbc2bca12591d8f8fdbe692ae4a89b343423d7dc6922beaf017ae13d4f2738b0

Observation a11b7aff-1b05-483f-b640-8515b7c6c8e7 · outbound

This paper cites BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.299955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.299955Z digest=sha256:4de0cd68026b3b48149cc334c30ba6173ae714f1e7c7b0d118470c2041c1b45e

Observation db639b68-38c1-4584-87dd-41e5125fb336 · outbound

This paper cites SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.361951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.361951Z digest=sha256:504d7141b88aa1ae56b976be218af825f5e301633be49eb3b503db9890421511

Observation 417313c1-cc50-42f6-b690-1c93f4b9ea54 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Crosslingual Generalization through Multitask Finetuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.446833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.446833Z digest=sha256:a00c800aae5cfc7b5148455aba14506ae2f64d0eb5331848f6df800d5c264fa1

Observation 7a09cff2-1343-4168-b394-1863624cde60 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.937127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.544759Z digest=sha256:055a45f893969263174b4cd99cd1380b535af9d039556e162122b6778aa2fc41

Observation b6627754-7a28-409a-a819-fe97787bc1b0 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.621695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.621695Z digest=sha256:9f90006e21bf046b9548cc753c57629fa56aac0ba56e6a9075fe281d74ac1bf7

Observation 5e829491-d525-4fe6-9ff4-02dc8b388965 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.784466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:31.672619Z digest=sha256:0ae20b6e0a7c269af83668e53076bb4521d7c390b4707d94fb47ad07578da4ce

Observation 0d910d80-1073-4ae0-a48c-5777ee2eb970 · outbound

This paper cites GPT-4 Technical Report.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.753236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.753236Z digest=sha256:347807e79ac5daa158087a637426c96cfd1d9243224c3d32ab7b91527e9039e1

Observation 274dd890-6f2b-4fe5-abfa-88d2fc3bf5c4 · outbound

This paper cites Qwen2.5 Technical Report.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Qwen2.5 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.893975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.893975Z digest=sha256:289349e09a10397fd5239d7c5960bc8bbdaa4486e97a565e2967214911c1ad92

Observation 1740e2e0-9e3d-4055-8641-d8456358748a · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Language Models are Multilingual Chain-of-Thought Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.015576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.015576Z digest=sha256:b39df08f2ea6d1bdce16d8e9f48bd1971df3dcc13ecdefbd9a7307ee3b90c34c

Observation 3e978c99-7ad8-4c14-b83a-5dc943c536eb · outbound

This paper cites Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.258791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.258791Z digest=sha256:96707f9e8766e4cc8ec4ec6cbcab708f19a7127aa1f7a37490225ab88f9835fe

Observation 9e974ea5-6772-40b4-af04-873cf1c17ffc · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.331555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.331555Z digest=sha256:f726af21d84c58019ca422fd9868061715ad4c4e9de58f135b3fd081709e48ac

Observation 248d82b0-cc93-41e4-b373-26ab7dfe7199 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.597931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:32.434992Z digest=sha256:08c567a90fb2fed4e48af0f5ca53aefed12b72d4cb1720854ba17a24106c1f4e

Observation 34fba750-ab18-4cf6-8e74-9eccf946f6b6 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.530994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.530994Z digest=sha256:3a6cc666f48d57f683eb0ea805a10f1b77b9caaded5410696a7b44ee167dbb67

Observation 2d4d9e63-387f-4f55-9ef2-9ed1152d6fc5 · outbound

This paper cites Do Llamas Work in English? On the Latent Language of Multilingual Transformers.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Do Llamas Work in English? On the Latent Language of Multilingual Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.626946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.626946Z digest=sha256:15f4166cae23ac118ab747857ea3de9744604aa045a1074a7dd76513e58ee2ab

Observation eb7bcc25-5d00-407a-95a0-8b0f2b563f27 · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.733838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.733838Z digest=sha256:839fb58995561e9255c304d566161b9cd93b5af1666e2b6d2fbd575742c3126c

Observation 5dd1b83d-7f22-4f67-b534-26abce526a3c · outbound

This paper cites an unresolved cited work.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:33.452076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:35:32.804553Z digest=sha256:0888a0225c8af095468885e1bd3f88bbe5324144276bb53814100eb6828e3bff

Observation 6e5cdfa0-728f-4af4-bc80-471c884a1b90 · outbound

This paper cites SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.877616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.877616Z digest=sha256:0df529457e50f2fe3a8323799a366a28abf235efc818aec6d248605f097e9528

Observation 11197da8-c143-4207-adcd-a3ab214724d4 · outbound

This paper cites How do Large Language Models Handle Multilingualism?.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models How do Large Language Models Handle Multilingualism?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:32.951876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:32.951876Z digest=sha256:e405e396c91ef2e05f998425f2a336792ab8d14671c3af9def2760e9b9d982ea

Observation ae484c60-081a-429b-bece-63d9e534e71d · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:33.032518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:33.032518Z digest=sha256:343e73f42fa9570f98bdd462b31b9befb2d2ef00d8a46c34bd467f1c2374fb81

Pith citing papers

No inbound Pith citation observations are available.