Pith. sign in

Paper Citation Record · LEDGER

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

As of 23 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.06298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06298 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:57:53.734445Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:31.361951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T10:14:36.195938Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fba6bad-c1d8-415b-95bb-b4788f6e0c73 · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Sailor: Open Language Models for South-East Asia

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.653020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.653020Z digest=sha256:8c0394641a582c39f3e55ab4a351d96b0d2f557518813edb3ba1bdcf2730c0a8

Observation d2201c71-6070-48e5-81bd-6d0c5e007e1e · outbound

This paper cites The Llama 3 Herd of Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.656975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.656975Z digest=sha256:6cd2e4c7e2ff83619ec59a7f45716c842c9156582b926ffdff8b5ee45d014114

Observation d06556c5-f0d9-43fc-86bf-36d016ae4cf4 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.660916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.660916Z digest=sha256:d4296afc279a425fb3c9d2492a39f576687a9d1b8f66075eeac04942b0fc8e09

Observation 4d5f4659-5a3c-4e58-9cb3-638d3a91b0a4 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.664955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.664955Z digest=sha256:e4d5053eff4ee5c56f31a05002eb9142a22a01af696e2bb9403fc82381e2559b

Observation 84945c91-ddf1-4e28-95c0-2ee6e97be30a · outbound

This paper cites Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.672568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.672568Z digest=sha256:fe3b8d670614f065a041102129865bf142bf7990861be86be755ea9a40d4ce24

Observation 1d7ddba1-1f0f-4e6c-a15c-c4e34e35f74c · outbound

This paper cites A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia A Survey on Large Language Models with Multilingualism: Recent Advances and New Frontiers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.676634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.676634Z digest=sha256:193a36bab57d789c7bf01a7aa0e3ee748ca40cf7da91146ba34c06be1470393d

Observation 4233678b-572b-407e-8669-689c803917de · outbound

This paper cites Mistral 7B.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.680596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.680596Z digest=sha256:abdb9e038e94749a4d4bb7a312adb9348dfdb0fed91e6a48beef6ea279de0b04

Observation 2be21d67-ecab-4335-9840-9de37aea88f2 · outbound

This paper cites ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.683952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.683952Z digest=sha256:54bdc2c988d9f70687f0e5725e6a33469053cb002a84b988f5b78358a54dc37a

Observation 6b66c004-bba1-4d15-a1f2-a93348aa2b50 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.687493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.687493Z digest=sha256:79badc66123c592cd03e895a3e7514ee6ecff0d819e155ae8d05a15799af2246

Observation 066afe2d-7b7e-4bfc-a3a2-bfe5989660d5 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.691083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.691083Z digest=sha256:ec977881d6a8f66ac2d6cdb7afead475ca4f5521301b9588e5b52a6cd5944d9b

Observation cec73c6e-b2f7-45f6-92da-f506fc3fd98e · outbound

This paper cites Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.694595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.694595Z digest=sha256:f51e64f2971915ad4c34eec5a03f0fde4d3cd72403756e0734f8c00e43f8a7ed

Observation 31e23f35-18db-4354-8f4f-4901da865df8 · outbound

This paper cites SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.698083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.698083Z digest=sha256:4463130207ae3ed7f3ad950e162f81c1128134ab4caf17a755f508610e4be1fa

Observation d4e32805-fc9e-41be-877a-a010c365033d · outbound

This paper cites SeaLLMs -- Large Language Models for Southeast Asia.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaLLMs -- Large Language Models for Southeast Asia

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.701632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.701632Z digest=sha256:91ec9c7c00187c9b4e6f56015a71ecd636dc60bee9e787e846c1e28bfe26130a

Observation f17dd090-94b8-4591-95fe-8829167c72fb · outbound

This paper cites GPT-4 Technical Report.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.705379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.705379Z digest=sha256:861fc5fb86d3f4da9623bf46b706005f501c7be33afb0b74ec4ea38037a3727f

Observation dc4e8bb6-94e1-4ea8-b9ad-52d80770ca42 · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.709076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.709076Z digest=sha256:c6b1d7f75ee7aad0e71d835153a5ed951cce1e7f5fc2fa17e61785e4847efaf4

Observation 71a1e99b-ab90-4460-a9d6-258414d52ecf · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Gemma 2: Improving Open Language Models at a Practical Size

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.716047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.716047Z digest=sha256:38175d6f4a43bcdc043d7f6b6cb23061bca89349441bf3fa519224cf491599d7

Observation 126e4584-5d4f-46d2-b2ab-5b5ecc7e9029 · outbound

This paper cites SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.719725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.719725Z digest=sha256:f3a98d5ff4d2af29bed3bc8f04b98d76d65765f4ee911e168881f1db39ca7ab1

Observation 811a771a-b6f4-4a47-ab5e-0ef39f32182e · outbound

This paper cites Qwen2 Technical Report.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.723180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.723180Z digest=sha256:6a8c5d588051f7e32b7fe1bb97832c5cbabf4abde3552cd3c76c11efabd04e5f

Observation 2dd5f617-5e2c-459c-a56f-1ba1820be171 · outbound

This paper cites M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.726650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.726650Z digest=sha256:1df3cad2b72f4be17237ab1e9a32b1512e0e8440dd3e59e177bdb063ecbf2547

Observation b1265146-5025-4a62-9ade-7d01bbc92b3d · outbound

This paper cites SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.730210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.730210Z digest=sha256:a217a806ff00fcd1d86ff534716351fc4546534b7a326bbd5926539514c3058c

Observation 5b32dc12-74e6-4fa0-a14c-4bbdb6dde627 · outbound

This paper cites The categorization follows the practice in M3Exam (Zhang et al., 2023).

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia The categorization follows the practice in M3Exam (Zhang et al., 2023)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:53.986362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T15:57:53.734445Z digest=sha256:88afc0b5a8f852af4e7142a14c94bcab9ed3ef12a82508e8bc6ca82aa77cbdb4

Observation d3cf97c9-a3d1-436d-bc0b-29966329a1a4 · outbound

This paper cites In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia In Proceedings of the 2018 Conference on Empirical Methods in Nat- ural Language Processing, pages 2475–2485, Brus- sels, Belgium

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:53.998075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T15:57:53.648649Z digest=sha256:de8686d5ac43e1a3ed17e420ff44f4e0ed76a66297b31db0a4c209ba0070d8fd

Observation f3654095-6c62-4dee-86dd-3e9d77811fd2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.668748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.668748Z digest=sha256:984ecbb1ca647f4140469e5277f4a39295b96f4f71e6dd28214648ffddc2b539

Observation 9fb5a892-f8a3-472e-8d93-226ed0e29ac1 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Language Models are Multilingual Chain-of-Thought Reasoners

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.712567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.712567Z digest=sha256:10278ef6c485e844434da1e642bf13c98700e0f9a4505f4be3a6a52cee965364

Observation b89f303a-0c04-415f-8aea-e2d14b6bf51d · outbound

This paper cites In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing , pages 4232–4267, Singapore

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:57:54.008737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T15:57:53.639218Z digest=sha256:25cee71b66f6edcbd154fe3ea0594f79486590af74664773db2297783e12ee1a

Observation cacd0365-b2e3-4aa9-a438-c87297933b29 · outbound

This paper cites Aya 23: Open Weight Releases to Further Multilingual Progress.

SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia Aya 23: Open Weight Releases to Further Multilingual Progress

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T15:57:53.643963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:57:53.643963Z digest=sha256:0b0fc4fad83a5976f8bd4a6adfc3df6ad5b0fde05d5761aa4dd0fd9b45bfe073

Pith citing papers

Observation db639b68-38c1-4584-87dd-41e5125fb336 · inbound

Disentangling Language and Culture for Evaluating Multilingual Large Language Models cites this paper.

Disentangling Language and Culture for Evaluating Multilingual Large Language Models SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.361951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.361951Z digest=sha256:e10dfc481583d074ddcb02683fd394149d702b3d75ab8a962c95e010fbb5ce80

Observation 85887f63-adf8-4e4d-875a-20fe503a91f5 · inbound

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages cites this paper.

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:14:36.197621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T10:12:45.090257Z digest=sha256:79d51bc1003e56dc3b2af56afcddd6bc165be649d501b1691f06b3d51da33368