Pith. sign in

Paper Citation Record · LEDGER

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs

As of 13 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2411.09492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09492 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:37:54.170067Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a583e0b2-fff4-43bc-a625-9b19763b4f7f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.104239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.104239Z digest=sha256:d33a4c4a4fb443fa4a796ec4f9ff2e765832bf35ca0ce07cecdd121e9317563d

Observation 3d2f28a7-7d32-4941-8856-37071717a3eb · outbound

This paper cites In 9th Inter- national Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7,.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In 9th Inter- national Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.400214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.113503Z digest=sha256:4015fcee632c48f4038c714e89dd1095808884ff91764e0dce87582573fb9da5

Observation 2fb805f6-ffcb-4c88-a22a-05b1a3061333 · outbound

This paper cites In Forty- first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In Forty- first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.388444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.117708Z digest=sha256:fb17c170bb27783b008c4e9c661a50ac806b185521381541b7c52c50a65df825

Observation cb63b25f-2bbb-4426-b3e7-4107b98c1173 · outbound

This paper cites In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 , pages 15670–15693.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 , pages 15670–15693

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.376270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.122331Z digest=sha256:6a32c5a1ae17bb22fdc403ebb39251cc5d3fa5c7468acbdb423a60add6de8705

Observation ad58dcd4-a2d0-4d59-ba8a-fb84796bdf88 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.131256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.131256Z digest=sha256:4358c99173cde4890f584f588903b16e8abdba5a9c6200450b8b3c57ce01bcd6

Observation 7605f53c-3c0a-46cd-9332-dc91925f1006 · outbound

This paper cites In Proceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 1913–.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In Proceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 1913–

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.364381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.135748Z digest=sha256:b358e2839c68393e68ce76a69f54d4970224f195a54f00a79499fc86c147bd25

Observation f70cf994-8ab0-4101-ae3e-27dc5f1c983a · outbound

This paper cites In The Eleventh International Conference on Learning Representa- tions, ICLR 2023, Kigali, Rwanda, May 1-5,.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In The Eleventh International Conference on Learning Representa- tions, ICLR 2023, Kigali, Rwanda, May 1-5,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.351780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.140824Z digest=sha256:555598e0163e6f784a9b095df6e5bbb271c304deb809648de2de04da00602e57

Observation 7379018f-5b3a-4b18-9505-ad6ef93d5d10 · outbound

This paper cites an unresolved cited work.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:37:54.339779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.144146Z digest=sha256:b9d4b93904bdad49409693339f5d3e076ff80f82df898695e2fe572c2aefce5b

Observation e383a4ab-8d89-4145-8b6b-890585a9fea3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.147432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.147432Z digest=sha256:95050fa830ab955dbc78f5b198661c152912321025347b832ba14a35681447b2

Observation 956f0f73-2f92-44af-87d0-be59b52faf94 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.151127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.151127Z digest=sha256:26fb5388c8d13f90a1af13fa94d7f28716cf58a10aa7bb033fcfd5978af56b90

Observation 0246c4e8-d247-481b-88e9-1aac8e0c8e65 · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.155352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.155352Z digest=sha256:79e4ac2499fed6b2a22657dab828363b16555d16de114a0bc44814a0502f35cf

Observation d17697b9-b177-4fa2-81d9-132ac127f217 · outbound

This paper cites In Forty-first In- ternational Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In Forty-first In- ternational Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.326717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.158900Z digest=sha256:99c25ce86e6d8e52092c7e37f18a77fbf9ca611bf39684f1b85bcaf7821bd598

Observation 09c81f42-75d0-428d-a6fb-c03d622948c1 · outbound

This paper cites Qwen2 Technical Report.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Qwen2 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.162015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.162015Z digest=sha256:2e939025c25aa304fe4228f94ff5729c78360a59231577855848c1f7da4fa6fc

Observation ed924ecb-3d29-45f3-8873-19acf9d66081 · outbound

This paper cites In The Eleventh Inter- national Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In The Eleventh Inter- national Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.299173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.170067Z digest=sha256:eab0ad59c1e80f69f9b9cb4b5246d65e0b04f1a1611f4f7ab8b5e4a85e8e06e9

Observation 692d58da-d245-4dad-a178-50b5306446e4 · outbound

This paper cites In Proceedings of the 54th Annual Meet- ing of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Vol- ume 2: Short Papers.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs In Proceedings of the 54th Annual Meet- ing of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berlin, Germany, Vol- ume 2: Short Papers

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:37:54.312363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.166154Z digest=sha256:4889ad5527552daad61cac6bcee64ca862a94bea1d48bc87f597292ee9dffe8e

Observation 23b9a928-9cce-46b4-8be9-bac2dae23443 · outbound

This paper cites an unresolved cited work.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:37:54.423820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.099706Z digest=sha256:dc3608094a56bb2c185294220922539cda5388301d27e7620fe68eab6804963c

Observation 026deebf-ed8e-4443-920f-e1c366e1d565 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.094498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.094498Z digest=sha256:4bbdc09be1d3c92243e3cf37a7b3448b472b2442e186e9f46d85d3d6b70499df

Observation ecc7abe4-6737-4f1a-8d3f-e168d375c609 · outbound

This paper cites GPT-4 Technical Report.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T20:37:54.126879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:37:54.126879Z digest=sha256:0eb0fc25cbdeed10ec39431caeffec4714d49f66f3ca6b3e6542d11a0004113d

Observation 6cc66206-08b5-4892-a127-27a2e8caff87 · outbound

This paper cites an unresolved cited work.

MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:37:54.411886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T20:37:54.108897Z digest=sha256:229dcf52bcae3080969d07b71cac9119e025c14c4bc404e137ae100dda5f78db

Pith citing papers

No inbound Pith citation observations are available.