Pith. sign in

Paper Citation Record · LEDGER

OpenCompass: A Universal Evaluation Platform for Large Language Models

As of 3 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2605.19276.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19276 v3

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:41:34.878542Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:59:49.787962Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact10
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d592d2f6-3689-4cd5-a00e-b686cf03132e · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

OpenCompass: A Universal Evaluation Platform for Large Language Models Longbench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.842784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:64a0ce9af396b81b5be49e8f3cb849a4397748c7707861c2be4ed96d6fa53f1d

Observation 71843c89-91ee-46e1-b745-5a7c8035d78a · outbound

This paper cites ARC Prize 2024: Technical Report.

OpenCompass: A Universal Evaluation Platform for Large Language Models ARC Prize 2024: Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.502117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:a77269c867246d15a9425f09166f81976333e37357c6ebe7ca4f16db148e99c5

Observation 75a1b657-6649-4020-9ce1-65c629aeccc8 · outbound

This paper cites Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy.

OpenCompass: A Universal Evaluation Platform for Large Language Models Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.840887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:91389ec65a4c9d7bac323b0331df56bc052892bf96f179da358b86f795d91ea5

Observation 0a0d9803-7ce2-4924-b3d8-ac51f9fb4e28 · outbound

This paper cites MMEngine: Openmmlab foundational library for training deep learning models.

OpenCompass: A Universal Evaluation Platform for Large Language Models MMEngine: Openmmlab foundational library for training deep learning models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.831285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:74885413822dd8e4089e6bc78920f22e703cae8878a45a52881a88dbd1801df5

Observation ae30e3c5-e26f-4805-b60b-d8e64ac5360a · outbound

This paper cites Physics: Benchmarking foundation models on university-level physics problem solving.

OpenCompass: A Universal Evaluation Platform for Large Language Models Physics: Benchmarking foundation models on university-level physics problem solving

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.838966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:66a9d4d1f1eb69a7ec562724c3c7da05dc9d684ae8519876bc182d2113f94064

Observation f71df103-5455-45a5-b596-6a14c78bacfb · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring Massive Multitask Language Understanding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.508507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:8ad5d1e647403420015e1a6d94ce9602b84cebd358394b9be2ac735673334b51

Observation 70f6b27c-8ebd-45ff-9323-15ff9bc55762 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring mathematical problem solving with the math dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.822755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:154b1ab64b364e650c89ca9aac3a12c129204452dac415ffe4bb65623df309f3

Observation fd4e5296-2d2d-4cfc-a779-b68347ca344a · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

OpenCompass: A Universal Evaluation Platform for Large Language Models RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.505348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:62c17252065d224480ce661d5068a68c6db9c2d6c55f1ee14c349c7d05ebb5e6

Observation 84eeaa5b-4d63-4434-92fd-42e42ae12c6d · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

OpenCompass: A Universal Evaluation Platform for Large Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.513123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:7667dfa8fb46d12cb42b7a6e8c521e299b9052dfa7b361eb1737e0aeef9933d0

Observation d4a2c6f4-899c-40a1-9da6-64e7ebe254a9 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

OpenCompass: A Universal Evaluation Platform for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.837020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:480359d59acb485713762ed152c4c42ec55c355ecb9f02e216f9894ea9df52e6

Observation 8fa6604c-aee4-4044-a3d5-dc3a1c4a6f20 · outbound

This paper cites ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models.

OpenCompass: A Universal Evaluation Platform for Large Language Models ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.510578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:41b4eac1bdffbb35ca6e9ba0340fa707d20d2e8ffc5b2e6259ac5030ba55ad00

Observation 9b8e936d-3c4c-4ce4-bb92-ef402ad1dd3c · outbound

This paper cites Humanity's Last Exam.

OpenCompass: A Universal Evaluation Platform for Large Language Models Humanity's Last Exam

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.519875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:a4cafe33897292d400546d23ed31e7cccee6c29e891260ef9f16bf7f26867b81

Observation 04cfaca0-f851-48e4-a4ea-333a348a2415 · outbound

This paper cites Generalizing Verifiable Instruction Following.

OpenCompass: A Universal Evaluation Platform for Large Language Models Generalizing Verifiable Instruction Following

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.522603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:5cc7137f20d981881644a85e2753c869855d36d2a96f249625ebbf63eb841fea

Observation c8e155bb-869b-40e6-a463-0ed768a3a19c · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

OpenCompass: A Universal Evaluation Platform for Large Language Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.844694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:86602216e96bd0a5e032cb1952a66bdef3cb2ce3d9c8f3970938cd27295ef5d3

Observation 00e1dd4e-1e7b-4a83-a767-4a376a74bd9e · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

OpenCompass: A Universal Evaluation Platform for Large Language Models Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.835142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:06e159594d0f640e707cde03c9a5398cf1972d2c7e1b1458e7c7f6c28c07d54b

Observation f8ec25ec-0c30-4789-b8ab-79d0c0058f19 · outbound

This paper cites Measuring short-form factuality in large language models.

OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring short-form factuality in large language models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.504164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:d9bdcba0f9a70ad996c48bce43993e979780ff38ae8263473d63a1099c50f424

Observation 69ed31f7-26d3-426e-a763-6ef3bf3c226b · outbound

This paper cites Openicl: An open-source framework for in-context learning.

OpenCompass: A Universal Evaluation Platform for Large Language Models Openicl: An open-source framework for in-context learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.827268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:1b718135898bdeba8b76e1587dd10b76b91fecf91423945f52629200e438a935

Observation 79721af4-5f87-4c63-b9df-a67a3c703a6a · outbound

This paper cites LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset.

OpenCompass: A Universal Evaluation Platform for Large Language Models LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:45:00.507659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:01cae19347a8a893da30190b5cc1569c1e68a25bf012bc50c51a52aa54744977

Observation 4342b2be-848a-402f-a365-b2b45a0d70d9 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL).

OpenCompass: A Universal Evaluation Platform for Large Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.824937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:a8c6061688a52ecb72322d06a101dcd90b6fcfd0c88273441d7844b9c21aa245

Observation a0f67ffe-b832-4361-b884-a0bb14748663 · outbound

This paper cites P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms.

OpenCompass: A Universal Evaluation Platform for Large Language Models P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.829275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:f745d8d46e48e80575ea11114403b8bb1b83514f1be4a1e7252976245721de83

Observation 0652a02d-29a8-4341-afe0-b8e1c143abd8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

OpenCompass: A Universal Evaluation Platform for Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.527674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:fc7b10af6baf89003fcf80c69471a4d1b91337ffe6c6b66640ff4b6d10192a4f

Observation b523fd04-a9c8-4360-9e61-e2e1b58b08f1 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

OpenCompass: A Universal Evaluation Platform for Large Language Models BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.525070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:db568e215f5eeabb252625ee57b37e3860434f0840987608cab8c7464698ae51

Observation 49ce5992-0ec6-407e-a9eb-ef2370e8097e · outbound

This paper cites an unresolved cited work.

OpenCompass: A Universal Evaluation Platform for Large Language Models Unresolved cited work

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T03:34:32.833117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:a923d7d1711868ba676a639698cb1f15809b950c673a4edae406d94bbcf21496

Pith citing papers

Observation d9c2fd53-3dc8-4f36-93bc-8a52f7067c9a · inbound

MemSFT: Mitigating Alignment Tax with an External Parametric Memory cites this paper.

MemSFT: Mitigating Alignment Tax with an External Parametric Memory OpenCompass: A Universal Evaluation Platform for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T01:59:49.787962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:59:49.787962Z digest=sha256:790b28a082bb2f84a4687bb2b3077945deb2ee8a415e8bcdd0e23a40783c2865