Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:41:34.878542Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2605.19276.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:41:34.878542Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:59:49.787962Z
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d592d2f6-3689-4cd5-a00e-b686cf03132e · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Longbench: A bilingual, multitask benchmark for long context understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 71843c89-91ee-46e1-b745-5a7c8035d78a · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models ARC Prize 2024: Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 75a1b657-6649-4020-9ce1-65c629aeccc8 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0a0d9803-7ce2-4924-b3d8-ac51f9fb4e28 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models MMEngine: Openmmlab foundational library for training deep learning models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ae30e3c5-e26f-4805-b60b-d8e64ac5360a · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Physics: Benchmarking foundation models on university-level physics problem solving
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f71df103-5455-45a5-b596-6a14c78bacfb · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring Massive Multitask Language Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 70f6b27c-8ebd-45ff-9323-15ff9bc55762 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring mathematical problem solving with the math dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd4e5296-2d2d-4cfc-a779-b68347ca344a · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 84eeaa5b-4d63-4434-92fd-42e42ae12c6d · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d4a2c6f4-899c-40a1-9da6-64e7ebe254a9 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8fa6604c-aee4-4044-a3d5-dc3a1c4a6f20 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9b8e936d-3c4c-4ce4-bb92-ef402ad1dd3c · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Humanity's Last Exam
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 04cfaca0-f851-48e4-a4ea-333a348a2415 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Generalizing Verifiable Instruction Following
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c8e155bb-869b-40e6-a463-0ed768a3a19c · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Gpqa: A graduate-level google-proof q&a benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 00e1dd4e-1e7b-4a83-a767-4a376a74bd9e · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Challenging big-bench tasks and whether chain-of-thought can solve them
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f8ec25ec-0c30-4789-b8ab-79d0c0058f19 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring short-form factuality in large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 69ed31f7-26d3-426e-a763-6ef3bf3c226b · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Openicl: An open-source framework for in-context learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 79721af4-5f87-4c63-b9df-a67a3c703a6a · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4342b2be-848a-402f-a365-b2b45a0d70d9 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a0f67ffe-b832-4361-b884-a0bb14748663 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0652a02d-29a8-4341-afe0-b8e1c143abd8 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Instruction-Following Evaluation for Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b523fd04-a9c8-4360-9e61-e2e1b58b08f1 · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 49ce5992-0ec6-407e-a9eb-ef2370e8097e · outbound
OpenCompass: A Universal Evaluation Platform for Large Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d9c2fd53-3dc8-4f36-93bc-8a52f7067c9a · inbound
MemSFT: Mitigating Alignment Tax with an External Parametric Memory OpenCompass: A Universal Evaluation Platform for Large Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.