Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2407.12844.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:23:53.806181Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T21:28:58.087704Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e90a0b89-3763-4ded-b0ac-f717c25c9ddd · inbound
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d40505a-3946-4f22-8c5e-6efaccecf898 · inbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd35b406-d089-423a-ade4-f0b92e37ae66 · inbound
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6133d1-c2e2-4143-8d7f-6fd5889bdb46 · inbound
Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1005c1-f974-401e-b4b1-1b1040bab0af · inbound
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1e313c25-af52-415a-8245-306e89ebdefc · inbound
Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 471cbdb5-6644-4b77-8c5b-0595ff536f0d · inbound
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d6229044-6503-4080-99b2-6ee5ecc8fcca · inbound
Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 51e162ca-9c1a-4458-8237-0c2b8b11f6a3 · inbound
Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 39807472-4816-4985-b804-9c3523e29efa · inbound
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9f99d01b-c72f-486b-ac57-0e4c82e19be4 · inbound
AGC-Bench: Measuring Artificial General Creativity metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 23705eed-4c7b-4b28-81e6-563dc571beeb · inbound
AGC-Bench: Measuring Artificial General Creativity metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation aa4f8dd3-6f04-4252-9483-c0a11dc34989 · inbound
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9241b1ca-3a78-41bb-9ad6-6c35e8772641 · inbound
Item Response Theory for AI Safety metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.