Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:31.461050Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2504.14039.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:31.461050Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-08T16:21:18.483029Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T18:16:11.503799Z
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 243db047-b1bd-4a26-bc73-269e32d69c53 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c51126b-cc54-49cb-9df2-06b5469e42f7 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2430c4b-1e30-4de9-b5a6-fa571ef65659 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a87dbfc-f786-4f7c-a054-802783b232cb · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks SECURE: Benchmarking Large Language Models for Cybersecurity
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c38c05b-8733-47b6-8c6d-18a5a200348c · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 313abe81-bff7-4440-8cac-2b226a726fee · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88dcf26a-c69d-4da8-9115-19410e580af6 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Evaluating Superhuman Models with Consistency Checks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acff226d-cf73-4073-8d60-97dd577a17d3 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks A framework for few-shot language model evaluation, 12 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d3bdb1-28b8-4a23-90a8-a4f6c45bb64e · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Measurement and Fairness
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ceda8b9-4bb1-46af-9560-e7485d4d4f86 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks SEvenLLM: Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05a05ba-6c23-4f4f-9569-6de44eb14441 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Seceval: A comprehensive benchmark for eval- uating cybersecurity knowledge of foundation models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf77ecd3-503e-4bc6-aa72-f110d08e7625 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a79429-96d1-48da-961d-a9b8afb933b7 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks ROUGE: A package for automatic evaluation of summaries
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 89da18fe-18d3-4128-ab03-f233e45f5cc0 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692c1e35-aa40-4370-8630-f4ad00f0e8bf · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e0382f-c17c-4226-8a95-c333b5fe7729 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks State of What Art? A Call for Multi-Prompt LLM Evaluation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807142c3-38b0-45f1-b62e-61073d9f8e5e · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e30fb3-fd28-4cf1-bb9d-2484d3b7cd8f · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Jinja2 Documentation, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4bc161a4-aab5-429b-b3e9-b68d7cf27dfe · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e52be142-08b5-4129-a223-2cbdae6225f2 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt for- matting, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c826add8-ad35-48ae-a152-2a3d399737ab · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Large Language Models are Inconsistent and Biased Evaluators
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c2b4058-a6ac-46fe-b104-ad97b75c80a2 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks It Takes Two to Tango: Navigating Conceptualizations of NLP Tasks and Measurements of Performance
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d61b9a3-360a-440f-8743-57679b614f7f · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f2cb32-15ef-4bcf-818e-89b9aaee5a60 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3f79fc-2b08-4467-bf0a-ad11d353c58a · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426e0c94-8ccc-4abf-a0c1-bcdde1e06c57 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc37ec53-3f00-410b-9698-170f711e3da8 · outbound
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1873b43c-28bf-4f49-85f1-1054990fea01 · inbound
Evaluating the Reliability of Multiple Large Language Models in Risk Assessment: A CIS Controls Based Approach MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.