Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:37:41.353329Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2604.19809.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:37:41.353329Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T23:35:47.679312Z
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b9cade3-da4a-4a5f-8a95-e1f6a1c904bc · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3423b013-5f83-4a29-b54c-1058854f55ac · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d6892b09-76c7-454f-bcdc-e748c4d6b1cc · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Each question has a unique identifier, domain label, subcategory label (5 per domain, 40 total), difficulty rating, question text, 4 answer choices, and a verified correct answer
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation eb81bfdf-64d1-4d6f-9aa9-4204e6de1821 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models fixed” — “tailored
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f2874805-4f77-44d8-b850-80460e13eee6 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models strong”—“weak
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ebb146e6-1bec-4afb-8cc2-c098974efa12 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2c46ba44-9e2a-4796-9642-fbd5210221c2 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 216edc78-548e-4b08-826a-6f569014b74b · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models All claims are empirical
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9b8870bb-ef28-48ce-bcd4-3b7c9e756f0a · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Section 3.4 specifies infrastructure requirements (∼8,000 API calls, temperature=0)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 69c5cf0c-7fd7-4067-a623-77d0f182bb36 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Dataset files will be accessible there, and the Croissant metadata file is intended to be included as data/croissant metadata.json
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 38747487-bfd3-4b5a-a811-057232132f6f · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models All hyperparameters (temperature=0, wager scale 1–10, scoring rules) are specified
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation ac4a0f93-8426-4f5a-bf2d-4d75d8d2eba0 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Effect sizes (Cohen’sd) and bootstrap p-values are reported for the three escalation curve comparisons
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a75ade1b-5e89-4179-abd0-f65849b48119 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Infrastructure: NVIDIA NIM (free tier), DeepSeek API, Google AI Studio
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8d7b48ec-7dff-48b5-bcd7-a81b5dd671b2 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Questions are factual across 8 cognitive domains
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8530fa40-2926-4ca5-a151-edab2bfd1935 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Section 6.2 discusses Goodhart risk (models gaming MIRROR scores) and ecological validity limitations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation d012cbf9-5560-45ad-a922-515a3e56720c · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Versioned releases support periodic updates
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 61f93bfb-696a-4252-a5c0-9fbc143e9882 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models The repository is released under the MIT License, which is also reflected in the Croissant metadata
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e96fbb64-904e-4de1-b916-571c69d5ce8e · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 567ebbc7-feec-40b3-8e77-6814aaab0321 · outbound
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models The human audit (Appendix W) was conducted by the authors
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 182ee173-6991-4f5f-8031-52d41ecddc7c · inbound
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.