Pith. sign in

Paper Citation Record · LEDGER

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

As of 1 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2604.19809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19809 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:37:41.353329Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:35:47.679312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b9cade3-da4a-4a5f-8a95-e1f6a1c904bc · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:40:32.974213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:3121b6274aeb127e0f9c35bf7687dd7fb78fdf3a01f04c8b763c5e5641a64e70

Observation 3423b013-5f83-4a29-b54c-1058854f55ac · outbound

This paper cites MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-06-01T02:02:27.350423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:64fb1687e7d76b93d0c5aa649705d4b831da9c82b07014b95dac394bbf478686

Observation d6892b09-76c7-454f-bcdc-e748c4d6b1cc · outbound

This paper cites Each question has a unique identifier, domain label, subcategory label (5 per domain, 40 total), difficulty rating, question text, 4 answer choices, and a verified correct answer.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Each question has a unique identifier, domain label, subcategory label (5 per domain, 40 total), difficulty rating, question text, 4 answer choices, and a verified correct answer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.383028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:6272cf520df23ed57f101659a24c7f889cf14f46364e78df6a34e21540e1b444

Observation eb81bfdf-64d1-4d6f-9aa9-4204e6de1821 · outbound

This paper cites fixed” — “tailored.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models fixed” — “tailored

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.388045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:84b737f5aeecf78ce2b1ede99cdfccedf5f966074b9e576399759cb78fbc0556

Observation f2874805-4f77-44d8-b850-80460e13eee6 · outbound

This paper cites strong”—“weak.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models strong”—“weak

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.385458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:5acdf1e9dcdc46f56ead6e890ad62a39a65f96ef34a2e474cd28cf361b3e2d40

Observation ebb146e6-1bec-4afb-8cc2-c098974efa12 · outbound

This paper cites an unresolved cited work.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-18T23:22:53.427283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:07baa9a73f8914c3f06ca204423a239bbf2f71c2d8bb7dc1a37906070b2e8ae7

Observation 2c46ba44-9e2a-4796-9642-fbd5210221c2 · outbound

This paper cites an unresolved cited work.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-18T23:22:53.396683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:7cf379c87e1deb17eb19559241a028368a666f137c8fa0d4f5cf32a0b3b0ac28

Observation 216edc78-548e-4b08-826a-6f569014b74b · outbound

This paper cites All claims are empirical.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models All claims are empirical

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.407714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:dc84536f01149437695e7d256853ad269f2c24ec0571d01b8d366608e61d3cc6

Observation 9b8870bb-ef28-48ce-bcd4-3b7c9e756f0a · outbound

This paper cites Section 3.4 specifies infrastructure requirements (∼8,000 API calls, temperature=0).

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Section 3.4 specifies infrastructure requirements (∼8,000 API calls, temperature=0)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.416170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:f754c5f3bfdf9a96e2839d5c618d63c22f857144bd810110e34c722c87c411d7

Observation 69c5cf0c-7fd7-4067-a623-77d0f182bb36 · outbound

This paper cites Dataset files will be accessible there, and the Croissant metadata file is intended to be included as data/croissant metadata.json.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Dataset files will be accessible there, and the Croissant metadata file is intended to be included as data/croissant metadata.json

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.419839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:9f0ced4386290483289a31651251cf0a7192d33bfda37eb620b20ce35f324552

Observation 38747487-bfd3-4b5a-a811-057232132f6f · outbound

This paper cites All hyperparameters (temperature=0, wager scale 1–10, scoring rules) are specified.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models All hyperparameters (temperature=0, wager scale 1–10, scoring rules) are specified

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.422420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:737915f1f33a493caa7a19d08cc5d7268fbc313a03ccd94caa012279e635c83f

Observation ac4a0f93-8426-4f5a-bf2d-4d75d8d2eba0 · outbound

This paper cites Effect sizes (Cohen’sd) and bootstrap p-values are reported for the three escalation curve comparisons.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Effect sizes (Cohen’sd) and bootstrap p-values are reported for the three escalation curve comparisons

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.424824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:88d54e3c6b80085c2173b725fb81d81aa57038cc845f88de66d47e0a32a69290

Observation a75ade1b-5e89-4179-abd0-f65849b48119 · outbound

This paper cites Infrastructure: NVIDIA NIM (free tier), DeepSeek API, Google AI Studio.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Infrastructure: NVIDIA NIM (free tier), DeepSeek API, Google AI Studio

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.410330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:890c7bb4d70bb2e54c5b7fb1ac22eb9c5e8b4d37f637886b56927306518a5851

Observation 8d7b48ec-7dff-48b5-bcd7-a81b5dd671b2 · outbound

This paper cites Questions are factual across 8 cognitive domains.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Questions are factual across 8 cognitive domains

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.407118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:c142e099119a97eee9dcecaddd82a0891ea078eb16d461c87ba1a8c385a40a5c

Observation 8530fa40-2926-4ca5-a151-edab2bfd1935 · outbound

This paper cites Section 6.2 discusses Goodhart risk (models gaming MIRROR scores) and ecological validity limitations.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Section 6.2 discusses Goodhart risk (models gaming MIRROR scores) and ecological validity limitations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.403677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:2e5070c79fa3527a553ef2133388251015eaf77c3846f296e5a24f9bde9827a8

Observation d012cbf9-5560-45ad-a922-515a3e56720c · outbound

This paper cites Versioned releases support periodic updates.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Versioned releases support periodic updates

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.412894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:bf1113e77f63ccdd75e455cf93fce4e5a939a35aea3f53225a4cb4dddee8c2d4

Observation 61f93bfb-696a-4252-a5c0-9fbc143e9882 · outbound

This paper cites The repository is released under the MIT License, which is also reflected in the Croissant metadata.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models The repository is released under the MIT License, which is also reflected in the Croissant metadata

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.395646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:e9abe36a9f3c918d5f0e44ab68bad0a2703579b42607c9465605771d87cfbb19

Observation e96fbb64-904e-4de1-b916-571c69d5ce8e · outbound

This paper cites an unresolved cited work.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-18T23:22:53.390472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:129ecbb4ec633bf46c54aa7f5e25b699486aafddffe05b0771afe6febade122b

Observation 567ebbc7-feec-40b3-8e77-6814aaab0321 · outbound

This paper cites The human audit (Appendix W) was conducted by the authors.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models The human audit (Appendix W) was conducted by the authors

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.410811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:4e017c8c1edd42063ae9676d4e99267ba9ccd85cbdca6d2b3997e8f1098492ce

Pith citing papers

Observation 182ee173-6991-4f5f-8031-52d41ecddc7c · inbound

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory cites this paper.

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:47.679312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:47.679312Z digest=sha256:e05f8d20afd7d9d99ed5d8f15fc0e57c237eb70082fa40c2d98dce6c1b898b85