Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:46:10.247199Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.22554.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:46:10.247199Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 10914893-6e48-421d-8505-a2ccc6fbdcb3 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596eada7-4f2b-4972-8b9a-163f2522c0b4 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Alzahrani, H
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b717b53f-5b5b-4b0d-915d-6614c39ef06a · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Language Models are Few-Shot Learners
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6424a8-e04f-4a2d-87e9-96493113ac65 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Burnell, W
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f3e8c84-92fa-4261-a5a1-1d303965e839 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Carlini and D
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b9a763-c9ee-4d0b-92be-9ff4c9e18048 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Cheng, W
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 27735564-e6ed-48e1-a759-572d2726ad83 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75f8ed8-dbda-489a-9afd-a7c8969dd37e · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy PaLM: Scaling Language Modeling with Pathways
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7ab7e7-a642-4352-a0ba-3bb180369ff7 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e15cc938-0404-4057-a0ce-3e460d47cdd6 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Dekoninck, M
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8438c37c-3c5b-4570-a8e3-5b4c99f9686b · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63130074-c28e-4cb6-b638-6a99935d06cb · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Elazar, N
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec13489-cdbb-4739-93ea-1fc06ef9c218 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45883ebb-322f-4835-8f46-ae0ae11a5bde · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Logical Consistency of Large Language Models in Fact-checking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03de9f84-b2c6-493e-aa62-e3ce2dedda00 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Time Travel in LLMs: Tracing Data Contamination in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77061f62-a06e-4d80-bd86-f1785eb60333 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd217cc6-ce1d-47e8-bc2f-da8adea0a031 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Explaining and Harnessing Adversarial Examples
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88cf836e-e13c-461c-a86b-5b685da159d5 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71156f81-cd47-4aa2-8671-d2dbfa5f6667 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a307188b-73ca-4202-a9f2-5062074b491e · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d10879-6694-4b7d-8fa7-fe7a5584f7d2 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Measuring Mathematical Problem Solving With the MATH Dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f6235d-2149-47cb-afba-c2ce1ccacc46 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Jiang, Y
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1f02f3-d142-4be7-8be9-6d5ec32cfebb · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy UnifiedQA: Crossing Format Boundaries With a Single QA System
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb68d0b-8058-43ab-b64d-3343b1fb4f96 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d0d7f3-5acf-4317-b775-995494556946 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9caf502-53cd-465a-a0b2-1d1934c03339 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142f522a-1065-4fcf-9eb0-6359d5b97258 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Meier, J
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee9152c8-6f06-4730-8337-399b16bf8f74 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e57383b0-e86a-4ba8-965d-9ab2c244edd0 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Mizrahi, G
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdfa029c-8a03-42d6-9bc1-f9cb2f82fea3 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d209073-0d78-4e4a-9706-38b2acb7d5d8 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Pezeshkpour and E
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63d0e612-7a57-4e37-ac73-d86f7f36e6f1 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Efficient multi-prompt evaluation of LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7627d3f-aa13-4067-bcf7-f25315bb8d17 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Qwen2.5 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3d187d-8d6c-4ee4-a253-926d23321023 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d8d600-75b0-452f-ae02-e1d3973a07b1 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Improving Consistency in Large Language Models through Chain of Guidance
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48282414-8bb1-4cee-aeb7-6aafdb4a4c1f · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Beyond Accuracy: Behavioral Testing of NLP models with CheckList
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5a835c-c89d-4baa-bb83-3058e788e72b · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a46123-ac97-4ffe-a532-9b4ac9c3f628 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80b08e6-75a1-4926-a977-c62b1f877d6d · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Intriguing properties of neural networks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03b4b3e-f46e-4d6e-b014-7b5eda569423 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa52227-d9e8-44d8-83c7-ca29fde41da8 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Measuring short-form factuality in large language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3b26e5-b154-4bb5-a587-369a4003a05a · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Qwen3 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82aee369-6288-4d86-8fb9-fd435c89823d · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab95698f-49c7-42ce-9d2d-7c085669a96c · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f4f713-cf7f-4e13-9fbe-200e975f2b6e · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f37bc3e1-d55c-4f1d-bbdf-f919d32e6830 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy any-correct
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b58037e-10e7-4ba6-83cb-fbfcc93841f3 · outbound
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.