Pith. sign in

Paper Citation Record · LEDGER

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

As of 21 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.22554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22554 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:46:10.247199Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10914893-6e48-421d-8505-a2ccc6fbdcb3 · outbound

This paper cites GPT-4 Technical Report.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:06.913577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:06.913577Z digest=sha256:5e2e9eda731bae6430108c9763c35ba36298bc8e2235b2cfe28af8b7654809b8

Observation 596eada7-4f2b-4972-8b9a-163f2522c0b4 · outbound

This paper cites Alzahrani, H.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Alzahrani, H

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:06.981595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:06.981595Z digest=sha256:02afd6f6f3f3c2d7b20950ab5ccdda7c27c1f5b37aabb0ce3a8e6f8ef7e2632a

Observation b717b53f-5b5b-4b0d-915d-6614c39ef06a · outbound

This paper cites Language Models are Few-Shot Learners.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.073994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.073994Z digest=sha256:571d3f5f26d36bc774421b7631028463cad04fb1d083dc66702eb5256435645d

Observation ce6424a8-e04f-4a2d-87e9-96493113ac65 · outbound

This paper cites Burnell, W.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Burnell, W

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.169312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.169312Z digest=sha256:d6cfcdbd7eafd663c885050a61e9d02ee7b75e6afa5500f1f25312a3b8e9221d

Observation 4f3e8c84-92fa-4261-a5a1-1d303965e839 · outbound

This paper cites Carlini and D.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Carlini and D

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.237953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.237953Z digest=sha256:75800cdc3035597601e2b3c76cf0be864c4888de94268c4eec0d20cd01d3c3ad

Observation d2b9a763-c9ee-4d0b-92be-9ff4c9e18048 · outbound

This paper cites Cheng, W.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Cheng, W

Reference 6

Resolution
verified exact
doi, observed 2026-08-02T13:48:30.686990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-02T13:46:07.329004Z digest=sha256:6797fdcf8e1187d469c47d99f2dbcf9e73c267d40a32b33c1318a538d8a47e4b

Observation 27735564-e6ed-48e1-a759-572d2726ad83 · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.413422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.413422Z digest=sha256:b5965585713c0832b57167522ac04c506dccf88170b776d2bb5934d793507860

Observation f75f8ed8-dbda-489a-9afd-a7c8969dd37e · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.475799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.475799Z digest=sha256:65d0ee7b57c6b4cf3924d47185f3685306633c045cba99cc03a0e95fcc54ba1d

Observation 2f7ab7e7-a642-4352-a0ba-3bb180369ff7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.542054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.542054Z digest=sha256:71db8d83c03239f27c8f13acabed6fdd38cdb7369a9f326a2a3c22af45d4fb73

Observation e15cc938-0404-4057-a0ce-3e460d47cdd6 · outbound

This paper cites Dekoninck, M.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Dekoninck, M

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.607889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.607889Z digest=sha256:57a2b01aa20c4b3f94aeb13c81ddbc1ccc6a0fcdd60349d627a9053728426de8

Observation 8438c37c-3c5b-4570-a8e3-5b4c99f9686b · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.660124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.660124Z digest=sha256:9edd16a9e4d2502f76828a31aced68a6f5779574d0c10ae85bf9cac8d5ec91fb

Observation 63130074-c28e-4cb6-b638-6a99935d06cb · outbound

This paper cites Elazar, N.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Elazar, N

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.729513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.729513Z digest=sha256:3bf31100973cd9361f22a695a82947e6ed38acacaf33aeb29a3f94e6a4a652bf

Observation 7ec13489-cdbb-4739-93ea-1fc06ef9c218 · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.790012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.790012Z digest=sha256:9e1975968a21c0f0776a11eadbb07f507d0224796d1d078c73c7381933c0d6cb

Observation 45883ebb-322f-4835-8f46-ae0ae11a5bde · outbound

This paper cites Logical Consistency of Large Language Models in Fact-checking.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Logical Consistency of Large Language Models in Fact-checking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.839786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.839786Z digest=sha256:44801fc9eea1d82f98f67511fcbaf88239af5b238b4363bdf3f3bad5e7e0c0bb

Observation 03de9f84-b2c6-493e-aa62-e3ce2dedda00 · outbound

This paper cites Time Travel in LLMs: Tracing Data Contamination in Large Language Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Time Travel in LLMs: Tracing Data Contamination in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.903416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.903416Z digest=sha256:6937a4aaf0e978d685e2b41a00d80b64044c2012f4899ea1d7870a7caa362fc7

Observation 77061f62-a06e-4d80-bd86-f1785eb60333 · outbound

This paper cites Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:07.961355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:07.961355Z digest=sha256:1f5de630bf5b817738e549af337f49abfad8f3bc2a6ace9fa58bd3dd593cf31f

Observation bd217cc6-ce1d-47e8-bc2f-da8adea0a031 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Explaining and Harnessing Adversarial Examples

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.024564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.024564Z digest=sha256:0cf539a3763d6963cb3e6f8ee055b2c4a7a32363ae6bdcd643f3b0bdc6ddf1b6

Observation 88cf836e-e13c-461c-a86b-5b685da159d5 · outbound

This paper cites The Llama 3 Herd of Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.091732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.091732Z digest=sha256:79e448ba58bd7da22f2556cf5260ae4527bb93afd5c377b3159b6a8dd19d607a

Observation 71156f81-cd47-4aa2-8671-d2dbfa5f6667 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.188766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.188766Z digest=sha256:4a414930071ece61d06b6e144fb2c0432fcbf3f363722daad8cd11b4932c85bc

Observation a307188b-73ca-4202-a9f2-5062074b491e · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.272905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.272905Z digest=sha256:5ca707d74249e9d4a6579503f6d38da40aa255a8b97894f7e9cb034b7de6b837

Observation c6d10879-6694-4b7d-8fa7-fe7a5584f7d2 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.384334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.384334Z digest=sha256:076649f35ae84c391ac6d29d7f9609eac31ae2464a9bd4ac68d4048799caf9ba

Observation 48f6235d-2149-47cb-afba-c2ce1ccacc46 · outbound

This paper cites Jiang, Y.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Jiang, Y

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.447963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.447963Z digest=sha256:2ccc54777be7bda4f32e9eb8ab1661ef2814ab29d830643c2e68fa8d4f304006

Observation 8e1f02f3-d142-4be7-8be9-6d5ec32cfebb · outbound

This paper cites UnifiedQA: Crossing Format Boundaries With a Single QA System.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy UnifiedQA: Crossing Format Boundaries With a Single QA System

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.528174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.528174Z digest=sha256:1af536e9d7ea8ea921fe41b2a10ade18a527aebeaa7abd5b00805a3007ac1f4d

Observation 5cb68d0b-8058-43ab-b64d-3343b1fb4f96 · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.592377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.592377Z digest=sha256:b594afce3237c4fb954759f310342f1d2677189c28cc6ac3fca385fb3909873a

Observation 86d0d7f3-5acf-4317-b775-995494556946 · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.684247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.684247Z digest=sha256:956c3187fd240546d394ce42e0a96b64d4c3c34fed02637d292340a6b937b51c

Observation b9caf502-53cd-465a-a0b2-1d1934c03339 · outbound

This paper cites On Robustness and Reliability of Benchmark-Based Evaluation of LLMs.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.844007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.844007Z digest=sha256:e51e23eab95d3f772c66ee344d27045ba13009057414ef211fbc416477316742

Observation 142f522a-1065-4fcf-9eb0-6359d5b97258 · outbound

This paper cites Meier, J.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Meier, J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.920095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.920095Z digest=sha256:dc5704f5536de5ea16c90c7193fffc958a6cf63e16e2c9d5bac7035c944ef1c2

Observation ee9152c8-6f06-4730-8337-399b16bf8f74 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.001294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.001294Z digest=sha256:4947f0c704ab6d6b422225dbf04c5ac8f1d3ec5e5c1be4fff4f85e9752d6dd42

Observation e57383b0-e86a-4ba8-965d-9ab2c244edd0 · outbound

This paper cites Mizrahi, G.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Mizrahi, G

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.055517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.055517Z digest=sha256:a8997adeae4b07a641e78bda62d57aa9ff7aeaeca7f9511bfc46168d972df844

Observation bdfa029c-8a03-42d6-9bc1-f9cb2f82fea3 · outbound

This paper cites SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.135848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.135848Z digest=sha256:050d812d831a1bf9a5a270014cde9fe2571bdff98e8dfa9b1551c3ee977f19f0

Observation 3d209073-0d78-4e4a-9706-38b2acb7d5d8 · outbound

This paper cites Pezeshkpour and E.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Pezeshkpour and E

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.189491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.189491Z digest=sha256:437d49191e50bbce931c47cef828f4d889402753e72e27c79cb1e0741efa1ebc

Observation 63d0e612-7a57-4e37-ac73-d86f7f36e6f1 · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Efficient multi-prompt evaluation of LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.251987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.251987Z digest=sha256:e289be7d340d2da915717161acb1deb48406ee8ef8f8806f543d063d8b0e2a3f

Observation d7627d3f-aa13-4067-bcf7-f25315bb8d17 · outbound

This paper cites Qwen2.5 Technical Report.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Qwen2.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.314188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.314188Z digest=sha256:c3d6f2c4b42e7cfd2ed2811c925579a1952f41bde00ed573d2ff5ff7e5360cc7

Observation 1f3d187d-8d6c-4ee4-a253-926d23321023 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.399690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.399690Z digest=sha256:9165dcf5aab0b05a7755d6d54bd1af89506d8d3646c44e5c2edc9e3e16efaf54

Observation 39d8d600-75b0-452f-ae02-e1d3973a07b1 · outbound

This paper cites Improving Consistency in Large Language Models through Chain of Guidance.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Improving Consistency in Large Language Models through Chain of Guidance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.473023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.473023Z digest=sha256:6cf3d1ebe7419f1ce1bbf2e0aeeeb8595c6dce33e0efb8519430fd618b762efa

Observation 48282414-8bb1-4cee-aeb7-6aafdb4a4c1f · outbound

This paper cites Beyond Accuracy: Behavioral Testing of NLP models with CheckList.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.525133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.525133Z digest=sha256:2bf9c0aafaa956d380971f9e762ee979e9eb23f135f7437c062f27cdef60241b

Observation 4b5a835c-c89d-4baa-bb83-3058e788e72b · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.599956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.599956Z digest=sha256:58e5cc5c26354317b1554f5c471f56725189c89416e8cad5e67708c1e0ae3525

Observation e9a46123-ac97-4ffe-a532-9b4ac9c3f628 · outbound

This paper cites Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.669375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.669375Z digest=sha256:e748521a15d6b26328607b0ae96751fb68253f5730f1a8bd9c45320ab7e62378

Observation a80b08e6-75a1-4926-a977-c62b1f877d6d · outbound

This paper cites Intriguing properties of neural networks.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Intriguing properties of neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.723900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.723900Z digest=sha256:ace32eab23f8fee2bd866c183b05f5f716d62365fe32ff1af55fc188944e02e1

Observation c03b4b3e-f46e-4d6e-b014-7b5eda569423 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.775824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.775824Z digest=sha256:1e8a56c50d7a409401aa99997e85e59ecb04b77b4158b220e7d342d30a13698b

Observation dfa52227-d9e8-44d8-83c7-ca29fde41da8 · outbound

This paper cites Measuring short-form factuality in large language models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Measuring short-form factuality in large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.835881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.835881Z digest=sha256:d1abc4abeca90cc01737ec8b5a463a66c8d63be09ee27aa793344432fecc0448

Observation ee3b26e5-b154-4bb5-a587-369a4003a05a · outbound

This paper cites Qwen3 Technical Report.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Qwen3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.897940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.897940Z digest=sha256:90580c76056c8872ecbbf88242159a27589373cbb898ac7ed27b865512689dcc

Observation 82aee369-6288-4d86-8fb9-fd435c89823d · outbound

This paper cites Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-08-02T13:46:09.953784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.953784Z digest=sha256:7c9de9aacf194cbde0f3020de5e5c3c8475a8e4465d92a3c96a0c9762372eaaf

Observation ab95698f-49c7-42ce-9d2d-7c085669a96c · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:10.012221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:10.012221Z digest=sha256:87a21b201ec09eedef2578ed621fb84e6e3787303f3e4ecfcc89f971604f28de

Observation 35f4f713-cf7f-4e13-9fbe-200e975f2b6e · outbound

This paper cites an unresolved cited work.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:10.107243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:10.107243Z digest=sha256:a737d105998709a3a99c460aab8f564cba86137e69667c737f482c4334fa33f9

Observation f37bc3e1-d55c-4f1d-bbdf-f919d32e6830 · outbound

This paper cites any-correct.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy any-correct

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-02T13:46:10.247199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:10.247199Z digest=sha256:35b74eb0c5fc7d0d09b65de0e7ab4961ddcc7ef342c05f5c94041b4af183a6f6

Observation 9b58037e-10e7-4ba6-83cb-fbfcc93841f3 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:08.764054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:08.764054Z digest=sha256:7829431a5fe01dcb810e95b2b2c8dcbe2da47d3d20e586ed4664af4b4f4e83f2

Pith citing papers

No inbound Pith citation observations are available.