Pith. sign in

Paper Citation Record · LEDGER

Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2311.08596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08596 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:36:32.025197Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2def5796-26b9-438e-8497-2612da600ef3 · inbound

Sycophancy in Large Language Models: Causes and Mitigations cites this paper.

Sycophancy in Large Language Models: Causes and Mitigations Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:36:32.025197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:36:32.025197Z digest=sha256:834e45bdc1ba377e0936d43c7749304fa7a5d91f2a604c903f18951965f33bfc

Observation fc3d6c4b-c7a7-4956-9545-80b57ffb5c06 · inbound

Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies cites this paper.

Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T04:43:27.307708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:43:27.307708Z digest=sha256:d389633a2c33abfdf23ebad8d34d80a8c108cb7697a4f167e2823ce255edcbde

Observation f1d13927-1dc7-4c22-9040-94661f54fb49 · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.232010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:332773966129de55d14f749f44b1f7b43e119f33bfa8a242253be193d98313c5

Observation bc8217e7-49bb-4a50-85c4-b5747c275eff · inbound

B-score: Detecting biases in large language models using response history cites this paper.

B-score: Detecting biases in large language models using response history Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:49.174083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:33:49.174083Z digest=sha256:333d2ac5bcaf0490066eb30ac7a041e31785aa8270f0a334acc662c03eae4d32

Observation c4df5e8a-0f8c-4991-89c9-85255de74185 · inbound

Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems cites this paper.

Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:25.333973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:25.333973Z digest=sha256:0b0459d1bfc0a3b75c2a2f399d26d9b3cd6e3065775007ff63277325f54dca0e

Observation 1e87feb2-7423-4e88-8b1b-29f2c3fdc83d · inbound

BASIL: Bayesian Assessment of Sycophancy in LLMs cites this paper.

BASIL: Bayesian Assessment of Sycophancy in LLMs Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:56:52.000793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T21:55:45.714195Z digest=sha256:9b591878ce1cdc5eaa72a363dca15e276216dae134fa59d88d7efd202d9414f8

Observation 5c16d096-67af-4732-89ab-b2a341a18c89 · inbound

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI cites this paper.

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:07:58.764297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T14:03:49.595868Z digest=sha256:5eb2bceb21998e33829fd564c17072bf30050112a6a71a90e51084c9cd2bc3fb

Observation 5558a29a-fe74-4a63-9907-b2d985e4f019 · inbound

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care cites this paper.

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T08:39:30.412601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:39:30.412601Z digest=sha256:71fe3dd5afd0d36503c6ad71425838ba7699f3c527e8ee8b3a671e56a9f25d8e

Observation c482b472-01bd-4493-ab47-f2ca510ea8c5 · inbound

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems cites this paper.

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:13:13.995294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T20:09:14.217158Z digest=sha256:4ffb514c55fd27926355633b3904c325c3ddf539fb8db3402209fe6d6916d756

Observation 18e89a7e-ef12-4a18-9363-609cf5f14d5c · inbound

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems cites this paper.

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T16:57:39.838460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:57:39.838460Z digest=sha256:26a87dcbf0fb342ebf4aba30781a42842f22c5dae77b2329ce2d0fd47e64c97a

Observation 0ae3a3d9-84c2-4bd2-a46d-96905dd72ccb · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.414046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:38541b17c2ff57cc2004ca487c460a232f0edeecd2de6067797ab769fc6a9b56

Observation ef697d0d-7b1c-4878-80d9-aa4b018aad85 · inbound

Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts cites this paper.

Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:12.899263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T10:17:35.694775Z digest=sha256:1fecc792d834c3647146f4e520c186cbb827febc0878d7ca56c4bb3e469b3e07

Observation 535b8f9a-eaf0-4bfe-971f-a1c5c592ab12 · inbound

How LLMs Are Persuaded: A Few Attention Heads, Rerouted cites this paper.

How LLMs Are Persuaded: A Few Attention Heads, Rerouted Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:26:24.301302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T04:18:09.600353Z digest=sha256:ff5036a54fa1fb58c5623117179abd79466d4894d6518fcde87a6ed38fd95beb

Observation 1c1c0c47-da5b-4c5b-a3e6-a4f793afc945 · inbound

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct cites this paper.

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:51:18.405646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T08:47:51.191086Z digest=sha256:898e414b3730153d58868d095bca75c48de00210f74c5514359a101628484c83

Observation 902fba5c-3174-42de-b380-d99934bb858c · inbound

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience cites this paper.

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.449174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T01:29:12.725865Z digest=sha256:8563e0fe00677abb03151b5c964716067968245ab73df964293a93438a16d1ae

Observation 4179136a-eaea-4438-9913-19f8746a1566 · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:07.569477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:65f6c9b73828f3f03b29928e19ce24093116b724bf48292c6ab08f9740abbab5

Observation 2a291b7c-4e4b-4143-85b8-bc784d7045f3 · inbound

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks cites this paper.

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T06:02:15.946864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T06:02:15.946864Z digest=sha256:396a07fdaf6ac75a987a89297de313f5f89c9c6e08b686226c925b29b4b7ada1

Observation 11a93628-9b4a-4d3f-ad51-483a63abe15b · inbound

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs cites this paper.

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T05:39:51.084130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:39:51.084130Z digest=sha256:fe43b65b981c83fca8bb09e8d5dac0809d7aac7e017b74cc5cac880415fefa74