Pith. sign in

Paper Citation Record · LEDGER

Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

As of 1 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2311.08596.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08596 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T06:02:15.946864Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f1d13927-1dc7-4c22-9040-94661f54fb49 · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.232010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:714b790932901a3a13c558405b3520874a786e1310b29df0f6d686eedbc0076d

Observation 1e87feb2-7423-4e88-8b1b-29f2c3fdc83d · inbound

BASIL: Bayesian Assessment of Sycophancy in LLMs cites this paper.

BASIL: Bayesian Assessment of Sycophancy in LLMs Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:56:52.000793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T21:55:45.714195Z digest=sha256:32ef312a5a0f3cde4bef1a84b472a772f36a171ec38b333c46f2b8fc47e671fb

Observation 5c16d096-67af-4732-89ab-b2a341a18c89 · inbound

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI cites this paper.

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:07:58.764297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-16T14:03:49.595868Z digest=sha256:5b25185fe9f484337a3822bda7438c5cc2e967987a254439fb8775a22475d630

Observation c482b472-01bd-4493-ab47-f2ca510ea8c5 · inbound

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems cites this paper.

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:13:13.995294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-13T20:09:14.217158Z digest=sha256:d035f9cfbecdc8463a537589311f0a6e68764618840ae668a97e52725d1cc378

Observation 0ae3a3d9-84c2-4bd2-a46d-96905dd72ccb · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.414046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:76a6644fd957625e3fb443f951fa7b5c6c4caf6f6443bbb1bfcaa8c1b51b9632

Observation ef697d0d-7b1c-4878-80d9-aa4b018aad85 · inbound

Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts cites this paper.

Towards Emotion Consistency Analysis of Large Language Models in Emotional Conversational Contexts Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:12.899263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-08T10:17:35.694775Z digest=sha256:89767f32fa316bb2b138150a2f5533c06f4d4de1e6923e9329edc6c80cf1d90b

Observation 535b8f9a-eaf0-4bfe-971f-a1c5c592ab12 · inbound

How LLMs Are Persuaded: A Few Attention Heads, Rerouted cites this paper.

How LLMs Are Persuaded: A Few Attention Heads, Rerouted Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:26:24.301302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-05-12T04:18:09.600353Z digest=sha256:a3a3f6250675183ac7de08db0023750fcb4dcfe6c671524ebab09c175ed8e598

Observation 1c1c0c47-da5b-4c5b-a3e6-a4f793afc945 · inbound

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct cites this paper.

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:51:18.405646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-22T08:47:51.191086Z digest=sha256:c5a07612c034343115f9b7f64317657001ff73519e0de2bb7b6f1d1620b9c59e

Observation 902fba5c-3174-42de-b380-d99934bb858c · inbound

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience cites this paper.

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.449174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-06-27T01:29:12.725865Z digest=sha256:8c3332636b1842c0a5a891b9efc5107f513ed2be5d3539846a11cfb9d36effbe

Observation 4179136a-eaea-4438-9913-19f8746a1566 · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:07.569477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:50f34f8707fd9ccadb48553d0ad9d5bdf2ac5798037943e5610e700aee80c618

Observation 2a291b7c-4e4b-4143-85b8-bc784d7045f3 · inbound

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks cites this paper.

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T06:02:15.946864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T06:02:15.946864Z digest=sha256:afa3d99cd3eda5385a4e1c724675d6a92906d3b4edf064809fd62b585a1af3fa