Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:53:06.503806Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 7 inbound Pith citation observations for arXiv:2412.00967.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:53:06.503806Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T05:19:12.781624Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:47:29.997154Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b710cbf-ab86-4b3f-bcfb-c14b113e4887 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Understanding intermediate layers using linear classifier probes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38caefae-1b8a-487c-b65e-8e9ae02132f5 · outbound
Linear Probe Penalties Reduce LLM Sycophancy A positivity bias in written and spoken english and its moderation by personality and gender
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 760992be-f392-4d6b-a37a-56f54cb1acc4 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c036c875-d4e1-40e7-b017-e0cde7eb336d · outbound
Linear Probe Penalties Reduce LLM Sycophancy On the Opportunities and Risks of Foundation Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38d39b0-c898-4b59-9a53-8d25da4c4240 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30d288e-1db0-4b2d-92f3-eb718c909ab5 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19cbed6-4c23-4db1-a75c-5969f5f28716 · outbound
Linear Probe Penalties Reduce LLM Sycophancy UltraFeedback: Boosting language models with high-quality feedback,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 592b7e19-a67b-4d82-a0a3-ab34ed448b11 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Human language reveals a universal positivity bias
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d6e4e40f-f362-4a43-b8bf-052f85703564 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f8a815df-0e3e-4948-ab7e-b3ee4c8dc1c7 · outbound
Linear Probe Penalties Reduce LLM Sycophancy siebert/sentiment-roberta-large-english · hugging face, 2021
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30c074fc-d465-468d-a45e-390d2eb8d031 · outbound
Linear Probe Penalties Reduce LLM Sycophancy LLM Agents can Autonomously Hack Websites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4653fe7c-b576-44d6-9454-969a87b5df0c · outbound
Linear Probe Penalties Reduce LLM Sycophancy Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa25cec5-b362-4418-acba-4aa39a577d84 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Interference of the end: Why recency bias in memory determines when a food is consumed again
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dfc429e2-4a5b-4a68-bef5-59f788c8e62a · outbound
Linear Probe Penalties Reduce LLM Sycophancy More than a feeling: Accuracy and application of sentiment analysis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75072bf4-7d34-46c6-bbcd-23118cbf1548 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Large Language Models are Zero-Shot Reasoners
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c2431a-ea83-4f41-b054-08c60154f045 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd14188e-5db7-4734-b332-363ad7596ad8 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Emergent linear representations in world models of self-supervised sequence models, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0449222-c72d-426f-8eb3-21cfed7a75e3 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e62d59-20a2-4170-a8cf-3153355c4d0b · outbound
Linear Probe Penalties Reduce LLM Sycophancy Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2bec276-2a96-4c61-a557-1c69ab9b8471 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Prioritizing High-Consequence Biological Capabilities in Evaluations of Artificial Intelligence Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d66ed55-9aec-4604-b8d9-da3a17c68fa6 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Discovering Language Model Behaviors with Model-Written Evaluations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a24925-d63e-4101-8aa4-2293805d0540 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Steering Llama 2 via Contrastive Activation Addition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b701c6-153e-4978-b58a-63b786c416cb · outbound
Linear Probe Penalties Reduce LLM Sycophancy Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4051584-aa8f-4bdd-856a-678f4f1f69ca · outbound
Linear Probe Penalties Reduce LLM Sycophancy Immediate and delayed primacy and recency effects in performance evaluation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b7a91471-2f7f-4fb4-9e85-fa3fb203032b · outbound
Linear Probe Penalties Reduce LLM Sycophancy Towards Understanding Sycophancy in Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b824f1-c9d9-4d28-8a48-a1739e65bf59 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Simple synthetic data reduces sycophancy in large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dc03a1-f3fd-411f-87d3-e659b81c1d3d · outbound
Linear Probe Penalties Reduce LLM Sycophancy OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3549866-c168-4a52-9ae9-47e7247b269b · outbound
Linear Probe Penalties Reduce LLM Sycophancy as a judge
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cfdd1245-6589-4265-8565-8991d0fb87cb · outbound
Linear Probe Penalties Reduce LLM Sycophancy Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda8f650-92d4-4d7a-b2df-01d42b317860 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9ecb7f78-a3bf-454b-9a14-2dda37db3fa4 · outbound
Linear Probe Penalties Reduce LLM Sycophancy (B) Positive
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dfc90fe1-cd57-4c9f-b170-393e68afd637 · outbound
Linear Probe Penalties Reduce LLM Sycophancy Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 682f048f-637d-4093-9789-ae5e8c50fb52 · outbound
Linear Probe Penalties Reduce LLM Sycophancy The distribution is now much more balanced
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d8a41a7-74b7-47c1-85ae-02c0e8f156ef · outbound
Linear Probe Penalties Reduce LLM Sycophancy UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c36c31-3b73-457e-92c9-4919fbb3cb87 · inbound
LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs Linear Probe Penalties Reduce LLM Sycophancy
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3323b6-488b-486c-a837-1b926db7d428 · inbound
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Linear Probe Penalties Reduce LLM Sycophancy
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2441dda6-046a-401f-a95b-2d7c6ac10835 · inbound
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs Linear Probe Penalties Reduce LLM Sycophancy
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da1a77b-b09a-4827-a2df-25dcd9cec6b9 · inbound
Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition Linear Probe Penalties Reduce LLM Sycophancy
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 56a958e8-e3e7-4eb0-964a-754c28de3fec · inbound
Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation Linear Probe Penalties Reduce LLM Sycophancy
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 31dbaf3a-b5c2-4511-b779-91f71efdcd61 · inbound
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Linear Probe Penalties Reduce LLM Sycophancy
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5594707-a557-43eb-a7e3-9b2adc876ff0 · inbound
Measuring and Detecting Harmful AI Sycophancy Linear Probe Penalties Reduce LLM Sycophancy
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.