Pith. sign in

Paper Citation Record · LEDGER

Linear Probe Penalties Reduce LLM Sycophancy

As of 23 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 7 inbound Pith citation observations for arXiv:2412.00967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00967 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:53:06.503806Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:19:12.781624Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:47:29.997154Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b710cbf-ab86-4b3f-bcfb-c14b113e4887 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Linear Probe Penalties Reduce LLM Sycophancy Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.372154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.372154Z digest=sha256:40e6847e37151b1a3509dacd45bb8825de51df00e27b1e697452989462195a63

Observation 38caefae-1b8a-487c-b65e-8e9ae02132f5 · outbound

This paper cites A positivity bias in written and spoken english and its moderation by personality and gender.

Linear Probe Penalties Reduce LLM Sycophancy A positivity bias in written and spoken english and its moderation by personality and gender

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.890407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.376698Z digest=sha256:f7769c1243f5a7217f6f5a8da9f72b16cb4e9fdaf84b66618f4fa05b4442780b

Observation 760992be-f392-4d6b-a37a-56f54cb1acc4 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Linear Probe Penalties Reduce LLM Sycophancy Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.380667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.380667Z digest=sha256:94409cfefa4cb6530f70d8cd552e52dbec10b546c5e008935a95e67ded825f76

Observation c036c875-d4e1-40e7-b017-e0cde7eb336d · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Linear Probe Penalties Reduce LLM Sycophancy On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.384899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.384899Z digest=sha256:85a6dfe3fa3181d6117a11701ec26db144eb705ac5995e8404924c6f3ac9c50f

Observation b38d39b0-c898-4b59-9a53-8d25da4c4240 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Linear Probe Penalties Reduce LLM Sycophancy Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.389428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.389428Z digest=sha256:5cae00352ec729e0844207edd09f45404403b2d695cd290904fc673f34663290

Observation b30d288e-1db0-4b2d-92f3-eb718c909ab5 · outbound

This paper cites Deep reinforcement learning from human preferences.

Linear Probe Penalties Reduce LLM Sycophancy Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.393513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.393513Z digest=sha256:a8a4f37eae52d2cda2b9e1d368b37117d6c7eb5d9053f515342c76e492d818d1

Observation a19cbed6-4c23-4db1-a75c-5969f5f28716 · outbound

This paper cites UltraFeedback: Boosting language models with high-quality feedback,.

Linear Probe Penalties Reduce LLM Sycophancy UltraFeedback: Boosting language models with high-quality feedback,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.870440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.397879Z digest=sha256:0410b929b7295c9c20ce42c72bbca4bb4017ab70146a4fb6ec8634d2b5adb09d

Observation 592b7e19-a67b-4d82-a0a3-ab34ed448b11 · outbound

This paper cites Human language reveals a universal positivity bias.

Linear Probe Penalties Reduce LLM Sycophancy Human language reveals a universal positivity bias

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.858415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.405999Z digest=sha256:d659999270ad0d670e26e64c613d104ad1bd7c6826b6ce2d2888724a8c99dc54

Observation d6e4e40f-f362-4a43-b8bf-052f85703564 · outbound

This paper cites an unresolved cited work.

Linear Probe Penalties Reduce LLM Sycophancy Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:53:06.847453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.409804Z digest=sha256:f8267d760a472cb5626a2e2cb55b3747cdc5980fa840d89b8472a71953a8fdaf

Observation f8a815df-0e3e-4948-ab7e-b3ee4c8dc1c7 · outbound

This paper cites siebert/sentiment-roberta-large-english · hugging face, 2021.

Linear Probe Penalties Reduce LLM Sycophancy siebert/sentiment-roberta-large-english · hugging face, 2021

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.836017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.413760Z digest=sha256:0c3c93bd80d9347474e6d28cc91a29d9379c1aa9eae1e3c93bf74393a4ca6470

Observation 30c074fc-d465-468d-a45e-390d2eb8d031 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

Linear Probe Penalties Reduce LLM Sycophancy LLM Agents can Autonomously Hack Websites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.417443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.417443Z digest=sha256:533b8ebb92900b3cffae4a38e5e36cae808c55471ddaf3fae2ee5ff8fd55b7f6

Observation 4653fe7c-b576-44d6-9454-969a87b5df0c · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Linear Probe Penalties Reduce LLM Sycophancy Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.421716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.421716Z digest=sha256:6a640cefdd04db2b602fdd389d6a4e721c9d7cc6ca6c8d7d32cc1124935ec598

Observation fa25cec5-b362-4418-acba-4aa39a577d84 · outbound

This paper cites Interference of the end: Why recency bias in memory determines when a food is consumed again.

Linear Probe Penalties Reduce LLM Sycophancy Interference of the end: Why recency bias in memory determines when a food is consumed again

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.824191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.425448Z digest=sha256:22922a24fafca700ce8ccd4a5fdf3d8416c02c76db33aab4598aa874520f261f

Observation dfc429e2-4a5b-4a68-bef5-59f788c8e62a · outbound

This paper cites More than a feeling: Accuracy and application of sentiment analysis.

Linear Probe Penalties Reduce LLM Sycophancy More than a feeling: Accuracy and application of sentiment analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.428918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.428918Z digest=sha256:5975944c4c653154c0dcade8cf8456a6865622b9af7dd0b5d70872ab4fb64694

Observation 75072bf4-7d34-46c6-bbcd-23118cbf1548 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Linear Probe Penalties Reduce LLM Sycophancy Large Language Models are Zero-Shot Reasoners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.433407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.433407Z digest=sha256:e2a9f6cd58ab51c70964948cd3c04d7816c1586683915da0c42762a5db84eb44

Observation 27c2431a-ea83-4f41-b054-08c60154f045 · outbound

This paper cites Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks.

Linear Probe Penalties Reduce LLM Sycophancy Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.437344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.437344Z digest=sha256:9ab948ce6f90d1e1ceab905853b2b5641924b6ba520655242adce92d3412d8c7

Observation cd14188e-5db7-4734-b332-363ad7596ad8 · outbound

This paper cites Emergent linear representations in world models of self-supervised sequence models, 2023.

Linear Probe Penalties Reduce LLM Sycophancy Emergent linear representations in world models of self-supervised sequence models, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.813024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.441162Z digest=sha256:c0dd4cfab69051248bfcdf548868bd4d1827b0baf0efc657ce5cce8f17a25cfe

Observation c0449222-c72d-426f-8eb3-21cfed7a75e3 · outbound

This paper cites Training language models to follow instructions with human feedback.

Linear Probe Penalties Reduce LLM Sycophancy Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.444875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.444875Z digest=sha256:1120b8de4ce318ba9410d4e61de18efed448fc09ba0d77a14268e6e4a993a112

Observation 73e62d59-20a2-4170-a8cf-3153355c4d0b · outbound

This paper cites Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.

Linear Probe Penalties Reduce LLM Sycophancy Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.448384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.448384Z digest=sha256:48db67cb83af3a61816864fcaad7297e4e27c8f5c1e37ed38da8918fadb6bf98

Observation b2bec276-2a96-4c61-a557-1c69ab9b8471 · outbound

This paper cites Prioritizing High-Consequence Biological Capabilities in Evaluations of Artificial Intelligence Models.

Linear Probe Penalties Reduce LLM Sycophancy Prioritizing High-Consequence Biological Capabilities in Evaluations of Artificial Intelligence Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.452479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.452479Z digest=sha256:d532f4746f3a63d176749bf221e27c2e7dc0c94e88dc845098e9402cc86cc9de

Observation 6d66ed55-9aec-4604-b8d9-da3a17c68fa6 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Linear Probe Penalties Reduce LLM Sycophancy Discovering Language Model Behaviors with Model-Written Evaluations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.456194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.456194Z digest=sha256:08d5a0770c415bbbc69bd4fb51d097a744bb3fd68fa82ad426345e099b49bdf1

Observation 58a24925-d63e-4101-8aa4-2293805d0540 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Linear Probe Penalties Reduce LLM Sycophancy Steering Llama 2 via Contrastive Activation Addition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.459992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.459992Z digest=sha256:51efe6f5c25bb60e4cad1788ce01d9efe5c5e5441dfe4503b4bb4c331955cf5e

Observation 22b701c6-153e-4978-b58a-63b786c416cb · outbound

This paper cites Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R.

Linear Probe Penalties Reduce LLM Sycophancy Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.794428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.464532Z digest=sha256:891dca2242523659964f41d35d4df629d24267ee9d737e3d3b9594b013ee7b80

Observation c4051584-aa8f-4bdd-856a-678f4f1f69ca · outbound

This paper cites Immediate and delayed primacy and recency effects in performance evaluation.

Linear Probe Penalties Reduce LLM Sycophancy Immediate and delayed primacy and recency effects in performance evaluation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.782841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.472309Z digest=sha256:4fa6d52d60bdef18cf066f65629514816ca463d6f1389866fb78c1882fef2ea4

Observation b7a91471-2f7f-4fb4-9e85-fa3fb203032b · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

Linear Probe Penalties Reduce LLM Sycophancy Towards Understanding Sycophancy in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.468357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.468357Z digest=sha256:ac51106980e2227413a1d69fbe00a65281a09db376d033d23ac1470b492cba70

Observation 39b824f1-c9d9-4d28-8a48-a1739e65bf59 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

Linear Probe Penalties Reduce LLM Sycophancy Simple synthetic data reduces sycophancy in large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.479722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.479722Z digest=sha256:0612ba9936fe53f65769fbb531de897a339abcf9569352961bc6d1a109e9edba

Observation 31dc03a1-f3fd-411f-87d3-e659b81c1d3d · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

Linear Probe Penalties Reduce LLM Sycophancy OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.475789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.475789Z digest=sha256:27cbceba39d68c853d8dfb8684ac2af8e93c216fd88c56fb22a38d8e31b5469b

Observation f3549866-c168-4a52-9ae9-47e7247b269b · outbound

This paper cites as a judge.

Linear Probe Penalties Reduce LLM Sycophancy as a judge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.762121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.487388Z digest=sha256:1401ffc16f3e1a23f08334a16f7db461fb7223db6b685f81026a5f8903c7283b

Observation cfdd1245-6589-4265-8565-8991d0fb87cb · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023.

Linear Probe Penalties Reduce LLM Sycophancy Starling-7b: Improving llm helpfulness & harmlessness with rlaif, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.483661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.483661Z digest=sha256:80e6f1644868fe5b9dcefaf21c692cdc4da1547fecb8dc25f1c72738d15c42c0

Observation cda8f650-92d4-4d7a-b2df-01d42b317860 · outbound

This paper cites an unresolved cited work.

Linear Probe Penalties Reduce LLM Sycophancy Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:53:06.750568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.492133Z digest=sha256:d1216e16f50ec956453dc0a588e5ba4842080e98a1504ce3d1301dc4a8c7e651

Observation 9ecb7f78-a3bf-454b-9a14-2dda37db3fa4 · outbound

This paper cites (B) Positive.

Linear Probe Penalties Reduce LLM Sycophancy (B) Positive

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.738779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.495946Z digest=sha256:c5c3af86c99c3a8fcafae871cfd3c0b0dc1c6bca8b058d6314b34623d04e9da1

Observation dfc90fe1-cd57-4c9f-b170-393e68afd637 · outbound

This paper cites an unresolved cited work.

Linear Probe Penalties Reduce LLM Sycophancy Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:53:06.727473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.500197Z digest=sha256:31d071e0677d036e9617a6a1e69a5e8b78a3e2e0ced3a4b22b5979680cc25d3b

Observation 682f048f-637d-4093-9789-ae5e8c50fb52 · outbound

This paper cites The distribution is now much more balanced.

Linear Probe Penalties Reduce LLM Sycophancy The distribution is now much more balanced

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:53:06.715215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T04:53:06.503806Z digest=sha256:45c020c7765acc2fed796a25acecfd5e1ee6feb2af7980de3b7e5c17edfb1e9e

Observation 3d8a41a7-74b7-47c1-85ae-02c0e8f156ef · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Linear Probe Penalties Reduce LLM Sycophancy UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T04:53:06.401822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:53:06.401822Z digest=sha256:6f2730c6a423fcd578e355c537c14c5440507a13545bfdb3b2d207df38fbf398

Pith citing papers

Observation 36c36c31-3b73-457e-92c9-4919fbb3cb87 · inbound

LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs cites this paper.

LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs Linear Probe Penalties Reduce LLM Sycophancy

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:28:30.431171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:28:30.431171Z digest=sha256:c8d0ed57336319f99cf23f77824df332a7afc256cd4767e79290502b34c3eef3

Observation 4e3323b6-488b-486c-a837-1b926db7d428 · inbound

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework cites this paper.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Linear Probe Penalties Reduce LLM Sycophancy

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.959394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.959394Z digest=sha256:1ed6082b2c1067f5ed39d892aaed07cff092ab52f2eaff6856637a5692869548

Observation 2441dda6-046a-401f-a95b-2d7c6ac10835 · inbound

CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs cites this paper.

CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs Linear Probe Penalties Reduce LLM Sycophancy

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:09:49.308697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:09:49.308697Z digest=sha256:492aff0ad23e56bd148ca9dff929995fae1cb5aaa133df283567869fb2d372ef

Observation 3da1a77b-b09a-4827-a2df-25dcd9cec6b9 · inbound

Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition cites this paper.

Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition Linear Probe Penalties Reduce LLM Sycophancy

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:49.183370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T20:10:56.036361Z digest=sha256:2032f31ec63abf62b4267b058526abc6f924ecafd52ee46eab69e917aea3f15b

Observation 56a958e8-e3e7-4eb0-964a-754c28de3fec · inbound

Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation cites this paper.

Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation Linear Probe Penalties Reduce LLM Sycophancy

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:11.336744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T15:04:40.429929Z digest=sha256:2b02ab95b59a9968f76635f474c697cf5051caa3953b52fbfbead75d8d3d4d63

Observation 31dbaf3a-b5c2-4511-b779-91f71efdcd61 · inbound

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating cites this paper.

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Linear Probe Penalties Reduce LLM Sycophancy

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:29.998699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T17:03:33.199645Z digest=sha256:397750028d4db5684fe9337b6a36e6ff0617b207778e2ef1caae9d06385304a1

Observation f5594707-a557-43eb-a7e3-9b2adc876ff0 · inbound

Measuring and Detecting Harmful AI Sycophancy cites this paper.

Measuring and Detecting Harmful AI Sycophancy Linear Probe Penalties Reduce LLM Sycophancy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:19:12.781624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:19:12.781624Z digest=sha256:d0f08643e8e552dc4b04e360b251063bf7654175d1537ab64bb254482ea1ab34