Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:20:04.523814Z
Paper Citation Record · LEDGER
As of 23 July 2026, this Paper Citation Record lists 22 of 22 outbound references and 50 inbound Pith citation observations for arXiv:2212.09251.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:20:04.523814Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T19:20:54.974570Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2f1ea4aa-a5f0-47d1-874a-4c7906188796 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e59c4b75-dbc9-401d-aa74-dd2192767639 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1d2bd2d6-e0af-4fd2-b0d0-14d77d3b8c6e · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 388c26f5-503f-4420-9a82-a1a3516d6152 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Supervising strong learners by amplifying weak experts
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4c00bb12-1a90-43c0-9353-70969ff015ea · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Scaling Laws for Autoregressive Generative Modeling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b4449089-2d85-4a82-9608-7258a46e905f · outbound
Discovering Language Model Behaviors with Model-Written Evaluations In NIPS Deep Learning and Representation Learning Workshop
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 866b6f61-746c-44e5-95ae-7a5c903021d4 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Scaling Laws for Neural Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation bc6c3912-2a04-4465-a00b-fc831300ed41 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 7871–7880, Online
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3d1d4115-44fb-4fb3-8120-356ed509fff2 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5b256988-9e41-45bf-93a3-e06bda154b9a · outbound
Discovering Language Model Behaviors with Model-Written Evaluations In Advances in Neural Information Processing Systems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0f19c396-e119-4687-adf3-cd67401cc1eb · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Timo Schick and Hinrich Schütze
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ff0450c4-abc3-40fd-9f92-838a8883f979 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Finetuned Language Models Are Zero-Shot Learners
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 32818064-7d94-40f4-8d67-230f0e073f1a · outbound
Discovering Language Model Behaviors with Model-Written Evaluations scaling laws
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3193d05b-70ac-490d-a720-08c84cbdc68b · outbound
Discovering Language Model Behaviors with Model-Written Evaluations sandbagging
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 76404ff6-1f93-48a2-8ba1-32aa53432b9e · outbound
Discovering Language Model Behaviors with Model-Written Evaluations 19 describing the data creation task
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation caa0f341-974e-436d-b69c-f09410155d47 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Surround each question in blockquotes and append to the result from stage 1
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a43fb6f7-38e6-4beb-ae66-2a6aa2ee33eb · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Is the above a good question to ask?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2bad01b8-366d-4225-b60c-36c0540af87d · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a34347ba-bcb9-4aa9-8477-3d7c26c969d4 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7eaee57b-b64e-45a5-aab6-b9e2b9844917 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations he/she/they
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ebd52f59-1d01-4d46-9a3c-4acbd152b6b7 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 27cfcd0f-a7df-4eb8-8726-2c1663784813 · outbound
Discovering Language Model Behaviors with Model-Written Evaluations Directors, religious activities and education
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation da7f9794-a52f-4613-8c4a-b0e9bb2d28b2 · inbound
Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting Discovering Language Model Behaviors with Model-Written Evaluations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 328c4a2b-14b6-4157-a794-464028daa068 · inbound
Simple synthetic data reduces sycophancy in large language models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4fcd8c1d-56c7-43fc-a912-54916e035bf8 · inbound
Steering Llama 2 via Contrastive Activation Addition Discovering Language Model Behaviors with Model-Written Evaluations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2350b892-3ea6-4f1e-9b5f-bc481c8eb5c7 · inbound
TrustLLM: Trustworthiness in Large Language Models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation dc0518fe-76dd-4026-aaac-c51d61218bd7 · inbound
A Roadmap to Pluralistic Alignment Discovering Language Model Behaviors with Model-Written Evaluations
Reference 148
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3875bbfc-817b-4fdd-8631-49c9bc61ab78 · inbound
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 282
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f6c7dc91-efe5-455d-8128-3832e33e9a07 · inbound
Humanity's Last Exam Discovering Language Model Behaviors with Model-Written Evaluations
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7298320e-6c5f-45a9-834b-ef16d0015afc · inbound
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs Discovering Language Model Behaviors with Model-Written Evaluations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 615f375e-97b1-453d-9351-f9788e9ffb9e · inbound
Simulating the Evolution of Alignment and Values in Machine Intelligence Discovering Language Model Behaviors with Model-Written Evaluations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7e6afbc8-b483-4ce4-8fa5-3f6c43c9eb6c · inbound
Distributed Interpretability and Control for Large Language Models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5976ee83-34ba-45ae-ace2-7818f297e1f2 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures Discovering Language Model Behaviors with Model-Written Evaluations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0680d878-32b9-4a58-acde-b76e3bfd44e1 · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures Discovering Language Model Behaviors with Model-Written Evaluations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b5e6cf-9bf9-4388-895b-bf7f0f7d69b5 · inbound
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f7842cd5-bce5-47c4-bbec-80b7ed6c991a · inbound
IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development Discovering Language Model Behaviors with Model-Written Evaluations
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f297dd5e-4af3-4397-bf53-1a1a59495662 · inbound
QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Discovering Language Model Behaviors with Model-Written Evaluations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 02758d1a-5e8c-4025-9149-44593a9c6314 · inbound
M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation Discovering Language Model Behaviors with Model-Written Evaluations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 424068fc-8435-4618-ba86-9a22c50f30e8 · inbound
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion Discovering Language Model Behaviors with Model-Written Evaluations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5a2b8dd8-ec9b-4cb6-bf06-ab71d46b6c2c · inbound
Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor Discovering Language Model Behaviors with Model-Written Evaluations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6a7e15d3-a903-4883-8f56-f0a3c627dbaa · inbound
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning Discovering Language Model Behaviors with Model-Written Evaluations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 05cc5bff-365d-4c3a-8b91-2104ad0b219b · inbound
Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes Discovering Language Model Behaviors with Model-Written Evaluations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 654cf425-b5f1-4afa-967f-cd1dd45537e4 · inbound
Exploring the "Banality" of Deception in Generative AI Discovering Language Model Behaviors with Model-Written Evaluations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c94cde44-13ee-4091-8d1a-df207e0fcd28 · inbound
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Discovering Language Model Behaviors with Model-Written Evaluations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ed629ba0-fc73-4127-b324-04434fef0ec8 · inbound
Positive Alignment: Artificial Intelligence for Human Flourishing Discovering Language Model Behaviors with Model-Written Evaluations
Reference 154
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation bc8c6a96-be6b-4c21-9194-620015f320db · inbound
Overtrained, Not Misaligned Discovering Language Model Behaviors with Model-Written Evaluations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c5a55b48-47ae-4819-841e-45834a1a8978 · inbound
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Discovering Language Model Behaviors with Model-Written Evaluations
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ecac7ec3-4164-4831-aa83-21aeb958956b · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Discovering Language Model Behaviors with Model-Written Evaluations
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a1ba2909-ba42-4664-8252-8656cee99999 · inbound
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs Discovering Language Model Behaviors with Model-Written Evaluations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a25ecfbc-9c10-4dc0-a248-df4a3a93c4f3 · inbound
Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy Discovering Language Model Behaviors with Model-Written Evaluations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 81bc80ee-f4a9-4415-b269-2cec00e741c1 · inbound
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Discovering Language Model Behaviors with Model-Written Evaluations
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c265a198-3f9f-4d5e-b03d-84a13170b8d6 · inbound
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Discovering Language Model Behaviors with Model-Written Evaluations
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 039db126-dd31-4b85-9c37-6c2539ab59a8 · inbound
AMEL: Accumulated Message Effects on LLM Judgments Discovering Language Model Behaviors with Model-Written Evaluations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 86873db7-8d9d-436a-b3f7-a2a6c84e8d3b · inbound
AMEL: Accumulated Message Effects on LLM Judgments Discovering Language Model Behaviors with Model-Written Evaluations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 40f4787f-b372-4797-900c-7bac02656c03 · inbound
Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study Discovering Language Model Behaviors with Model-Written Evaluations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 36d620a4-f8c4-43ea-b1f2-ac1f8c2f8fc7 · inbound
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Discovering Language Model Behaviors with Model-Written Evaluations
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e49d709d-809e-452a-9c4a-5805cad2a173 · inbound
KARMA: Karma-Aligned Reward Model Adaptation Discovering Language Model Behaviors with Model-Written Evaluations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c37a6ecd-3af9-4d03-9778-e922c79b5ff7 · inbound
Toward Agentic Governance: What Shapes LLM-Agent Intervention in Public Forums? Discovering Language Model Behaviors with Model-Written Evaluations
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation dece125e-a744-4bd5-8b88-9921eec80927 · inbound
The Self-Correction Illusion: LLMs Correct Others but Not Themselves Discovering Language Model Behaviors with Model-Written Evaluations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6019ec37-2dcc-4c46-9b3b-ee68716bcaf5 · inbound
What Do People Actually Want From AI? Mapping Preference Plurality Discovering Language Model Behaviors with Model-Written Evaluations
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5ed35cb1-0dba-48c6-a6b9-e8bf75a5472a · inbound
Emergent alignment and the projectability of ethical personas Discovering Language Model Behaviors with Model-Written Evaluations
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1a351f00-4a3b-496a-94ec-ce6b6d6f2e7e · inbound
An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 62765607-ec45-4618-bcf1-32a7e0590963 · inbound
LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis Discovering Language Model Behaviors with Model-Written Evaluations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7c1977a9-9684-4d62-8e8f-15d69dcb38d5 · inbound
Channel Location Constrains the Auditability of Subliminal Learning Discovering Language Model Behaviors with Model-Written Evaluations
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d4581689-7660-4cd5-9c4f-1fef5252ccf3 · inbound
Reinforcement Learning Towards Broadly and Persistently Beneficial Models Discovering Language Model Behaviors with Model-Written Evaluations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7175abbf-a9d0-4319-94f4-c8f8986adfc0 · inbound
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs Discovering Language Model Behaviors with Model-Written Evaluations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation eeb9ca6d-d17b-49da-b8cb-7ef3cd874782 · inbound
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Discovering Language Model Behaviors with Model-Written Evaluations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c7d20993-6739-4b22-b4cf-a59045cdf089 · inbound
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Discovering Language Model Behaviors with Model-Written Evaluations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372ae0d1-3618-44c4-a602-3a53fd9e776c · inbound
Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety Discovering Language Model Behaviors with Model-Written Evaluations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5ca4b640-4c2f-411a-8d2f-5cf5e05cb288 · inbound
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents Discovering Language Model Behaviors with Model-Written Evaluations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1602572c-801e-402a-a547-396e3f07afc9 · inbound
Dissociating the Internal Representations of Sycophancy in LLMs Discovering Language Model Behaviors with Model-Written Evaluations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5900e437-b799-4432-b4ef-4f2ff79cf683 · inbound
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring Discovering Language Model Behaviors with Model-Written Evaluations
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.