Pith. sign in

Paper Citation Record · LEDGER

Discovering Language Model Behaviors with Model-Written Evaluations

As of 23 July 2026, this Paper Citation Record lists 22 of 22 outbound references and 50 inbound Pith citation observations for arXiv:2212.09251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.09251 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T06:20:04.523814Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T19:20:54.974570Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f1ea4aa-a5f0-47d1-874a-4c7906188796 · outbound

This paper cites Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al.

Discovering Language Model Behaviors with Model-Written Evaluations Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.662012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:3ce7069f1b6a77676848a3d8ef8d7e6eef85b0c0f6d741ffb4a1f27719cf2a49

Observation e59c4b75-dbc9-401d-aa74-dd2192767639 · outbound

This paper cites Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp.

Discovering Language Model Behaviors with Model-Written Evaluations Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.671784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:18388615d82457ea4a6c36c9c1b018038a03cccf29a0bbf4e12ae2edde79b8e0

Observation 1d2bd2d6-e0af-4fd2-b0d0-14d77d3b8c6e · outbound

This paper cites Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants.

Discovering Language Model Behaviors with Model-Written Evaluations Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.569294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:40b9032a3a30bddde1f6c5ffcb7c32f2c282fe236bae56b7177999e9fb67dc07

Observation 388c26f5-503f-4420-9a82-a1a3516d6152 · outbound

This paper cites Supervising strong learners by amplifying weak experts.

Discovering Language Model Behaviors with Model-Written Evaluations Supervising strong learners by amplifying weak experts

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:20:04.595734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:43e9b2d6568bd0cea1825f62e1c19d8f2fb11e283a6f9d5b98bc54d11a2b937a

Observation 4c00bb12-1a90-43c0-9353-70969ff015ea · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Discovering Language Model Behaviors with Model-Written Evaluations Scaling Laws for Autoregressive Generative Modeling

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:20:04.612704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:5bebd0a5ed30efa955d860c3e0fc9faf8c2a0fe576b488ed769f472aad5309c8

Observation b4449089-2d85-4a82-9608-7258a46e905f · outbound

This paper cites In NIPS Deep Learning and Representation Learning Workshop.

Discovering Language Model Behaviors with Model-Written Evaluations In NIPS Deep Learning and Representation Learning Workshop

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.676637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:20627363bcc99fbdba36bc38504cc4cfdb8c3952d1b641f2842a18c6ec9808bb

Observation 866b6f61-746c-44e5-95ae-7a5c903021d4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Discovering Language Model Behaviors with Model-Written Evaluations Scaling Laws for Neural Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:20:04.587007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:28d1a86484590d9f25de7af20c5e84505be7307c83eb565c443a98954dd78104

Observation bc6c3912-2a04-4465-a00b-fc831300ed41 · outbound

This paper cites In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 7871–7880, Online.

Discovering Language Model Behaviors with Model-Written Evaluations In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 7871–7880, Online

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.681421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:4f222b4191e11dcc9c42bdf3c84b2c14d74762ff5cb13fcc1be49245bc92ad60

Observation 3d1d4115-44fb-4fb3-8120-356ed509fff2 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Discovering Language Model Behaviors with Model-Written Evaluations UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:20:04.604492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:02cbfeeeed57b7fccdc627b81a43980578a1730949732955aaa24b001a7439fa

Observation 5b256988-9e41-45bf-93a3-e06bda154b9a · outbound

This paper cites In Advances in Neural Information Processing Systems.

Discovering Language Model Behaviors with Model-Written Evaluations In Advances in Neural Information Processing Systems

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.685921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:df9628cdb8d1404b04d62a26b438a17d47c07448ebe8e9b0ec50d8c9fd719858

Observation 0f19c396-e119-4687-adf3-cd67401cc1eb · outbound

This paper cites Timo Schick and Hinrich Schütze.

Discovering Language Model Behaviors with Model-Written Evaluations Timo Schick and Hinrich Schütze

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.690259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:7725d68b659415022adbef863e79c008e876ceaedfb900412acd5aa8d898baab

Observation ff0450c4-abc3-40fd-9f92-838a8883f979 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Discovering Language Model Behaviors with Model-Written Evaluations Finetuned Language Models Are Zero-Shot Learners

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:20:04.578392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:b2c899c1207844dd7b38cfb7c068350e84fde13a3e6d8439de032a16885f2c67

Observation 32818064-7d94-40f4-8d67-230f0e073f1a · outbound

This paper cites scaling laws.

Discovering Language Model Behaviors with Model-Written Evaluations scaling laws

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.618760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:ee420b2abac44b66643cf3180d07c50fef49b551b18639f719c2a544c7de98d2

Observation 3193d05b-70ac-490d-a720-08c84cbdc68b · outbound

This paper cites sandbagging.

Discovering Language Model Behaviors with Model-Written Evaluations sandbagging

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.623990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:bc24004bf135f48e1c5ee82fea9118867979348ce56299b68cb28690f889537c

Observation 76404ff6-1f93-48a2-8ba1-32aa53432b9e · outbound

This paper cites 19 describing the data creation task.

Discovering Language Model Behaviors with Model-Written Evaluations 19 describing the data creation task

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.629196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:1beda722ce11c4bcb15a6a62f1285b3b03fa20ce195689d79f16e012027a917c

Observation caa0f341-974e-436d-b69c-f09410155d47 · outbound

This paper cites Surround each question in blockquotes and append to the result from stage 1.

Discovering Language Model Behaviors with Model-Written Evaluations Surround each question in blockquotes and append to the result from stage 1

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.634118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:a849ee4fc67dfbc60facc5f7e5ddafe19e39067e1303468a3a0ed3adce5ae287

Observation a43fb6f7-38e6-4beb-ae66-2a6aa2ee33eb · outbound

This paper cites Is the above a good question to ask?.

Discovering Language Model Behaviors with Model-Written Evaluations Is the above a good question to ask?

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.638599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:e22f087dfabc1cf3999eeb8c0a5057fefb532aef7e402b2ddb6c73711b3d1ade

Observation 2bad01b8-366d-4225-b60c-36c0540af87d · outbound

This paper cites an unresolved cited work.

Discovering Language Model Behaviors with Model-Written Evaluations Unresolved cited work

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.643520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:bc764e49ba9876d8fd0e4b6fe95a0f244a48af2983430a5408b24908800b2d64

Observation a34347ba-bcb9-4aa9-8477-3d7c26c969d4 · outbound

This paper cites an unresolved cited work.

Discovering Language Model Behaviors with Model-Written Evaluations Unresolved cited work

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.648310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:f19949da4d8bf95855ef311d66ccc337d6757b9399fd8e5e8afbc45ffa309040

Observation 7eaee57b-b64e-45a5-aab6-b9e2b9844917 · outbound

This paper cites he/she/they.

Discovering Language Model Behaviors with Model-Written Evaluations he/she/they

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.652945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:85cec704cc74b81305caa4ae34580386721ff18eca986e9371446dc46a3a5954

Observation ebd52f59-1d01-4d46-9a3c-4acbd152b6b7 · outbound

This paper cites an unresolved cited work.

Discovering Language Model Behaviors with Model-Written Evaluations Unresolved cited work

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.657373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:fe7c6b358c8262921fd60fbfd046df309c78aad0f5409ee39017a4ab58f7e93f

Observation 27cfcd0f-a7df-4eb8-8726-2c1663784813 · outbound

This paper cites Directors, religious activities and education.

Discovering Language Model Behaviors with Model-Written Evaluations Directors, religious activities and education

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:20:04.667157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:20:04.523814Z digest=sha256:1c5f7a389095d1b6e6eabe155dfe8176ddb6891bfe4e272f7c9f63ee81a0e645

Pith citing papers

Observation da7f9794-a52f-4613-8c4a-b0e9bb2d28b2 · inbound

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting cites this paper.

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting Discovering Language Model Behaviors with Model-Written Evaluations

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:01:19.838726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:14892967e2363451f9ea65ff6448e2960885eba8347039e14ef77fa2854ca9df

Observation 328c4a2b-14b6-4157-a794-464028daa068 · inbound

Simple synthetic data reduces sycophancy in large language models cites this paper.

Simple synthetic data reduces sycophancy in large language models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:48:08.648099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-16T14:48:08.508109Z digest=sha256:cd8bb1151b7c8422dbb0cf180407790ff94052cd527c2d18f9ce4f017751d71f

Observation 4fcd8c1d-56c7-43fc-a912-54916e035bf8 · inbound

Steering Llama 2 via Contrastive Activation Addition cites this paper.

Steering Llama 2 via Contrastive Activation Addition Discovering Language Model Behaviors with Model-Written Evaluations

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-11T20:37:20.408376Z digest=sha256:ef0abfbe9fddfee5719871434677400828d045c4048f4fc578aa25975f47fc3e

Observation 2350b892-3ea6-4f1e-9b5f-bc481c8eb5c7 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:17:08.353781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:366d210d0e23c6ec41573d74812defb26db78eb04affaf8270c7f5ca70ca7087

Observation dc0518fe-76dd-4026-aaac-c51d61218bd7 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment Discovering Language Model Behaviors with Model-Written Evaluations

Reference 148

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:37:53.409241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-07-11T11:50:26.030339Z digest=sha256:0cbbfee557ef1d1e9ea6f1f027d0c10d33457a486d0729d1046a0aa43bdaec3f

Observation 3875bbfc-817b-4fdd-8631-49c9bc61ab78 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 282

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:37.141206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:aa635a3bbd23dc31ce27ce6b62fcb0606f7882747393eac72dfa8d81561ddb62

Observation f6c7dc91-efe5-455d-8128-3832e33e9a07 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Discovering Language Model Behaviors with Model-Written Evaluations

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:560bc8bba7cca89a583115863b79e26d38e3dbb5782e5d77464e4c799f07c1c1

Observation 7298320e-6c5f-45a9-834b-ef16d0015afc · inbound

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs cites this paper.

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs Discovering Language Model Behaviors with Model-Written Evaluations

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:07:15.425691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-19T11:03:59.222849Z digest=sha256:69e4861e1f8e06b45378dae6efac3ea2aa26967c9efb760ac2e0ff246169f88d

Observation 615f375e-97b1-453d-9351-f9788e9ffb9e · inbound

Simulating the Evolution of Alignment and Values in Machine Intelligence cites this paper.

Simulating the Evolution of Alignment and Values in Machine Intelligence Discovering Language Model Behaviors with Model-Written Evaluations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T20:15:46.311347Z digest=sha256:e09f36bef51422da04f45da2fe64324a24cf8e37b8e55dfbd5b593d148522982

Observation 7e6afbc8-b483-4ce4-8fa5-3f6c43c9eb6c · inbound

Distributed Interpretability and Control for Large Language Models cites this paper.

Distributed Interpretability and Control for Large Language Models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T19:15:43.444508Z digest=sha256:82a71365e32b1998e34cbc812f06b24f0a7b86abd2d46eedc5996d907f228f77

Observation 5976ee83-34ba-45ae-ace2-7818f297e1f2 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures Discovering Language Model Behaviors with Model-Written Evaluations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:570213fa6f51efd7cb691701138866f623be318e0313d4d59b25f2c6f8678672

Observation 0680d878-32b9-4a58-acde-b76e3bfd44e1 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures Discovering Language Model Behaviors with Model-Written Evaluations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:6a83338f0b83e646a714af7aaeb49bb4b11f262dc39f0aa41dc10c8894d13fa3

Observation 62b5e6cf-9bf9-4388-895b-bf7f0f7d69b5 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:f41fc6d20519aa3ab5557cab0b4f3a2431ac78ee851299233e9ef63b6debe89f

Observation f7842cd5-bce5-47c4-bbec-80b7ed6c991a · inbound

IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development cites this paper.

IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development Discovering Language Model Behaviors with Model-Written Evaluations

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-13T23:58:51.460784Z digest=sha256:add15911b65c200aba5ba741ce35971625645c13c932a69de97edca12ec88249

Observation f297dd5e-4af3-4397-bf53-1a1a59495662 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Discovering Language Model Behaviors with Model-Written Evaluations

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:fe22ccaadbedc6430e7eaf8b3da805dac94febddce2bd6a01d4698effb157d95

Observation 02758d1a-5e8c-4025-9149-44593a9c6314 · inbound

M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation cites this paper.

M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation Discovering Language Model Behaviors with Model-Written Evaluations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-14T22:36:22.123368Z digest=sha256:dff684d68c559987dd40e2aa5edf69c2ab8c34968341acb2dc86dbb95b33678a

Observation 424068fc-8435-4618-ba86-9a22c50f30e8 · inbound

Measuring Opinion Bias and Sycophancy via LLM-based Persuasion cites this paper.

Measuring Opinion Bias and Sycophancy via LLM-based Persuasion Discovering Language Model Behaviors with Model-Written Evaluations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-09T21:48:37.940162Z digest=sha256:65c48ffb10a1a16013b2208c684a76de66d592a1a664ad28a94e8dacb2282fa7

Observation 5a2b8dd8-ec9b-4cb6-bf06-ab71d46b6c2c · inbound

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor cites this paper.

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor Discovering Language Model Behaviors with Model-Written Evaluations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T06:35:59.474625Z digest=sha256:1a2d8c060f8366560e3c595e7c3de675464c5c09619a25bfff467251f270b7a2

Observation 6a7e15d3-a903-4883-8f56-f0a3c627dbaa · inbound

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning cites this paper.

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning Discovering Language Model Behaviors with Model-Written Evaluations

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T15:52:43.274993Z digest=sha256:428565da41f966cfbdc70e3ef46595994d11adea4c51fe173e1827a89efb967e

Observation 05cc5bff-365d-4c3a-8b91-2104ad0b219b · inbound

Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes cites this paper.

Pairwise matrices for sparse autoencoders: single-feature inspection mislabels causal axes Discovering Language Model Behaviors with Model-Written Evaluations

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T18:52:01.624279Z digest=sha256:d1d95eb59876f834203777b3336ae26675f42ecb21b9b3a5bc7ce85e653053a0

Observation 654cf425-b5f1-4afa-967f-cd1dd45537e4 · inbound

Exploring the "Banality" of Deception in Generative AI cites this paper.

Exploring the "Banality" of Deception in Generative AI Discovering Language Model Behaviors with Model-Written Evaluations

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T01:09:10.548791Z digest=sha256:0179bff569e9400b8d1b62b0f816f91244b24a77925ac72de5e2d4619228bad1

Observation c94cde44-13ee-4091-8d1a-df207e0fcd28 · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Discovering Language Model Behaviors with Model-Written Evaluations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:3bd58ce1436fa8b512d2340ab1e284b497600de51f55b5800f0986352f792805

Observation ed629ba0-fc73-4127-b324-04434fef0ec8 · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Discovering Language Model Behaviors with Model-Written Evaluations

Reference 154

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:cc0942b1bfdf05eb76a82fc7e5e2447d46f286fe3a794d4e295be7685bc8e7d6

Observation bc8c6a96-be6b-4c21-9194-620015f320db · inbound

Overtrained, Not Misaligned cites this paper.

Overtrained, Not Misaligned Discovering Language Model Behaviors with Model-Written Evaluations

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-13T06:45:52.544674Z digest=sha256:1c59d572817b92d08192ebe4a724e5b5c13ffedc4e0f034fe3ca1704403ad65f

Observation c5a55b48-47ae-4819-841e-45834a1a8978 · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Discovering Language Model Behaviors with Model-Written Evaluations

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:c39634aa822d26b3f712d79d68b1e2c1a2a8753a468b753328e58ab3c7380853

Observation ecac7ec3-4164-4831-aa83-21aeb958956b · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Discovering Language Model Behaviors with Model-Written Evaluations

Reference 209

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:59a6c9e860e06d9fc015097cb40869a01934ac62112ed7fa337a92b90656680b

Observation a1ba2909-ba42-4664-8252-8656cee99999 · inbound

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs cites this paper.

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs Discovering Language Model Behaviors with Model-Written Evaluations

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:20:04.692181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-14T20:20:04.306202Z digest=sha256:7d1f42394e3ea8bc3652861c13811e1c51ecc3dcfdfe466ff334b179b4100ed6

Observation a25ecfbc-9c10-4dc0-a248-df4a3a93c4f3 · inbound

Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy cites this paper.

Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy Discovering Language Model Behaviors with Model-Written Evaluations

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T04:33:57.956246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-21T04:32:36.723049Z digest=sha256:34c9aa44dc81e0396436928fe4306828b0f46ed0b8a7069c0d0ea83b9346dddc

Observation 81bc80ee-f4a9-4415-b269-2cec00e741c1 · inbound

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most cites this paper.

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Discovering Language Model Behaviors with Model-Written Evaluations

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T05:24:38.187939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-22T05:24:10.513564Z digest=sha256:8b0cbe38396a5fc570f266353c1dfcba75dab39d4bc222de9b1766f0144420ab

Observation c265a198-3f9f-4d5e-b03d-84a13170b8d6 · inbound

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most cites this paper.

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Discovering Language Model Behaviors with Model-Written Evaluations

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T06:05:26.146399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-25T06:04:28.277835Z digest=sha256:215dff64167063515d382f0ff34a91cbd33302f31d72d60b4cbe938d2dbfb8fc

Observation 039db126-dd31-4b85-9c37-6c2539ab59a8 · inbound

AMEL: Accumulated Message Effects on LLM Judgments cites this paper.

AMEL: Accumulated Message Effects on LLM Judgments Discovering Language Model Behaviors with Model-Written Evaluations

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:11:06.833222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-22T05:08:30.607268Z digest=sha256:36e6940a2522408fd8d859632c7509501f20bfc092d6d8bd469bbb1f0b148feb

Observation 86873db7-8d9d-436a-b3f7-a2a6c84e8d3b · inbound

AMEL: Accumulated Message Effects on LLM Judgments cites this paper.

AMEL: Accumulated Message Effects on LLM Judgments Discovering Language Model Behaviors with Model-Written Evaluations

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:04:56.431422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T17:04:22.688250Z digest=sha256:0f8abc09c101bc0cac7a5daebe9a4404e6c2bc11ddc6c25fdf616dc779f27ae2

Observation 40f4787f-b372-4797-900c-7bac02656c03 · inbound

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study cites this paper.

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study Discovering Language Model Behaviors with Model-Written Evaluations

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:10:22.397990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-25T05:07:12.832683Z digest=sha256:61b4ab0c0298c4bbc435ce538808c02daafe3ddf8b86bc9085c93044793d7cdb

Observation 36d620a4-f8c4-43ea-b1f2-ac1f8c2f8fc7 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Discovering Language Model Behaviors with Model-Written Evaluations

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:24:50.169567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:203584c50bbd61a3c342d5a25fbecfc3d8781f194bdd57b232edf5aecf303ef4

Observation e49d709d-809e-452a-9c4a-5805cad2a173 · inbound

KARMA: Karma-Aligned Reward Model Adaptation cites this paper.

KARMA: Karma-Aligned Reward Model Adaptation Discovering Language Model Behaviors with Model-Written Evaluations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:53:46.462474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-29T17:52:06.830441Z digest=sha256:fbcb820ade2e13de246989fb9339ce7d8e75637c509d8da85e52d1b1d15a29cd

Observation c37a6ecd-3af9-4d03-9778-e922c79b5ff7 · inbound

Toward Agentic Governance: What Shapes LLM-Agent Intervention in Public Forums? cites this paper.

Toward Agentic Governance: What Shapes LLM-Agent Intervention in Public Forums? Discovering Language Model Behaviors with Model-Written Evaluations

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:42:37.875395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T18:21:31.715046Z digest=sha256:c786e44d5dadbe3f8284bacbea9f3009cbd75fd566e6f47b34327988bd19a708

Observation dece125e-a744-4bd5-8b88-9921eec80927 · inbound

The Self-Correction Illusion: LLMs Correct Others but Not Themselves cites this paper.

The Self-Correction Illusion: LLMs Correct Others but Not Themselves Discovering Language Model Behaviors with Model-Written Evaluations

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-28T01:31:29.336519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T01:25:07.890796Z digest=sha256:c817c4da6bd3579295874e513664b975871898bf4c447b5926d7ce6996714ee0

Observation 6019ec37-2dcc-4c46-9b3b-ee68716bcaf5 · inbound

What Do People Actually Want From AI? Mapping Preference Plurality cites this paper.

What Do People Actually Want From AI? Mapping Preference Plurality Discovering Language Model Behaviors with Model-Written Evaluations

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:41:29.757995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-28T01:32:04.660400Z digest=sha256:f35739e0df55c6b58726a8cbf7938969536ec55d3d1e0dcd33e82d9b0c6a65e1

Observation 5ed35cb1-0dba-48c6-a6b9-e8bf75a5472a · inbound

Emergent alignment and the projectability of ethical personas cites this paper.

Emergent alignment and the projectability of ethical personas Discovering Language Model Behaviors with Model-Written Evaluations

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:47:31.790778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-27T16:16:48.545212Z digest=sha256:a08f4ec2f6a46a6bd525bb14a6077f5b14dc576f5d39f2da5c4c8f7d8647bc5c

Observation 1a351f00-4a3b-496a-94ec-ce6b6d6f2e7e · inbound

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models cites this paper.

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:09.970185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-07-04T19:46:24.142611Z digest=sha256:5870bfbc7bd0fd0272a91480cd0245b999a8e530f6b3d5871f4f48f3f5640807

Observation 62765607-ec45-4618-bcf1-32a7e0590963 · inbound

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis cites this paper.

LLM-as-an-Investigator: Evidence-First Reasoning for Robust Interactive Problem Diagnosis Discovering Language Model Behaviors with Model-Written Evaluations

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:38:29.240998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-27T06:58:42.823851Z digest=sha256:86a597806971018b16e5d39000e4aa522e532579eb81c43eb70147c1b31019e8

Observation 7c1977a9-9684-4d62-8e8f-15d69dcb38d5 · inbound

Channel Location Constrains the Auditability of Subliminal Learning cites this paper.

Channel Location Constrains the Auditability of Subliminal Learning Discovering Language Model Behaviors with Model-Written Evaluations

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.405788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-26T11:52:03.948568Z digest=sha256:f8364270ee9bd4ee10f98f2804e4fa6901dba9b1adac7b65173fca49f54d6a35

Observation d4581689-7660-4cd5-9c4f-1fef5252ccf3 · inbound

Reinforcement Learning Towards Broadly and Persistently Beneficial Models cites this paper.

Reinforcement Learning Towards Broadly and Persistently Beneficial Models Discovering Language Model Behaviors with Model-Written Evaluations

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T11:39:46.788198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-26T07:51:13.283619Z digest=sha256:c8d9db7491a07d091621817abfc5fa6c1c6d0dc6d96e27222b4eca087799883a

Observation 7175abbf-a9d0-4319-94f4-c8f8986adfc0 · inbound

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs cites this paper.

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs Discovering Language Model Behaviors with Model-Written Evaluations

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T17:09:58.913742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-25T23:54:36.233158Z digest=sha256:04dfc26c5c1d19d24d5f81c7bc8b9917ae0323df5facb8c66b7c67acaa1961f2

Observation eeb9ca6d-d17b-49da-b8cb-7ef3cd874782 · inbound

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training cites this paper.

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Discovering Language Model Behaviors with Model-Written Evaluations

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:25:33.144427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-01T08:21:41.008505Z digest=sha256:b810699f2ed8be255f05e20ba0a1b8dd211e5e7f63383c2358e0cd6f773536b8

Observation c7d20993-6739-4b22-b4cf-a59045cdf089 · inbound

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training cites this paper.

Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training Discovering Language Model Behaviors with Model-Written Evaluations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T19:20:54.974570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:20:54.974570Z digest=sha256:07466cbe58d503b91791bbb97b1a07b088f5a913fbf3630b3678cc0ef8b2d3d0

Observation 372ae0d1-3618-44c4-a602-3a53fd9e776c · inbound

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety cites this paper.

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety Discovering Language Model Behaviors with Model-Written Evaluations

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T00:52:55.781827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-29T00:46:03.210076Z digest=sha256:4cd03cf1d9313c6574a9b1d7ceccb1f299b56e3b64ebdd6b99141c574e6e2eb0

Observation 5ca4b640-4c2f-411a-8d2f-5cf5e05cb288 · inbound

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents cites this paper.

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents Discovering Language Model Behaviors with Model-Written Evaluations

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:44:41.466459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T05:48:15.990443Z digest=sha256:49b09513b1bea5f409811811c3bac94c3c3eeb531e3836faaefca53e44ee2f22

Observation 1602572c-801e-402a-a547-396e3f07afc9 · inbound

Dissociating the Internal Representations of Sycophancy in LLMs cites this paper.

Dissociating the Internal Representations of Sycophancy in LLMs Discovering Language Model Behaviors with Model-Written Evaluations

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:06:35.306164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-09T21:56:32.778349Z digest=sha256:254794ce868ced45d8dc2f71abde15fc670cec949bafbe6878e83f6d4e592e0c

Observation 5900e437-b799-4432-b4ef-4f2ff79cf683 · inbound

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring cites this paper.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring Discovering Language Model Behaviors with Model-Written Evaluations

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:56:41.052542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-07-10T00:52:47.537142Z digest=sha256:72e6d03c1a2f5c0c2e346cf9f747e418a286ce01e1a9d7b6c02d8892e8aea78e