Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T15:01:41.161487Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 78 inbound Pith citation observations for arXiv:2211.03540.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T15:01:41.161487Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:44:02.926095Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
42 of 42 outbound references displayed
External citation measurements
32
pith, observed 2026-08-05T02:28:24.338817Z
Observation e4174cb8-408e-4da0-b551-8d1f93f2ef5b · outbound
Measuring Progress on Scalable Oversight for Large Language Models The case for aligning narrowly superhuman models , url=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 43f90ad9-77ef-4b3d-8ef8-da54d4337960 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1fc12978-fdcc-4fdf-8f40-3a7f0935865b · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e61c81fe-cc1d-41c0-ba01-0f0747937c12 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Weld , journal=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6c904c71-1b06-4520-a528-11806f77ca97 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Advances in neural information processing systems , volume=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e38530dd-4572-4342-b936-787ee2a1ad4f · outbound
Measuring Progress on Scalable Oversight for Large Language Models Advances in Neural Information Processing Systems , volume=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 68d621d7-74b2-410e-a62b-39fdb16acce7 · outbound
Measuring Progress on Scalable Oversight for Large Language Models 2014 , isbn =
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7f12dfe8-a82b-4e83-ae87-a8f1fa0e797c · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0e030b36-4d66-468e-904f-da803314b8ce · outbound
Measuring Progress on Scalable Oversight for Large Language Models Organizational behavior and human performance , volume=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 56ee0fd4-2e0e-40eb-ab9f-be6117d81079 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ebf72b88-e288-4613-9b41-1c5355bd2a76 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Submitted to The Eleventh International Conference on Learning Representations , year=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 48075c22-a281-43ab-b94b-91eba9ddc028 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Concrete Problems in AI Safety
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af034535-36a8-4349-aee9-185ad063330c · outbound
Measuring Progress on Scalable Oversight for Large Language Models A General Language Assistant as a Laboratory for Alignment
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 068e9cb5-7109-4594-a000-6428372bbef8 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 63ac058c-2294-42fd-a11c-5fe3c6dd8d34 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f41ae8f-6573-400f-b128-5808eaa93162 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 770ab2e0-d5bd-4de9-b653-926b1649a85f · outbound
Measuring Progress on Scalable Oversight for Large Language Models Supervising strong learners by amplifying weak experts
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d627035e-63ff-4ce6-a56d-6bd6c823a23a · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d83291cd-833c-4417-aacd-50548c71d4c3 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2630f559-5066-4a0a-b0ac-fdb86d473c89 · outbound
Measuring Progress on Scalable Oversight for Large Language Models In: 26th Inter- national Conference on Intelligent User Interfaces, pp
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2082097a-2392-4c89-b1bc-af4e73fabce2 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ab1ac030-0194-48be-8b88-645a90aa4394 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Measuring Massive Multitask Language Understanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 77c61ac6-aff5-4770-b065-45719dbc9a48 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 219991db-4899-45af-b29f-a371f39ce014 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a16fa239-61ad-4915-8a27-069ed6c94bb6 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 616e3a07-4f6d-4b1a-b950-b71a2e5d4011 · outbound
Measuring Progress on Scalable Oversight for Large Language Models AI safety via debate
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c268239-7b4f-4d82-b546-4bdca85a5e61 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Language Models (Mostly) Know What They Know
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d9aba1e-ff68-465f-8c5b-fd7298f3886a · outbound
Measuring Progress on Scalable Oversight for Large Language Models Large Language Models are Zero-Shot Reasoners
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb04b6c5-6c29-4d52-ad20-af9f0fb2442f · outbound
Measuring Progress on Scalable Oversight for Large Language Models Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a056445e-b407-4ff3-a9f0-721f4559b0bd · outbound
Measuring Progress on Scalable Oversight for Large Language Models Bach, and Jure Leskovec
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bc09f87c-8d8a-4b89-85ed-780447feea9e · outbound
Measuring Progress on Scalable Oversight for Large Language Models Scalable agent alignment via reward modeling: a research direction
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bad5145c-a75a-4df8-b393-d975c9763547 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 253a3c6d-02c9-40ab-b123-de61f5ba7ff9 · outbound
Measuring Progress on Scalable Oversight for Large Language Models URLhttps://doi.org/10.18653/v1/2022.acl-long.229
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51b8fa22-8dd0-4bcd-b1cc-1ac700418a71 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c94a520-7597-45c9-96ba-0b3323890368 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Show Your Work: Scratchpads for Intermediate Computation with Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 90bfc13f-24b3-4db8-a064-b31d33d74640 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Q u ALITY : Question Answering with Long Input Texts, Yes!
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c03a36e3-1208-41f4-9048-799a93ffb05e · outbound
Measuring Progress on Scalable Oversight for Large Language Models Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5c48183f-f710-4029-920c-05f742cf5ff8 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Single-Turn Debate Does Not Help Humans Answer Hard Reading-Comprehension Questions
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e3ebd67e-6b71-48ec-9f1b-4f3fc098f26e · outbound
Measuring Progress on Scalable Oversight for Large Language Models Self-critiquing models for assisting human evaluators
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5d3449ba-ecfb-45eb-b5ca-99e8eb70b6d8 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fe0086a4-287f-4f52-b76b-75c9e0f80dde · outbound
Measuring Progress on Scalable Oversight for Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ccb6a507-96d6-47ff-926c-880bb6cb9849 · outbound
Measuring Progress on Scalable Oversight for Large Language Models Recursively Summarizing Books with Human Feedback
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 310a7225-23d5-403c-b16f-07a60dbc2a9a · inbound
Measuring Faithfulness in Chain-of-Thought Reasoning Measuring Progress on Scalable Oversight for Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 727345d7-2b3b-41a1-9510-ab5cc62c3528 · inbound
Simple synthetic data reduces sycophancy in large language models Measuring Progress on Scalable Oversight for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 677afb01-6fb3-4c8a-9183-1496bbf70eb3 · inbound
Llemma: An Open Language Model For Mathematics Measuring Progress on Scalable Oversight for Large Language Models
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b8b401bf-361d-43d4-993c-33504914f439 · inbound
Towards Understanding Sycophancy in Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e7055ce4-1ef6-49b0-a8c9-a3b6dba9cd98 · inbound
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions Measuring Progress on Scalable Oversight for Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8601443d-1632-4d69-907c-c0eeff835971 · inbound
TrustLLM: Trustworthiness in Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 901d418f-491f-47db-8464-82980c25dbad · inbound
A Roadmap to Pluralistic Alignment Measuring Progress on Scalable Oversight for Large Language Models
Reference 231
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2223dc72-1d69-4bda-a082-004977665a4a · inbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3e608c7d-1aef-462d-9636-e631bea76776 · inbound
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 267
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bb63dc13-39a3-4e88-93b4-b7a2e01e40bf · inbound
AI Safety Landscape for Large Language Models: Taxonomy, State-of-the-art, and Future Directions Measuring Progress on Scalable Oversight for Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c4551775-ebc5-4bce-950f-c5d9e3dcf410 · inbound
Preference Optimization for Reasoning with Pseudo Feedback Measuring Progress on Scalable Oversight for Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a6d641-dcb1-4ee2-95f4-e69ed53ebdf8 · inbound
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision Measuring Progress on Scalable Oversight for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfec8ed9-71c9-4079-821b-9a38feb8ac92 · inbound
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies Measuring Progress on Scalable Oversight for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20742aa9-b3f4-421b-8794-8abe8b5f63a3 · inbound
ProcessBench: Identifying Process Errors in Mathematical Reasoning Measuring Progress on Scalable Oversight for Large Language Models
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32a8254-b368-4624-95fa-9520e269f644 · inbound
The Superalignment of Superhuman Intelligence with Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 236f959f-c4e5-4f57-8810-375533c6a15a · inbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Measuring Progress on Scalable Oversight for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5871bfc-dba2-466e-92b8-e35000899a04 · inbound
The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment Measuring Progress on Scalable Oversight for Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72d5970-6bf3-4fb9-bedc-9c2f0c23309a · inbound
Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0909d486-5115-43f5-9e71-8c6f749473e3 · inbound
Governing AI Agents Measuring Progress on Scalable Oversight for Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75af3184-376c-4a08-a8e3-6b1f7060ada3 · inbound
Debate Helps Weak-to-Strong Generalization Measuring Progress on Scalable Oversight for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da23bdd6-e48f-4ab9-af14-37dbdf74fe4a · inbound
Automated Capability Discovery via Foundation Model Self-Exploration Measuring Progress on Scalable Oversight for Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde3f0b1-fc9c-4a43-8dc6-79aff4c43612 · inbound
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance Measuring Progress on Scalable Oversight for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5401f7-c3a6-44e1-b232-ae7e53c35dca · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Measuring Progress on Scalable Oversight for Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a86f4835-6b31-4184-9d06-8fc9cb6b7139 · inbound
Super Co-alignment of Human and AI for Sustainable Symbiotic Society Measuring Progress on Scalable Oversight for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97709941-4e7e-403b-8b0e-d25b482973ae · inbound
DeepCritic: Deliberate Critique with Large Language Models Measuring Progress on Scalable Oversight for Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a405042c-10b5-4bb6-a6f2-d243d2ed560e · inbound
Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers Measuring Progress on Scalable Oversight for Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6b4ab2-641c-4bc8-b438-0f2efaac871f · inbound
What Is AI Safety? What Do We Want It to Be? Measuring Progress on Scalable Oversight for Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa68ffa3-cc7d-4d91-b5c8-109bf5f91d26 · inbound
An alignment safety case sketch based on debate Measuring Progress on Scalable Oversight for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e35d027e-89f9-4d94-82ad-e3bb9aee02a8 · inbound
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48e7a18c-07af-44b3-b642-2af018737cc0 · inbound
Benchmarking Misuse Mitigation Against Covert Adversaries Measuring Progress on Scalable Oversight for Large Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 329a6965-01d5-4c3c-85e1-d48057160d83 · inbound
Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) Measuring Progress on Scalable Oversight for Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bde6da0-f4f3-4a8c-9f28-1afffd67c247 · inbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Progress on Scalable Oversight for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab5e77a-ee45-4595-9196-49652ab17900 · inbound
Architecting Human-AI Cocreation for Technical Services -- Interaction Modes and Contingency Factors Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d23f02-0625-4652-9ef7-650e7721348c · inbound
ADEPTS: A Capability Framework for Human-Centered Agent Design Measuring Progress on Scalable Oversight for Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10a1c5d-7669-4c92-aa99-9091e123183d · inbound
Reliable Weak-to-Strong Monitoring of LLM Agents Measuring Progress on Scalable Oversight for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17980314-f36b-4b8d-b475-9a95744ea767 · inbound
ACE and Diverse Generalization via Selective Disagreement Measuring Progress on Scalable Oversight for Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507cc51e-c822-4262-8068-1acda3b91d92 · inbound
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment Measuring Progress on Scalable Oversight for Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be9b6ac-ece3-479a-b547-8d6ba6fe8cfc · inbound
Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts Measuring Progress on Scalable Oversight for Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ca93db08-baeb-405b-b3ec-e9cc1a04fb04 · inbound
Learning When to Trust in Contextual Social Bandits Measuring Progress on Scalable Oversight for Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800a08c5-2321-45c7-838a-1a34b6af4792 · inbound
Extrapolating Volition with Recursive Information Markets Measuring Progress on Scalable Oversight for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3665eacb-86e0-463d-b7d0-64c798b4d4bc · inbound
Auditing and Controlling AI Agent Actions in Spreadsheets Measuring Progress on Scalable Oversight for Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 50419c15-4d9e-43c2-949d-5d1335d125af · inbound
Building a Precise Video Language with Human-AI Oversight Measuring Progress on Scalable Oversight for Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a05e8937-e2ad-401a-9a3e-907f6cbceaf4 · inbound
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture Measuring Progress on Scalable Oversight for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b25ff9da-9ea7-4955-b92d-3df7b37fdc41 · inbound
Agentic-imodels: Evolving agentic interpretability tools via autoresearch Measuring Progress on Scalable Oversight for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6b7f1291-242a-4bdd-9a84-c26283012f02 · inbound
Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e88a34e5-5099-4b7d-b469-1edb27a7e224 · inbound
Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 52ba6201-1c33-4492-8663-24ecf82b4e34 · inbound
Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75d69348-83be-4026-bc55-6aff42c4b851 · inbound
Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 96615e44-6514-4e9a-aaa8-e9727849299b · inbound
Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6265a63d-b96a-473b-91b3-fdccb56ba11f · inbound
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight Measuring Progress on Scalable Oversight for Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a7e5db1f-8aee-4d85-8204-4e1a350632ed · inbound
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight Measuring Progress on Scalable Oversight for Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a000633a-0eea-4ccf-98db-9324d8d58369 · inbound
LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight Measuring Progress on Scalable Oversight for Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7e6c204e-c786-4fb9-ac09-3ce41289aba7 · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Measuring Progress on Scalable Oversight for Large Language Models
Reference 211
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cbbff684-9dc2-4824-9d4a-d88844acec80 · inbound
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy Measuring Progress on Scalable Oversight for Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f17d7638-63a1-41b3-a6bd-51446d02fd38 · inbound
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy Measuring Progress on Scalable Oversight for Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8053ac46-4d95-4aba-8a75-6439a1ad7f3b · inbound
How to Interpret Agent Behavior Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 760313a2-9ef8-4878-92e5-3c63ea3fc415 · inbound
The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Measuring Progress on Scalable Oversight for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation efa228ea-f0fe-4621-a88a-75d94cf44080 · inbound
The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Measuring Progress on Scalable Oversight for Large Language Models
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada9729b-3f19-4131-b615-3ed97031409a · inbound
ReasonOps: Operator Segmentation for LLM Reasoning Traces Measuring Progress on Scalable Oversight for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b082e248-ed58-4ef8-9852-9153da2db123 · inbound
Civilizational Metamaterials: Engineering Coordination Under Capability Gradients and Structural Turbulence Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 888bdece-b8ad-438b-ba15-21aa95d4705e · inbound
When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Measuring Progress on Scalable Oversight for Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fee3e933-890f-416d-82ff-4a0a763fc890 · inbound
Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning Measuring Progress on Scalable Oversight for Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b38cc38a-d24e-42aa-8656-65f8a2ee65ed · inbound
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Measuring Progress on Scalable Oversight for Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 340197aa-715d-48ad-8231-e0ea82a8d34f · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Measuring Progress on Scalable Oversight for Large Language Models
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d21ccf50-e626-4102-b7c9-9ccce5a3d31b · inbound
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3bd557f4-6cc1-46f9-858a-d0aee14bd328 · inbound
HANSEL: Extracting Breadcrumbs from Web Agent Trajectories for Interactive Verification Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 03ece39d-ccb5-4cef-bbaa-cb848f64945b · inbound
Regulating AI: Where U.S. State Policy and HCI (Mis)align Measuring Progress on Scalable Oversight for Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f09301-6d3f-449b-852e-96cf1488e19c · inbound
Attention Limited Reward Learning Measuring Progress on Scalable Oversight for Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af3d5c1-69af-47a7-b4d8-1a4627d85359 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Measuring Progress on Scalable Oversight for Large Language Models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation efc4af27-2ee3-4a29-9127-9092046319d2 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Measuring Progress on Scalable Oversight for Large Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08d1af3-c0a9-48b4-9d10-950748120de9 · inbound
Measuring Intelligence Beyond Human Scale Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 63bcabe1-a474-4187-a067-0c06b3a8122b · inbound
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety Measuring Progress on Scalable Oversight for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d790367-a3a2-4cfc-9264-d510447bdbd1 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Measuring Progress on Scalable Oversight for Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 699c0921-412f-408c-af41-6900b077ad48 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Measuring Progress on Scalable Oversight for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9267c136-32a2-45fd-bbcc-21459b50cce3 · inbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Measuring Progress on Scalable Oversight for Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6873bfa-1982-4cee-a43a-d17336c645ae · inbound
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning Measuring Progress on Scalable Oversight for Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a26895-fad4-47a8-8063-427b7b116ce1 · inbound
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation Measuring Progress on Scalable Oversight for Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aacf96d-82b4-4b27-aabd-7c71d9a5ae21 · inbound
Training AI Scientists to Replicate Research Measuring Progress on Scalable Oversight for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.