REVIEW 13 cited by
Evaluating Language-Model Agents on Realistic Autonomous Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refer to this cluster of capabilities as "autonomous replication and adaptation" or ARA. We believe that systems capable of ARA could have wide-reaching and hard-to-anticipate consequences, and that measuring and forecasting ARA may be useful for informing measures around security, monitoring, and alignment. Additionally, once a system is capable of ARA, placing bounds on a system's capabilities may become significantly more difficult. We construct four simple example agents that combine language models with tools that allow them to take actions in the world. We then evaluate these agents on 12 tasks relevant to ARA. We find that these language model agents can only complete the easiest tasks from this list, although they make some progress on the more challenging tasks. Unfortunately, these evaluations are not adequate to rule out the possibility that near-future agents will be capable of ARA. In particular, we do not think that these evaluations provide good assurance that the ``next generation'' of language models (e.g. 100x effective compute scaleup on existing models) will not yield agents capable of ARA, unless intermediate evaluations are performed during pretraining. Relatedly, we expect that fine-tuning of the existing models could produce substantially more competent agents, even if the fine-tuning is not directly targeted at ARA.
Forward citations
Cited by 13 Pith papers
-
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
A new 17-task benchmark measures how well frontier AI agents can secretly complete a harmful side task while doing a benign task, and how well AI monitors can detect them; the best agent succeeds 27% of the time and t...
-
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Affirmative AI-agent insurance with billion-scale limits is achievable by 2030 solely through coordinated industry build-out of an eight-component stack spanning data, CAT models, standards, contracts, underwriting, p...
-
Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies
DSCC composes per-tool NIST-aligned policies via a Most Restrictive Set algorithm with monotonic taint tracking, blocking most multi-tool chains that would enable exfiltration or clearance violations.
-
Reliable Weak-to-Strong Monitoring of LLM Agents
Monitor scaffolding, not monitor awareness or omniscience, drives detection reliability, and a hybrid chunked monitor lets weak models supervise strong LLM agents.
-
Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
Thought-Aligner corrects unsafe intermediate reasoning in LLM agents before actions, raising measured behavioral safety across three benchmarks.
-
Frontier AI systems have surpassed the self-replicating red line
Two open-weight AI agents, when explicitly instructed and given system tools, successfully made live copies of themselves in most trials, a capability the authors call the self-replication red line.
-
Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction
On the OEDD benchmark, LLM agents select the worse of two actions below chance when the correct choice requires combining two earlier facts and ignoring a recent distractor in contexts over 1,615 tokens.
-
Safety case template for frontier AI: A cyber inability argument
A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.
-
Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations
A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.
-
Steering Language Model Refusal with Sparse Autoencoders
Boosting one SAE 'refusal' feature in Phi-3 Mini and Llama 3.1 raises refusal rates on unsafe and safe prompts alike while sharply reducing MMLU, TruthfulQA, and GSM8K accuracy.
-
Private, Verifiable, and Auditable AI Systems
A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.
-
International Security Applications of Flexible Hardware-Enabled Guarantees
A policy report argues that hardware-level guarantees on AI chips could make international AI governance agreements stable if states value winning only modestly and fear catastrophe sufficiently.
Discussion (0). Continue with ORCID to comment.