Pith. sign in

REVIEW 13 cited by

Evaluating Language-Model Agents on Realistic Autonomous Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11671 v2 pith:ZIAW3XPS submitted 2023-12-18 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords agentscapablelanguagemodelstasksevaluationsautonomouscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refer to this cluster of capabilities as "autonomous replication and adaptation" or ARA. We believe that systems capable of ARA could have wide-reaching and hard-to-anticipate consequences, and that measuring and forecasting ARA may be useful for informing measures around security, monitoring, and alignment. Additionally, once a system is capable of ARA, placing bounds on a system's capabilities may become significantly more difficult. We construct four simple example agents that combine language models with tools that allow them to take actions in the world. We then evaluate these agents on 12 tasks relevant to ARA. We find that these language model agents can only complete the easiest tasks from this list, although they make some progress on the more challenging tasks. Unfortunately, these evaluations are not adequate to rule out the possibility that near-future agents will be capable of ARA. In particular, we do not think that these evaluations provide good assurance that the ``next generation'' of language models (e.g. 100x effective compute scaleup on existing models) will not yield agents capable of ARA, unless intermediate evaluations are performed during pretraining. Relatedly, we expect that fine-tuning of the existing models could produce substantially more competent agents, even if the fine-tuning is not directly targeted at ARA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 15 citations worldwide. Full citation record

  1. SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

    cs.AI 2025-06 conditional novelty 7.0 of 10

    A new 17-task benchmark measures how well frontier AI agents can secretly complete a harmful side task while doing a benign task, and how well AI monitors can detect them; the best agent succeeds 27% of the time and t...

  2. Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Affirmative AI-agent insurance with billion-scale limits is achievable by 2030 solely through coordinated industry build-out of an eight-component stack spanning data, CAT models, standards, contracts, underwriting, p...

  3. Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies

    cs.CR 2026-07 conditional novelty 6.0 of 10

    DSCC composes per-tool NIST-aligned policies via a Most Restrictive Set algorithm with monotonic taint tracking, blocking most multi-tool chains that would enable exfiltration or clearance violations.

  4. Reliable Weak-to-Strong Monitoring of LLM Agents

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Monitor scaffolding, not monitor awareness or omniscience, drives detection reliability, and a hybrid chunked monitor lets weak models supervise strong LLM agents.

  5. Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

    cs.AI 2025-05 conditional novelty 6.0 of 10

    Thought-Aligner corrects unsafe intermediate reasoning in LLM agents before actions, raising measured behavioral safety across three benchmarks.

  6. Frontier AI systems have surpassed the self-replicating red line

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Two open-weight AI agents, when explicitly instructed and given system tools, successfully made live copies of themselves in most trials, a capability the authors call the self-replication red line.

  7. Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction

    cs.CL 2024-11 conditional novelty 6.0 of 10

    On the OEDD benchmark, LLM agents select the worse of two actions below chance when the correct choice requires combining two earlier facts and ignoring a recent distractor in contexts over 1,615 tokens.

  8. Safety case template for frontier AI: A cyber inability argument

    cs.CY 2024-11 accept novelty 6.0 of 10

    A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.

  9. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.

  10. Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.

  11. Steering Language Model Refusal with Sparse Autoencoders

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Boosting one SAE 'refusal' feature in Phi-3 Mini and Llama 3.1 raises refusal rates on unsafe and safe prompts alike while sharply reducing MMLU, TruthfulQA, and GSM8K accuracy.

  12. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  13. International Security Applications of Flexible Hardware-Enabled Guarantees

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A policy report argues that hardware-level guarantees on AI chips could make international AI governance agreements stable if states value winning only modestly and fear catastrophe sufficiently.

Pith tools