Pith. sign in

REVIEW 10 cited by

AI capabilities can be significantly improved without expensive retraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07413 v1 pith:4EJKX63F submitted 2023-12-12 cs.AI cs.LG

classification cs.AIcs.LG
keywords enhancementspost-trainingtrainingdifferentimproveperformancecomputeexpensive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

State-of-the-art AI systems can be significantly improved without expensive retraining via "post-training enhancements"-techniques applied after initial training like fine-tuning the system to use a web browser. We review recent post-training enhancements, categorizing them into five types: tool-use, prompting methods, scaffolding, solution selection, and data generation. Different enhancements improve performance on different tasks, making it hard to compare their significance. So we translate improvements from different enhancements into a common currency, the compute-equivalent gain: how much additional training compute would be needed to improve performance by the same amount as the enhancement. Our non-experimental work shows that post-training enhancements have significant benefits: most surveyed enhancements improve benchmark performance by more than a 5x increase in training compute, some by more than 20x. Post-training enhancements are relatively cheap to develop: fine-tuning costs are typically <1% of the original training cost. Governing the development of capable post-training enhancements may be challenging because frontier models could be enhanced by a wide range of actors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Compute Requirements for Algorithmic Innovation in Frontier AI Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Estimated development compute for 36 LLM pretraining innovations shows half would remain possible under GPT-2-level or 8-H100 compute caps.

  2. AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A MIRI governance agenda argues for an internationally coordinated halt to dangerous AI development and catalogs around 400 research questions across four strategic scenarios.

  3. Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress

    cs.LG 2025-05 conditional novelty 6.0 of 10

    The authors introduce a compute-dependent versus compute-independent framework and report nanoGPT experiments showing compute-independent algorithms such as LayerNorm and RoPE give compute-equivalent gains up to 1.9x,...

  4. Towards Responsible Governing AI Proliferation

    cs.CY 2024-12 conditional novelty 6.0 of 10

    The paper proposes a 'Proliferation' paradigm of AI, where small, hidden, augmented, decentralized, and open-weight models challenge compute-centric governance.

  5. Safety case template for frontier AI: A cyber inability argument

    cs.CY 2024-11 accept novelty 6.0 of 10

    A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.

  6. Multi-Head Attention Residuals

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Splitting the depth-routing query into per-subspace heads (a parameter-free reshape) improves Transformer validation loss at 100M–1B and mid-training at 8B.

  7. Bare Minimum Mitigations for Autonomous AI Development

    cs.CY 2025-04 conditional novelty 5.0 of 10

    A position paper proposing two thresholds and four minimum safeguards for frontier AI labs before AI agents automate AI R&D.

  8. Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation

    cs.AI 2024-11 conditional novelty 5.0 of 10

    AI evaluation-based regulation should require developers to state and justify key assumptions, and halt development when those justifications are inadequate.

  9. What AI evaluations for preventing catastrophic risks can and cannot do

    cs.CY 2024-11 conditional novelty 4.0 of 10

    AI evaluations can establish lower bounds on capabilities but cannot establish upper bounds, forecast future capabilities robustly, or assess misalignment risk, so they should not be the primary basis for AI safety decisions.

  10. A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A narrative review of behavioral and representational Theory of Mind in LLMs, with a taxonomy of safety risks and mitigation directions.

Pith tools