Pith. sign in

Title resolution pending

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 5

roles

background 1

polarities

background 1

representative citing papers

The Last Harness You'll Ever Build

cs.AI · 2026-04-22 · unverdicted · novelty 6.0

A two-level evolution framework automates the design of task-specific harnesses for AI agents by optimizing both per-task performance and a reusable meta-blueprint that enables adaptation to new domains without human engineering.

VeRO: A Harness for Agents to Optimize Agents

cs.AI · 2026-02-25 · conditional · novelty 5.0

A new harness and benchmark, VeRO, lets coding agents be scored on how much they improve target agents under a fixed evaluation budget.

citing papers explorer

Showing 5 of 5 citing papers.

  • VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents cs.RO · 2026-06-03 · unverdicted · none · ref 22

    VASO is a verification-guided self-evolution framework for LLM robot skill contracts that reaches 97.2% formal-specification compliance on Jackal and quadcopter tasks using under 100 samples.

  • The Last Harness You'll Ever Build cs.AI · 2026-04-22 · unverdicted · none · ref 12

    A two-level evolution framework automates the design of task-specific harnesses for AI agents by optimizing both per-task performance and a reusable meta-blueprint that enables adaptation to new domains without human engineering.

  • Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cs.AI · 2026-04-20 · unverdicted · none · ref 12

    LLM agents trained with a task-success reward on self-generated knowledge can spontaneously explore and adapt to new environments without any rewards or instructions at inference, yielding 20% gains on web tasks and allowing a 14B model to beat Gemini-2.5-Flash.

  • Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models cs.LG · 2026-04-09 · unverdicted · none · ref 14

    A feedforward graph of heterogeneous frozen LLMs linked by linear projections in a shared latent space outperforms single models on ARC-Challenge, OpenBookQA, and MMLU using just 17.6M trainable parameters.

  • VeRO: A Harness for Agents to Optimize Agents cs.AI · 2026-02-25 · conditional · none · ref 32

    A new harness and benchmark, VeRO, lets coding agents be scored on how much they improve target agents under a fixed evaluation budget.