Pith. sign in

hub Canonical reference

AdaEvolve: Adaptive LLM driven zeroth-order optimization

Canonical reference. 80% of citing Pith papers cite this work as background.

26 Pith papers citing it
Background 80% of classified citations

hub tools

citation-role summary

background 4 dataset 1

citation-polarity summary

years

2026 26

representative citing papers

What Do Evolutionary Coding Agents Evolve?

cs.NE · 2026-05-19 · unverdicted · novelty 7.0

Evolutionary coding agents achieve most benchmark gains through a small subset of edit types and by cycling previously deleted code lines rather than developing new algorithmic structures.

Meta-Harness: End-to-End Optimization of Model Harnesses

cs.AI · 2026-03-30 · unverdicted · novelty 7.0

Meta-Harness discovers improved harness code for LLMs via agentic search over prior execution traces, yielding 7.7-point gains on text classification with 4x fewer tokens and 4.7-point gains on math reasoning across held-out models.

When Does Continual Learning Require Learning

cs.LG · 2026-07-08 · conditional · novelty 6.0

Different patterns of environmental change (space vs time) require different LLM update behaviors; no single family of methods—prompts, distillation, RL, or compression—handles all regimes.

Building Agent Harnesses for Scientific Curation from Multimodal Sources

cs.AI · 2026-06-19 · conditional · novelty 6.0

An agent harness combining staged task decomposition, multimodal evidence tooling, and artifact-grounded self-improvement scores 81.0 GRAS on multimodal scientific curation, 22.4 points above the strongest baseline — with the caveat that 8 of 23 evaluation papers were used for optimization.

VESTA: Visual Exploration with Statistical Tool Agents

cs.AI · 2026-05-29 · unverdicted · novelty 6.0

VESTA introduces dynamic tool creation for VLMs that outperforms static-tool and no-tool baselines on distribution fitting, time series, and astronomy tasks in the new DAWN benchmark.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

cs.AI · 2026-05-25 · unverdicted · novelty 6.0

ScientistOne introduces Chain-of-Evidence and an audit system that achieves zero hallucinated references, perfect score verification, and top method-code alignment while matching or beating human experts on five frontier tasks and generalizing to six more.

TuxBot: Semantic-Aware Online OS Tuning with Large Language Models

cs.OS · 2026-05-14 · reject · novelty 6.0

An LLM-driven dual-loop controller claims 72.5% stable-phase improvement over default and 153.3% over the strongest non-LLM baseline, but the comparison protocol inflates the gaps by scoring baselines during continued exploration.

Evolutionary Ensemble of Agents

cs.NE · 2026-05-09 · conditional · novelty 6.0

A dual-population evolutionary ensemble of coding agents discovers a rescale-then-interpolate PE for ICON example-count generalization and outperforms static-agent baselines via stage-dependent adaptation.

Evaluation-driven Scaling for Scientific Discovery

cs.LG · 2026-04-21 · unverdicted · novelty 6.0

SimpleTES scales test-time evaluation in LLMs to discover state-of-the-art solutions on 21 scientific problems across six domains, outperforming frontier models and optimization pipelines with examples like 2x faster LASSO and new Erdos constructions.

citing papers explorer

Showing 26 of 26 citing papers.