Pith. sign in

hub Canonical reference

Travelplanner: A benchmark for real-world planning with language agents

Canonical reference. 83% of citing Pith papers cite this work as background.

25 Pith papers citing it
Background 83% of classified citations

hub tools

citation-role summary

background 5 dataset 1

citation-polarity summary

representative citing papers

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns

cs.SE · 2026-06-17 · unverdicted · novelty 7.0

StaminaBench evaluates coding agents over 100 procedurally generated change requests to a REST API, finding that tested models fail within 5-6 turns without feedback but improve up to 12x with test feedback and good harnesses.

State-Centric Decision Process

cs.AI · 2026-05-12 · unverdicted · novelty 7.0

SDP constructs a task-induced state space from raw text by having agents commit to and certify natural-language predicates as states, enabling structured planning and analysis in unstructured language environments.

COMPASS: Benchmarking Constrained Optimization in LLM Agents

cs.LG · 2025-10-08 · unverdicted · novelty 7.0

COMPASS benchmark shows LLM agents reach 70-90% feasibility but only 20-60% optimality on constrained travel planning tasks, attributing the gap to insufficient search space exploration rather than tool use.

Trip+: Benchmarking Agents in Personalized Interactive Travel Planning

cs.AI · 2026-06-19 · unverdicted · novelty 6.0

Trip+ benchmark evaluates language model agents on generating and revising personalized minute-level travel itineraries under dynamic interactions, finding consistent gaps where models produce feasible but exhausting plans that ignore traveler profiles.

Scaling Diffusion Language Models via Adaptation from Autoregressive Models

cs.CL · 2024-10-23 · conditional · novelty 6.0

Adapting autoregressive models via continual pre-training yields diffusion language models from 127M to 7B parameters that outperform prior diffusion models and compete with their autoregressive counterparts on language, reasoning, and commonsense benchmarks.

Interactive Evaluation Requires a Design Science

cs.AI · 2026-05-18 · unverdicted · novelty 5.0

Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.

Agentic AI for Trip Planning Optimization Application

cs.AI · 2026-04-30 · unverdicted · novelty 5.0

An orchestrated multi-agent AI framework for trip planning optimization paired with a new ground-truth dataset achieves 77.4% accuracy on the TOP Benchmark, outperforming single-agent and workflow baselines.

citing papers explorer

Showing 25 of 25 citing papers.