Pith. sign in

REVIEW 2 cited by

Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05307 v1 pith:YUINIT4X submitted 2024-03-08 cs.AI

classification cs.AI
keywords agentsdataanalysisinteractivetapilot-crossingactionchallengesevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Interactive Data Analysis, the collaboration between humans and LLM agents, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic interactive logs for data analysis hinder the quantitative evaluation of Large Language Model (LLM) agents in this task. To mitigate this issue, we introduce Tapilot-Crossing, a new benchmark to evaluate LLM agents on interactive data analysis. Tapilot-Crossing contains 1024 interactions, covering 4 practical scenarios: Normal, Action, Private, and Private Action. Notably, Tapilot-Crossing is constructed by an economical multi-agent environment, Decision Company, with few human efforts. We evaluate popular and advanced LLM agents in Tapilot-Crossing, which underscores the challenges of interactive data analysis. Furthermore, we propose Adaptive Interaction Reflection (AIR), a self-generated reflection strategy that guides LLM agents to learn from successful history. Experiments demonstrate that Air can evolve LLMs into effective interactive data analysis agents, achieving a relative performance improvement of up to 44.5%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

    cs.AI 2026-08 conditional novelty 7.0 of 10

    P-Bench and Fisher-R1 show that small LLM agents trained with outcome-grounded reinforcement learning on synthetic hypothesis-testing tasks can outperform frontier models at statistically valid p-value reporting and d...

  2. Aviary: training language agents on challenging scientific tasks

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A small open-source LLM trained in the new Aviary environments with expert iteration and majority voting matches or exceeds a frontier LLM agent on SeqQA and LitQA2 at far lower inference cost.

Pith tools