Pith. sign in

REVIEW 2 cited by

Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15701 v6 pith:2EOSHQLH submitted 2024-12-20 cs.AI cs.CLcs.HC

classification cs.AIcs.CLcs.HC
keywords agentscollaborativecollaborationframeworkhumanstaskconditionenvironments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While the advancement of large language models has spurred the development of AI agents to automate tasks, numerous use cases inherently require agents to collaborate with humans due to humans' latent preferences, domain expertise, or the need for control. To facilitate the study of human-agent collaboration, we introduce Collaborative Gym (Co-Gym), an open framework for developing and evaluating collaborative agents that engage in bidirectional communication with humans while interacting with task environments. We describe how the framework enables the implementation of new task environments and coordination between humans and agents through a flexible, non-turn-taking interaction paradigm, along with an evaluation suite that assesses both collaboration outcomes and processes. Our framework provides both a simulated condition with a reliable user simulator and a real-world condition with an interactive web application. Initial benchmark experiments across three representative tasks -- creating travel plans, writing related work sections, and analyzing tabular data -- demonstrate the benefits of human-agent collaboration: The best-performing collaborative agents consistently outperform their fully autonomous counterparts in task performance, achieving win rates of 86% in Travel Planning, 74% in Tabular Analysis, and 66% in Related Work when evaluated by real users. Despite these improvements, our evaluation reveals persistent limitations in current language models and agents, with communication and situational awareness failures observed in 65% and 40% of cases in the real condition, respectively. Released under the permissive MIT license, Co-Gym supports the addition of new task environments and can be used to develop collaborative agent applications, while its evaluation suite enables assessment and improvement of collaborative agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparks of Science: Hypothesis Generation Using Structured Paper Data

    cs.CL 2025-04 conditional novelty 6.0 of 10

    The authors built HypoGen, 5,478 Bit-Flip-Spark hypothesis triples with reasoning chains from NeurIPS 2023 and ICLR 2024 papers, and fine-tuned LLaMA models on it, reporting higher feasibility but lower diversity in g...

  2. MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering

    cs.LG 2025-05 conditional novelty 5.0 of 10

    An open Gym-style environment running 200+ Kaggle competitions lets LLM agents iterate on ML solutions and provides a benchmark for training and evaluating them.

Pith tools