Pith. sign in

REVIEW 18 cited by

AgentSims: An Open-Source Sandbox for Large Language Model Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.04026 v1 pith:ERIKW5NU submitted 2023-08-08 cs.AI

classification cs.AI
keywords evaluationagentsimsagentslanguagelargeresearcherstaskstest
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With ChatGPT-like large language models (LLM) prevailing in the community, how to evaluate the ability of LLMs is an open question. Existing evaluation methods suffer from following shortcomings: (1) constrained evaluation abilities, (2) vulnerable benchmarks, (3) unobjective metrics. We suggest that task-based evaluation, where LLM agents complete tasks in a simulated environment, is a one-for-all solution to solve above problems. We present AgentSims, an easy-to-use infrastructure for researchers from all disciplines to test the specific capacities they are interested in. Researchers can build their evaluation tasks by adding agents and buildings on an interactive GUI or deploy and test new support mechanisms, i.e. memory, planning and tool-use systems, by a few lines of codes. Our demo is available at https://agentsims.com .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Research Bench: Evaluating AI Web Research Agents

    cs.AI 2025-05 conditional novelty 7.0 of 10

    A benchmark of 89 human-verified web research tasks with a frozen web snapshot shows frontier AI agents reach about half the score of skilled human researchers under simple prompting.

  2. Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies

    cs.CR 2026-07 conditional novelty 6.0 of 10

    An adaptive stochastic offer policy cuts adversary inference of private negotiation constraints by 43–50% on synthetic traces while keeping success and utility above 90%.

  3. DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

    cs.AI 2025-09 conditional novelty 6.0 of 10

    An RL framework with a travel sandbox, verifier-based rewards, and failure replay produces a deployed travel-planning agent that outperforms frontier LLMs on internal benchmarks.

  4. Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks

    cs.SI 2025-07 conditional novelty 6.0 of 10

    LLM-based generative agents simulate personality-driven temperature negotiations in a shared residential building, with positive personality traits associated with higher happiness and stronger friendships.

  5. MobiVerse: Scaling Urban Mobility Simulation with Hybrid Lightweight Domain-Specific Generator and Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A hybrid generator and LLM framework scales adaptive urban mobility simulation to 53,000 agents, demonstrated in a Westwood, Los Angeles case study.

  6. Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.

  7. Can Language Models Represent the Past without Anachronism?

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Fine-tuned GPT-4o-mini still betrays its present-day training to human readers, while prompting alone fails to shift style, evidence that period pretraining may be required for historical simulation.

  8. Agent Identity Evals: Measuring Agentic Identity

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.

  9. SimSpark: Interactive Simulation of Social Media Behaviors

    cs.HC 2025-06 conditional novelty 5.0 of 10

    SimSpark combines LLM-powered agents with interactive visualizations to simulate believable, customizable social media behaviors such as posting, liking, following, and replying.

  10. Interpretable Locomotion Prediction in Construction Using a Memory-Driven LLM Agent With Chain-of-Thought Reasoning

    cs.RO 2025-04 conditional novelty 5.0 of 10

    An LLM agent with short-term and long-term memory improved weighted F1 for construction locomotion prediction from 0.73 to 0.90 on a self-collected dataset of 226 multimodal samples.

  11. TrendSim: Simulating Trending Topics in Social Media Under Poisoning Attacks with LLM-based Multi-agent System

    cs.SI 2024-12 reject novelty 5.0 of 10

    TrendSim simulates trending social-media topics with LLM-based user agents and prototype attackers, but its conclusions about poisoning impacts are encoded in the agent prompts.

  12. Crowd: A Social Network Simulation Framework

    cs.SI 2024-12 conditional novelty 5.0 of 10

    Crowd provides a configuration-driven, GUI-supported Python framework for social network agent-based simulations, demonstrated on epidemic, influence maximization, and trust game case studies.

  13. Evaluation and Benchmarking of LLM Agents: A Survey

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A review that proposes a two-dimensional taxonomy for evaluating LLM agents and highlights enterprise-specific evaluation gaps.

  14. Generative to Agentic AI: Survey, Conceptualization, and Challenges

    cs.AI 2025-04 conditional novelty 4.0 of 10

    Agentic AI is characterized over Generative AI by iterative reasoning, environment interaction, memory, and tool use, with autonomy as the defining difference.

  15. Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks

    cs.CL 2025-04 conditional novelty 4.0 of 10

    A survey that frames the core problem of LLM evaluation as 'evaluation generalization': finite test sets cannot scale with unbounded model capabilities, and proposes two transitions in evaluation design.

  16. MADP: Multi-Agent Deductive Planning for Enhanced Cognitive-Behavioral Mental Health Question Answer

    cs.CL 2025-01 conditional novelty 4.0 of 10

    The MADP framework uses Explorer, Empathizer, and Interpreter agents to plan LLM mental health responses, reporting roughly 4-5% improvements that may be fragile due to weak evaluation.

  17. A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application

    cs.CL 2024-12 conditional novelty 4.0 of 10

    This survey organizes recent LLM-based multi-agent research into task-solving, simulation, and agent-evaluation applications, and identifies efficiency and evaluation gaps as key open problems.

  18. Practical Considerations for Agentic LLM Systems

    cs.AI 2024-12 conditional novelty 3.0 of 10

    This paper is a practical survey that organizes research on LLM-based agents into design considerations for planning, memory, tools, and control flow.

Pith tools