REVIEW 18 cited by
AgentSims: An Open-Source Sandbox for Large Language Model Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With ChatGPT-like large language models (LLM) prevailing in the community, how to evaluate the ability of LLMs is an open question. Existing evaluation methods suffer from following shortcomings: (1) constrained evaluation abilities, (2) vulnerable benchmarks, (3) unobjective metrics. We suggest that task-based evaluation, where LLM agents complete tasks in a simulated environment, is a one-for-all solution to solve above problems. We present AgentSims, an easy-to-use infrastructure for researchers from all disciplines to test the specific capacities they are interested in. Researchers can build their evaluation tasks by adding agents and buildings on an interactive GUI or deploy and test new support mechanisms, i.e. memory, planning and tool-use systems, by a few lines of codes. Our demo is available at https://agentsims.com .
Forward citations
Cited by 18 Pith papers
-
Deep Research Bench: Evaluating AI Web Research Agents
A benchmark of 89 human-verified web research tasks with a frozen web snapshot shows frontier AI agents reach about half the score of skilled human researchers under simple prompting.
-
Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies
An adaptive stochastic offer policy cuts adversary inference of private negotiation constraints by 43–50% on synthetic traces while keeping success and utility above 90%.
-
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
An RL framework with a travel sandbox, verifier-based rewards, and failure replay produces a deployed travel-planning agent that outperforms frontier LLMs on internal benchmarks.
-
Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks
LLM-based generative agents simulate personality-driven temperature negotiations in a shared residential building, with positive personality traits associated with higher happiness and stronger friendships.
-
MobiVerse: Scaling Urban Mobility Simulation with Hybrid Lightweight Domain-Specific Generator and Large Language Models
A hybrid generator and LLM framework scales adaptive urban mobility simulation to 53,000 agents, demonstrated in a Westwood, Los Angeles case study.
-
Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.
-
Can Language Models Represent the Past without Anachronism?
Fine-tuned GPT-4o-mini still betrays its present-day training to human readers, while prompting alone fails to shift style, evidence that period pretraining may be required for historical simulation.
-
Agent Identity Evals: Measuring Agentic Identity
Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.
-
SimSpark: Interactive Simulation of Social Media Behaviors
SimSpark combines LLM-powered agents with interactive visualizations to simulate believable, customizable social media behaviors such as posting, liking, following, and replying.
-
Interpretable Locomotion Prediction in Construction Using a Memory-Driven LLM Agent With Chain-of-Thought Reasoning
An LLM agent with short-term and long-term memory improved weighted F1 for construction locomotion prediction from 0.73 to 0.90 on a self-collected dataset of 226 multimodal samples.
-
TrendSim: Simulating Trending Topics in Social Media Under Poisoning Attacks with LLM-based Multi-agent System
TrendSim simulates trending social-media topics with LLM-based user agents and prototype attackers, but its conclusions about poisoning impacts are encoded in the agent prompts.
-
Crowd: A Social Network Simulation Framework
Crowd provides a configuration-driven, GUI-supported Python framework for social network agent-based simulations, demonstrated on epidemic, influence maximization, and trust game case studies.
-
Evaluation and Benchmarking of LLM Agents: A Survey
A review that proposes a two-dimensional taxonomy for evaluating LLM agents and highlights enterprise-specific evaluation gaps.
-
Generative to Agentic AI: Survey, Conceptualization, and Challenges
Agentic AI is characterized over Generative AI by iterative reasoning, environment interaction, memory, and tool use, with autonomy as the defining difference.
-
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
A survey that frames the core problem of LLM evaluation as 'evaluation generalization': finite test sets cannot scale with unbounded model capabilities, and proposes two transitions in evaluation design.
-
MADP: Multi-Agent Deductive Planning for Enhanced Cognitive-Behavioral Mental Health Question Answer
The MADP framework uses Explorer, Empathizer, and Interpreter agents to plan LLM mental health responses, reporting roughly 4-5% improvements that may be fragile due to weak evaluation.
-
A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application
This survey organizes recent LLM-based multi-agent research into task-solving, simulation, and agent-evaluation applications, and identifies efficiency and evaluation gaps as key open problems.
-
Practical Considerations for Agentic LLM Systems
This paper is a practical survey that organizes research on LLM-based agents into design considerations for planning, memory, tools, and control flow.
Discussion (0). Continue with ORCID to comment.