{"total":14,"items":[{"citing_arxiv_id":"2607.06815","ref_index":2,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies","primary_cat":"cs.CR","submitted_at":"2026-07-07T21:22:28+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"An adaptive stochastic offer policy cuts adversary inference of private negotiation constraints by 43–50% on synthetic traces while keeping success and utility above 90%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.26883","ref_index":8,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"EconSimulacra: A Digital Twin Platform of Socio-Economic Systems Powered by LLM Agents","primary_cat":"cs.DL","submitted_at":"2026-06-25T11:13:13+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"EconSimulacra is a multi-agent LLM simulator that couples economy, mobility, and social networks through shared internal states to reproduce nonlinear relationships between online attention and offline popularity.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.08367","ref_index":78,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy","primary_cat":"cs.MA","submitted_at":"2026-06-06T22:59:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Emergence World is a model-agnostic multi-agent simulation platform integrating live data, 120+ tools, persistent memory, and democratic governance, illustrated by a 15-day study showing divergent outcomes across five LLM models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02293","ref_index":38,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AI as a Tool for Simulation-Based Experiments in Literary Studies","primary_cat":"cs.CL","submitted_at":"2026-06-01T14:16:13+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Proposes AI-driven simulations for literary-historical experiments and reports preliminary text-generation results claiming the first limited in-distribution outputs matching human novels.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25815","ref_index":30,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network","primary_cat":"cs.AI","submitted_at":"2026-05-25T13:12:27+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.07462","ref_index":68,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment","primary_cat":"cs.CL","submitted_at":"2026-05-08T09:10:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"An AI-agent social platform generated mostly neutral content whose use in fine-tuning reduced model truthfulness comparably to human Reddit data, suggesting limited unique harm but flagging tail risks like secret leaks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2506.23978","ref_index":55,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LLM Agents Are the Antidote to Walled Gardens","primary_cat":"cs.LG","submitted_at":"2025-06-30T15:45:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"LLM agents enable universal interoperability by serving as automatic translators and adapters between proprietary digital services.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2505.11336","ref_index":79,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration","primary_cat":"cs.CL","submitted_at":"2025-05-16T15:02:19+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"XtraGPT is a suite of 1.5B-14B parameter open-source LLMs fine-tuned on 140,000 revision pairs from 7,000 top-tier papers to support controllable, context-aware academic paper editing.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2410.07283","ref_index":41,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems","primary_cat":"cs.MA","submitted_at":"2024-10-09T11:01:29+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"Prompt injection attacks can self-replicate across LLM agents in multi-agent systems, enabling data theft, misinformation, and system disruption while propagating silently.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2406.13352","ref_index":31,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents","primary_cat":"cs.CR","submitted_at":"2024-06-19T08:55:56+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"AgentDojo introduces an extensible evaluation framework populated with realistic agent tasks and security test cases to measure prompt injection robustness in tool-using LLM agents.","context_count":1,"top_context_role":"baseline","top_context_polarity":"baseline","context_text":"https://lakeraai.github.io/chainguard/. 2024. [29] LangChain. Hugging Face prompt injection identification . https://python.langchain. com/v0.1/docs/guides/productionization/safety/hugging _face_prompt_injection/. 2024. [30] Learn Prompting. Sandwich Defense . https://learnprompting.org/docs/prompt _ hacking/defensive_measures/sandwich_defense. 2024. [31] Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen. AgentSims: An Open-Source Sandbox for Large Language Model Evaluation. 2023. arXiv:2308.04026 [cs.AI]. [32] Xiao Liu et al. AgentBench: Evaluating LLMs as Agents. 2023. arXiv:2308.03688 [cs.AI]. [33] Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang,"},{"citing_arxiv_id":"2401.05561","ref_index":137,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"TrustLLM: Trustworthiness in Large Language Models","primary_cat":"cs.CL","submitted_at":"2024-01-10T22:07:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"TrustLLM defines eight trustworthiness principles, creates a six-dimension benchmark, and evaluates 16 LLMs showing proprietary models generally lead but some open-source ones are close while over-calibration can hurt utility.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"in mathematics [122, 123], general science [29, 124], and engineering [125, 126] domains. In the medical field, LLMs have been evaluated for their proficiency in addressing medical queries [ 127, 128], medical examinations [129, 130], and functioning as medical assistants [131, 132]. In addition, some benchmarks are designed to evaluate specific language abilities of LLMs like Chinese [133, 134, 135, 136]. Besides, agent applications [137] underline their capabilities for interaction and using tools [ 138, 139, 140, 141]. Beyond these areas, LLMs contribute to different domains, such as education [ 142], finance [143, 144, 145, 146], search and recommendation [147, 148], personality testing [149]. Other specific applications, such as game design [150] and log parsing [151], illustrate the broad scope of the application and evaluation of LLMs."},{"citing_arxiv_id":"2311.12983","ref_index":54,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GAIA: a benchmark for General AI Assistants","primary_cat":"cs.CL","submitted_at":"2023-11-21T20:34:47+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"GAIA benchmark shows humans at 92% accuracy on simple real-world questions far outperform current AI systems at 15%, proposing this gap as a key milestone for general AI.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2309.07864","ref_index":175,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Rise and Potential of Large Language Model Based Agents: A Survey","primary_cat":"cs.AI","submitted_at":"2023-09-14T17:12:03+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":4.0,"formal_verification":"none","one_line_summary":"The paper surveys the origins, frameworks, applications, and open challenges of AI agents built on large language models.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Raising the length limit of Transformers BART [163], Park et al. [164], LongT5 [165], CoLT5 [166], Ruoss et al. [167], etc. Summarizing memory Generative Agents [22], SCM [168], Reflexion [169], Memory- bank [170], ChatEval [171], etc. Compressing mem- ories with vectors or data structures ChatDev [109], GITM [172], RET-LLM [173], AgentSims [174], ChatDB [175], etc. Memory retrieval Automated retrieval Generative Agents [22], Memory- bank [170], AgentSims [174], etc. Interactive retrieval Memory Sandbox[176], ChatDB [175], etc. Reasoning & Planning §3.1.4 Reasoning CoT [95], Zero-shot-CoT [96], Self-Consistency [97], Self- Polish [99], Selection-Inference [177], Self-Refine [178], etc. Planing Plan formulation"},{"citing_arxiv_id":"2308.11432","ref_index":34,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"A Survey on Large Language Model based Autonomous Agents","primary_cat":"cs.AI","submitted_at":"2023-08-22T13:30:37+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A survey of LLM-based autonomous agents that proposes a unified framework for their construction and reviews applications in social science, natural science, and engineering along with evaluation methods and future directions.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"solidates important information over time. For in- stance, Generative Agent [20] employs a hybrid memory structure to facilitate agent behaviors. The short-term memory contains the context informa- tion about the agent current situations, while the long-term memory stores the agent past behaviors and thoughts, which can be retrieved according to the current events. AgentSims [34] also implements a hybrid memory architecture. The information pro- vided in the prompt can be considered as short-term memory. In order to enhance the storage capac- ity of memory, the authors propose a long-term memory system that utilizes a vector database, fa- cilitating e fficient storage and retrieval. Specifi- cally, the agent's daily memories are encoded as"}],"limit":50,"offset":0}