{"total":13,"items":[{"citing_arxiv_id":"2606.21037","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers","primary_cat":"cs.CR","submitted_at":"2026-06-19T02:01:50+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.11678","ref_index":5,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment","primary_cat":"cs.CL","submitted_at":"2026-06-10T05:42:36+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.5,"formal_verification":"none","one_line_summary":"On UPBench, 25 LLMs show a non-monotonic planning curve—strong Remember/Analyze, weak Understand/Evaluate—with four failure modes that support differential, not blanket, AI delegation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.30036","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Teaching Values to Machines: Simulating Human-Like Behavior in LLMs","primary_cat":"cs.AI","submitted_at":"2026-05-28T14:56:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Value-prompted LLMs align with human value structures and value-behavior relationships, and incorporating human value distributions improves population-level simulations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29874","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension","primary_cat":"cs.MA","submitted_at":"2026-05-28T12:58:56+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Empirical tests on four new frontier LLMs show cooperative equilibria favored in most balanced conditions, with provider identity correlating more strongly with outcomes than model generation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.11789","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Beyond Inefficiency: Systemic Costs of Incivility in Multi-Agent Monte Carlo Simulations","primary_cat":"cs.AI","submitted_at":"2026-05-12T08:54:03+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Monte Carlo simulations of LLM agents confirm that toxic debates take 25% longer to converge, with larger delays in smaller models, and show a first-mover advantage independent of toxicity.","context_count":1,"top_context_role":"background","top_context_polarity":"unclear","context_text":"In Trevor Cohn, Yulan He, and Yang Liu, editors,Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356-3369, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL https://aclanthology.org/ 2020.findings-emnlp.301/. Nigel Gilbert and Pietro Terna. How to build and use agent-based models in social science.Mind & Society, 1(1):57-72, March 2000. ISSN 1860-1839. doi: 10.1007/BF02512229. URLhttps://doi.org/10.1007/BF02512229. Zhe Hu, Hou Pong Chan, Jing Li, and Yu Yin. Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation, January 2025. URLhttp://arxiv.org/abs/2406.19643. arXiv:2406.19643 [cs]. Yiming Huang, Biquan Bie, Zuqiu Na, Weilin Ruan, Songxin Lei, Yutao Yue, and Xinlei He."},{"citing_arxiv_id":"2605.06524","ref_index":37,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Process Matters more than Output for Distinguishing Humans from Machines","primary_cat":"cs.AI","submitted_at":"2026-05-07T16:30:35+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A new battery of 30 cognitive tasks demonstrates that process-level behavioral features distinguish humans from frontier AI agents better than performance metrics (mean AUC 0.88), with process-specific fine-tuning improving mimicry but limited cross-task transfer.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.27271","ref_index":1,"ref_count":3,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook","primary_cat":"cs.CY","submitted_at":"2026-04-29T23:54:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"In the Moltbook AI agent community, identity-claim production is highly concentrated among a few frame entrepreneurs, with event-driven attention not translating into broad claim-making.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.09581","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding","primary_cat":"cs.AI","submitted_at":"2026-02-25T18:59:42+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Avenir-UX automates web usability testing by using GUI-grounded simulation of user behavior to generate standardized reports with SUS, SEQ, and Think Aloud protocols.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2603.03295","ref_index":1,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task","primary_cat":"cs.CL","submitted_at":"2026-02-06T15:39:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LLMs diverge from human goal selection in self-directed learning by exploiting single solutions with low variability across instances.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2509.12626","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow","primary_cat":"cs.HC","submitted_at":"2025-09-16T03:43:13+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DoubleAgents shows that a distributed-cognition design with coordination agent, dashboard, and policy module increases user comfort and reliance on AI agents for coordination tasks over time.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2504.09662","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations","primary_cat":"cs.MA","submitted_at":"2025-04-13T17:26:35+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2404.04475","ref_index":49,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators","primary_cat":"cs.LG","submitted_at":"2024-04-06T02:29:02+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Length-controlled AlpacaEval applies regression adjustment to remove length bias from LLM auto-evaluations, raising Spearman correlation with Chatbot Arena from 0.94 to 0.98.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2305.09620","ref_index":3,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction","primary_cat":"cs.CL","submitted_at":"2023-05-16T17:13:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LLM embeddings enable strong retrodiction of masked GSS opinions via cross-validation and external validation but only modest performance on entirely unasked opinions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}