{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:JQK5FKKVWUSAYQJYIHHMMZAL7V","short_pith_number":"pith:JQK5FKKV","schema_version":"1.0","canonical_sha256":"4c15d2a955b5240c413841cec6640bfd5fde57b6b27972d2d01e44eb041c42e6","source":{"kind":"arxiv","id":"2506.11102","version":1},"attestation_state":"computed","paper":{"title":"Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Bo Chen, Congmin Zheng, Jiachen Zhu, Jianghao Lin, Menghui Zhu, Renting Rui, Rong Shan, Ruiming Tang, Weinan Zhang, Weiwen Liu, Yong Yu, Yunjia Xi","submitted_at":"2025-06-06T17:52:18Z","abstract_excerpt":"The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The transition from these traditional LLM chatbots to more advanced AI agents represents a pivotal evolutionary step. However, existing evaluation frameworks often blur the distinctions between LLM chatbots and AI agents, leading to confusion among researchers selecting appropriate benchmarks. To bridge this gap, this paper introduces a systematic analysis of current evaluation app"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.11102","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2025-06-06T17:52:18Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"98e3b0a7e3a7fe0f5d3e00770b7a4c31721100fce1b4df4f7f94195824854ce7","abstract_canon_sha256":"418b96fc7d06a153bdf1ccc6cdd918007299bdbcf8752b39fd51bec47c82dd51"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:20:59.476291Z","signature_b64":"QDC4UyNDR1T/+MgGWxAqDj3Ex9QxCOLt9ucU8kgZN7r62K/X5gVV/7wchi4SmgHR1WNN1cT+QPFAUD8EVuSzAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4c15d2a955b5240c413841cec6640bfd5fde57b6b27972d2d01e44eb041c42e6","last_reissued_at":"2026-07-05T11:20:59.475772Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:20:59.475772Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Bo Chen, Congmin Zheng, Jiachen Zhu, Jianghao Lin, Menghui Zhu, Renting Rui, Rong Shan, Ruiming Tang, Weinan Zhang, Weiwen Liu, Yong Yu, Yunjia Xi","submitted_at":"2025-06-06T17:52:18Z","abstract_excerpt":"The advent of large language models (LLMs), such as GPT, Gemini, and DeepSeek, has significantly advanced natural language processing, giving rise to sophisticated chatbots capable of diverse language-related tasks. The transition from these traditional LLM chatbots to more advanced AI agents represents a pivotal evolutionary step. However, existing evaluation frameworks often blur the distinctions between LLM chatbots and AI agents, leading to confusion among researchers selecting appropriate benchmarks. To bridge this gap, this paper introduces a systematic analysis of current evaluation app"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.11102","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.11102/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.11102","created_at":"2026-07-05T11:20:59.475834+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.11102v1","created_at":"2026-07-05T11:20:59.475834+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.11102","created_at":"2026-07-05T11:20:59.475834+00:00"},{"alias_kind":"pith_short_12","alias_value":"JQK5FKKVWUSA","created_at":"2026-07-05T11:20:59.475834+00:00"},{"alias_kind":"pith_short_16","alias_value":"JQK5FKKVWUSAYQJY","created_at":"2026-07-05T11:20:59.475834+00:00"},{"alias_kind":"pith_short_8","alias_value":"JQK5FKKV","created_at":"2026-07-05T11:20:59.475834+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.28158","citing_title":"OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15710","citing_title":"SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2507.21046","citing_title":"A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence","ref_index":154,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08224","citing_title":"Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering","ref_index":201,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V","json":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V.json","graph_json":"https://pith.science/api/pith-number/JQK5FKKVWUSAYQJYIHHMMZAL7V/graph.json","events_json":"https://pith.science/api/pith-number/JQK5FKKVWUSAYQJYIHHMMZAL7V/events.json","paper":"https://pith.science/paper/JQK5FKKV"},"agent_actions":{"view_html":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V","download_json":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V.json","view_paper":"https://pith.science/paper/JQK5FKKV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.11102&json=true","fetch_graph":"https://pith.science/api/pith-number/JQK5FKKVWUSAYQJYIHHMMZAL7V/graph.json","fetch_events":"https://pith.science/api/pith-number/JQK5FKKVWUSAYQJYIHHMMZAL7V/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V/action/timestamp_anchor","attest_storage":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V/action/storage_attestation","attest_author":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V/action/author_attestation","sign_citation":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V/action/citation_signature","submit_replication":"https://pith.science/pith/JQK5FKKVWUSAYQJYIHHMMZAL7V/action/replication_record"}},"created_at":"2026-07-05T11:20:59.475834+00:00","updated_at":"2026-07-05T11:20:59.475834+00:00"}