{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:7ZY6QH5ZPP7KAVL7XOKNK6KDHW","short_pith_number":"pith:7ZY6QH5Z","schema_version":"1.0","canonical_sha256":"fe71e81fb97bfea0557fbb94d579433d90a4b04dbe556e43d557f8dbc817143b","source":{"kind":"arxiv","id":"2507.02825","version":5},"attestation_state":"computed","paper":{"title":"Establishing Best Practices for Building Rigorous Agentic Benchmarks","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Andy Zhang, Antony Kellermann, Cozmin Ududec, Daniel Kang, Fazl Barez, Harry Coppock, Ion Stoica, Jacob Merizian, Jacob Steinhardt, Jasjeet Sekhon, Jwala Dhamala, Kevin Meng, Mario Giulianelli, Matei Zaharia, Percy Liang, Rahul Gupta, Rebecca Weiss, Sarah Schwettmann, Sasha Cui, Sayash Kapoor, Shayne Longpre, Shu Liu, Tengjun Jin, Yada Pruksachatkun, Yuxuan Zhu","submitted_at":"2025-07-03T17:35:31Z","abstract_excerpt":"Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench Verified uses insufficient test cases, while TAU-bench counts empty responses as successful. Such issues can lead to under- or overestimation of agents"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.02825","kind":"arxiv","version":5},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2025-07-03T17:35:31Z","cross_cats_sorted":[],"title_canon_sha256":"c8b1a22ca24f10df9de7351a2b661a61d126e1b681e62dc12d89f3fd59590a3d","abstract_canon_sha256":"de49248e866a9846d5aaded435610b6ba03831b7d62b6e0eb55a2f8fb9d882c0"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:49:56.309209Z","signature_b64":"slIUAE1IA3vlTmWsRTuOpCljDfcTMGnr9P/y6uFeMQdeJo5x/sXgiq1rOOR9yzaVifR0KXFhqTfaRpXqevuZDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fe71e81fb97bfea0557fbb94d579433d90a4b04dbe556e43d557f8dbc817143b","last_reissued_at":"2026-07-05T11:49:56.308409Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:49:56.308409Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Establishing Best Practices for Building Rigorous Agentic Benchmarks","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Andy Zhang, Antony Kellermann, Cozmin Ududec, Daniel Kang, Fazl Barez, Harry Coppock, Ion Stoica, Jacob Merizian, Jacob Steinhardt, Jasjeet Sekhon, Jwala Dhamala, Kevin Meng, Mario Giulianelli, Matei Zaharia, Percy Liang, Rahul Gupta, Rebecca Weiss, Sarah Schwettmann, Sasha Cui, Sayash Kapoor, Shayne Longpre, Shu Liu, Tengjun Jin, Yada Pruksachatkun, Yuxuan Zhu","submitted_at":"2025-07-03T17:35:31Z","abstract_excerpt":"Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench Verified uses insufficient test cases, while TAU-bench counts empty responses as successful. Such issues can lead to under- or overestimation of agents"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.02825","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.02825/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.02825","created_at":"2026-07-05T11:49:56.308548+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.02825v5","created_at":"2026-07-05T11:49:56.308548+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.02825","created_at":"2026-07-05T11:49:56.308548+00:00"},{"alias_kind":"pith_short_12","alias_value":"7ZY6QH5ZPP7K","created_at":"2026-07-05T11:49:56.308548+00:00"},{"alias_kind":"pith_short_16","alias_value":"7ZY6QH5ZPP7KAVL7","created_at":"2026-07-05T11:49:56.308548+00:00"},{"alias_kind":"pith_short_8","alias_value":"7ZY6QH5Z","created_at":"2026-07-05T11:49:56.308548+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":20,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.18532","citing_title":"AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework","ref_index":167,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05391","citing_title":"Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents","ref_index":157,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26321","citing_title":"Anchor: Mitigating Artifact Drift in Agent Benchmark Generation","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26195","citing_title":"CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly","ref_index":80,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":101,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22564","citing_title":"SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22643","citing_title":"Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety","ref_index":101,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22568","citing_title":"Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13950","citing_title":"Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07161","citing_title":"SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios","ref_index":92,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12673","citing_title":"Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack","ref_index":68,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13139","citing_title":"SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08468","citing_title":"PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10448","citing_title":"Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00927","citing_title":"BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07073","citing_title":"TeamBench: Evaluating Agent Coordination under Enforced Role Separation","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2601.11868","citing_title":"Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07161","citing_title":"SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios","ref_index":94,"is_internal_anchor":false},{"citing_arxiv_id":"2604.05229","citing_title":"From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19818","citing_title":"Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI","ref_index":3,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW","json":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW.json","graph_json":"https://pith.science/api/pith-number/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/graph.json","events_json":"https://pith.science/api/pith-number/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/events.json","paper":"https://pith.science/paper/7ZY6QH5Z"},"agent_actions":{"view_html":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW","download_json":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW.json","view_paper":"https://pith.science/paper/7ZY6QH5Z","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.02825&json=true","fetch_graph":"https://pith.science/api/pith-number/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/graph.json","fetch_events":"https://pith.science/api/pith-number/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/action/storage_attestation","attest_author":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/action/author_attestation","sign_citation":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/action/citation_signature","submit_replication":"https://pith.science/pith/7ZY6QH5ZPP7KAVL7XOKNK6KDHW/action/replication_record"}},"created_at":"2026-07-05T11:49:56.308548+00:00","updated_at":"2026-07-05T11:49:56.308548+00:00"}