{"id":"ae7a3b9d-0131-4335-924f-888176078bf9","arxiv_id":"2502.01635","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.","lead":"This paper introduces the AI Agent Index, a public database documenting 67 deployed agentic AI systems, each described across 33 structured fields covering components, applications, and safety practices. It finds that developers share ample information about capabilities but little about safety testing and risk management, a gap that matters for AI governance and transparency.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the qualitative finding of sparse public safety information is robust to the acknowledged subjectivity in inclusion criteria and coding.","rationale":"The paper is best read as introducing a structured, transparently documented registry rather than a statistically representative population estimate. Its headline claim is conditional on the authors' inclusion criteria, which are disclosed in Figure 3 and whose limitations are addressed in Section 6. The reader's weakest assumption correctly identifies the subjective agency threshold, but the threshold does not function as a load-bearing premise for the qualitative finding: the safety percentages are low across the board, and plausible boundary changes (adding customer-service agents, removing discretionary frontier-model entries, or restricting to industry systems) all preserve the direction of the result. The coding of safety indicators is less explicitly documented, but the agent cards cite sources and the raw data is public, so the percentages can be audited. No internal inconsistency in the quantitative claims was found. The verdict should remain ACCEPT, with the suggested sensitivity analysis as a useful follow-up verification.","tokens_in":17220,"tokens_out":10934,"duration_ms":115143,"concrete_test":"Recompute the three safety-transparency percentages (formal safety policy, external safety testing, and public safety evaluations) on the industry-only subsample and again after removing all entries admitted via the discretionary final node of Figure 3; if any rate exceeds 50% or the qualitative contrast with capability documentation reverses, the headline claim would require qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I could not identify a load-bearing technical error. The central claim is a contrast between relatively ample capability and usage documentation and limited public safety and risk documentation among the 67 indexed systems. The two obvious soft spots are the subjective 'meaningfully higher degree of agency than ChatGPT-4o' threshold in Figure 3 and the uncalibrated coding of indicators such as 'formal safety policy' used in Figure 2. Both are real limitations, and the paper explicitly acknowledges the inclusion subjectivity in Section 6. However, neither threatens the qualitative finding: even generous reclassifications would leave the safety percentages far below the capability documentation rates, and the index excludes internal systems, so any selection bias is toward more transparency, making the safety-transparency deficit conservative. The final discretionary inclusion node (e.g., OpenAI o3) is inconsistent with the stated exclusion of language models, but it affects only a handful of entries and would not change the aggregate conclusion. The released raw data makes independent verification possible. I therefore find no reason to alter the reader's ACCEPT verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces the AI Agent Index, a structured public database of 67 deployed agentic AI systems as of December 31, 2024. The authors develop inclusion criteria based on the four agency characteristics from Chan et al. (2023) (underspecification, directness of impact, goal-directedness, and long-term planning) plus an explicit decision graph, and they populate 33 fields per system from public sources and developer correspondence, with a reported 36% developer response rate. The main empirical finding is an asymmetry in public documentation: 70.1% of indexed systems have public documentation and 49.3% release code, while only 19.4% disclose a formal safety policy, 7.5% report external safety testing, and 9% report public safety evaluations by the developer. The paper also reports distributions across countries, developer types, and application domains, and it closes with governance recommendations. Limitations, including English-language bias, public-documentation bias, incomplete developer verification, and the subjectivity of the inclusion threshold, are acknowledged in Section 6.","tokens_in":17343,"tokens_out":6733,"duration_ms":64259,"significance":"If the index is accepted as representative, this is the first structured, system-level empirical evidence of a documentation gap between capability and usage information on the one hand and safety and risk-management information on the other hand for agentic AI. The contribution is timely and useful for users, auditors, researchers, and policymakers. The paper's strengths include the release of raw data and archived citations, a detailed sample agent card, an explicit decision graph for inclusion, and the decision not to use the index as a scorecard in order to reduce gaming incentives. The acknowledged subjectivity of the inclusion criteria and the reliance on public documentation limit precision, but the qualitative finding is robust: even a generous reclassification of marginal systems would leave safety disclosure far below capability disclosure, and the selection bias toward publicly documented systems makes the safety deficit conservative. The paper is therefore credible as a first empirical mapping of the field.","major_comments":[],"minor_comments":[{"comment":"The final discretionary node of the inclusion decision graph is not constrained by the system being deployed or open source, and the examples given (OpenAI o3, Project Mariner) are either not deployed as of the cutoff or are models rather than agentic systems; this sits awkwardly with the stated exclusions and with the 'currently deployed' framing, so the operational rule and the number of affected entries should be stated explicitly.","section":"Section 3 / Figure 3"},{"comment":"The threshold 'meaningfully higher degree of agency than ChatGPT-4o' is not operationalized beyond the footnote about ChatGPT-4o; a brief calibration example or sensitivity check would make the sample boundary more reproducible, although the Section 6 caveat already mitigates the risk to the main conclusion.","section":"Section 3 / Figure 3"},{"comment":"The paper should state explicitly in the main results that the safety percentages measure public availability of documented practices, not the absence of internal practices; Section 6 makes this point, but it is central enough to the interpretation of the headline numbers to appear alongside the findings.","section":"Section 5 / Figure 2"},{"comment":"The sample card for Magentic One lists the announcement date as November 4, 2023, which appears to be a typo for November 4, 2024, given the system's technical report and the paper's timeline.","section":"Appendix A"},{"comment":"The caption should clarify whether the bars count systems or unique organizations, since it states that some developers contribute multiple systems but the reader cannot tell from the figure alone whether the country totals are system-level counts.","section":"Figure 5"},{"comment":"Minor editorial issues: 'V onage' should be 'Vonage', 'Moatlesss' should likely be 'Moatless', and the CORE-Bench reference misspells 'Nadgir' as 'Nagdir'.","section":"Section 3 / References"}],"recommendation":"accept","confidential_remarks":"This is a strong empirical contribution for the journal's scope, and the minor comments can be addressed in a camera-ready revision without further external review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe AI Agent Index is the first systematic public registry of deployed agentic AI systems, and the core finding holds up: capabilities are documented, safety practices are not. The authors collected 67 systems, coded 33 fields each, and released the raw data. Their headline percentages (19.4% disclosing a formal safety policy, 7.5% external safety testing, 9% public safety evaluations) are direct counts from that database, not fitted claims.\n\nThe paper is unusually honest about its limits. It reports the 36% developer response rate, the English-language and public-documentation bias, and the subjective 'meaningfully higher degree of agency than ChatGPT-4o' inclusion threshold. The inclusion decision graph includes a discretionary final node that they only used for a few high-profile announced-but-not-yet-deployed systems like OpenAI o3, which does sit uneasily with their stated exclusion of non-deployed systems, but the aggregate conclusion is not sensitive to those few entries.\n\nThe real soft spot is the uncalibrated coding of indicators like 'formal safety policy.' The manual coding is plausible but not independently replicated, and the paper does not provide inter-coder reliability or a rubric that distinguishes a perfunctory policy statement from a substantive safety management system. That limits the precision of the percentages, though not the direction of the finding.\n\nI agree with the stress-test note: the qualitative conclusion that safety disclosure lags capability disclosure would survive even generous reclassification of the ambiguous cases. If anything, the public-documentation bias cuts toward overestimating transparency, so the safety-transparency deficit is a conservative estimate.\n\nThe citation pattern is appropriate. It builds on the Foundation Model Transparency Index, AI Incident Database, and the AI Risk Repository, and the novelty claim—that no prior framework catalogs deployed agentic systems as a class—is justified.\n\nThis paper is for policymakers, auditors, and researchers who need a baseline map of the agentic ecosystem. It does not resolve a scientific question, but it is a well-scoped empirical resource. It deserves a serious referee: the methodology is transparent, the data is reusable, and the limitations are stated rather than buried.\n\nRecommendation: send it to peer review. I would take a look at the agent card coding rubric in the supplementary materials during review, but I would not block acceptance on it.\n\nRegards.","headline":"A genuinely first-of-its-kind public registry of deployed agentic systems, with an honest limitations section and a robust headline finding about sparse safety disclosure.","tokens_in":17952,"tokens_out":2967,"would_cite":true,"duration_ms":24647,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The first public index of 67 agentic AI systems shows safety disclosure lags far behind capability information.","keywords":["agentic AI","AI Agent Index","transparency","safety disclosure","risk management","agent cards","AI governance","public database"],"falsifier":"Re-run the inclusion graph with the agency threshold anchored to a clearly weaker baseline, such as GPT-3.5, or drop the discretionary final node, and recompute the share of indexed systems disclosing a formal safety policy; if that share moves by more than a few percentage points, the reported 19.4% is an artifact of the chosen threshold rather than a property of the ecosystem. Alternatively, check the public documentation of the 43 developers who never replied to the index team; if a substantial fraction of them publish formal safety policies, the 'limited information' conclusion partly reflects non-response.","tokens_in":17005,"feed_emoji":"📊","tokens_out":6046,"duration_ms":48060,"temperature":0.7,"pith_summary":"The paper introduces the AI Agent Index, a public database of 67 currently deployed agentic AI systems that can plan and act with limited human involvement. For each system it documents 33 fields across six categories, from backend models and tool use to guardrails, safety evaluations, and developer information, drawing on public sources and developer feedback. The central finding is an asymmetry: developers publish substantial information about what their agents can do, but almost none about how they are managed for safety. The authors argue this is the first structured, cross-system evidence of that gap, and that it gives policymakers and auditors a concrete starting point for transparency and governance efforts.","feed_headline":"Safety disclosures lag in 67-system agentic-AI index","feed_subtitle":"Only 19% of indexed agents publish a formal safety policy; under 10% report external testing or public safety evaluations.","key_machinery":"The machinery is the inclusion decision graph plus the agent card template. The decision graph starts from a named, 'agentic' system and requires it to accomplish a diverse range of tasks with a meaningfully higher degree of agency than ChatGPT-4o, judged using the four characteristics of agency the paper adopts from its background review: underspecification, directness of impact, goal-directedness, and long-term planning. It excludes plain language models, development frameworks without a qualifying flagship system, and systems that cannot be used off the shelf; the final node lets the authors include important announced-but-not-yet-deployed systems at their discretion. Applied to each included system, the 33-field agent card standardizes what is recorded, and its use of 'None' or 'Unknown' is the mechanism that surfaces the safety-transparency gap.","core_discovery":"The paper's central discovery is the transparency asymmetry documented by the index: while 70.1% of the 67 indexed systems publicly release documentation and 49.3% release code, only 19.4% disclose a formal safety policy, 7.5% report external safety testing, and 9% report public safety evaluations by the developer. The 33-field agent cards record 'None' or 'Unknown' when information is absent, which is what turns the asymmetry into a measurable, citable finding. The paper frames this as the first public database of deployed agentic systems and the first structured evidence that the agentic AI ecosystem is transparent about capabilities and applications but opaque about safety and risk management.","pith_inferences":["The index likely understates true safety-practice disclosure: 64% of developers did not respond, and internal or unpublished safety processes are invisible by construction, so the headline percentages may be a lower bound on practice but an upper bound on public transparency.","The ChatGPT-4o anchor for agency will date quickly as frontier models become more agentic, making the December 31, 2024 snapshot hard to compare with later indices unless the threshold is re-anchored.","The 9% and 7.5% figures are small enough that a re-sampling with a slightly different inclusion rule could shift them materially; the paper's qualitative conclusion of a safety gap is more robust than its exact percentages.","A natural extension the paper does not build is a per-system transparency score separating capability disclosure from safety disclosure, letting users and regulators track whether the gap widens or closes over time."],"forward_implications":["The index gives policymakers a first evidence base: the deployment rate, geographic and institutional spread, and domain concentration of agentic systems are now documented rather than anecdotal.","Governance attention should focus on corporate developers, US-based organizations, and software-engineering and computer-use agents, which together dominate the index.","The transparency gap argues for disclosure mechanisms as an early intervention, including structured bug bounties, coordinated external testing of agents, and integration of indices into model registries.","Future documentation efforts should adopt the paper's method of explicit 'None' and 'Unknown' recording and should scope their selection criteria to reduce the subjectivity the authors acknowledge."],"supporting_citations":[{"why":"Supplies the four characteristics of agency that anchor the index's inclusion criteria.","marker":"(Chan et al., 2023)"},{"why":"The SWE-bench leaderboard used to identify software-engineering agentic systems for the index.","marker":"(Jimenez et al., 2023)"},{"why":"The GAIA leaderboard used to identify general-purpose agentic assistants for the index.","marker":"(Mialon et al., 2023b)"},{"why":"The Foundation Model Transparency Index, the closest prior transparency effort this index extends to agentic systems.","marker":"(Bommasani et al., 2023a)"},{"why":"The AI Risk Repository, a comparable public database whose documentation approach the index builds alongside.","marker":"(Slattery et al., 2024)"},{"why":"Datasheets, one of the documentation frameworks motivating the structured 33-field agent card format.","marker":"(Gebru et al., 2018)"},{"why":"Model cards, another documentation framework whose reporting-in-the-absence-of-information convention the agent cards adopt.","marker":"(Mitchell et al., 2019)"}],"fun_headline_variants":["Agentic AI index finds safety info scarce","Only 7% of agentic systems report external safety tests","AI agent index: capabilities transparent, safety opaque","Safety disclosures lag in 67-system agentic AI index","First agentic AI index exposes safety info gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire sample rests on the authors' judgment that a system has 'a meaningfully higher degree of agency than ChatGPT-4o,' plus a discretionary final inclusion step, so the 67-system list and every percentage derived from it depend on that subjective threshold.","fun_headline_variants_meta":{"raw":{"variants":["Agentic AI index finds safety info scarce","Only 7% of agentic systems report external safety tests","AI agent index: capabilities transparent, safety opaque","Safety disclosures lag in 67-system agentic AI index","First agentic AI index exposes safety info gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000425,"raw_usage":{"total_tokens":2126,"prompt_tokens":841,"completion_tokens":1285,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":1210}},"tokens_in":457,"tokens_out":1285,"duration_ms":9180,"temperature":1.0,"reasoning_tokens":1210,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T14:43:25.779094+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the inclusion graph with the agency threshold anchored to a clearly weaker baseline, such as GPT-3.5, or drop the discretionary final node, and recompute the share of indexed systems disclosing a formal safety policy; if that share moves by more than a few percentage points, the reported 19.4% is an artifact of the chosen threshold rather than a property of the ecosystem. Alternatively, check the public documentation of the 43 developers who never replied to the index team; if a substantial fraction of them publish formal safety policies, the 'limited information' conclusion partly reflects non-response.","supporting_citations":[],"review_version":1}