{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:BNQSVT3XLHJMNRJGH6HSF73VSM","short_pith_number":"pith:BNQSVT3X","schema_version":"1.0","canonical_sha256":"0b612acf7759d2c6c5263f8f22ff759337bdcfb745141dbff634dd808caefa4e","source":{"kind":"arxiv","id":"2310.07637","version":5},"attestation_state":"computed","paper":{"title":"OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.NI"],"primary_cat":"cs.AI","authors_text":"Bohan Chen, Changhua Pei, Dan Pei, Gaogang Xie, Haiming Zhang, Jianhui Li, Kun Wang, Longlong Xu, Minghua Ma, Mingze Sun, Shenglin Zhang, Xiaohui Nie, Xidao Wen, Yongqian Sun, Yuhe Liu, Zhirui Zhang","submitted_at":"2023-10-11T16:33:29Z","abstract_excerpt":"Information Technology (IT) Operations (Ops), particularly Artificial Intelligence for IT Operations (AIOps), is the guarantee for maintaining the orderly and stable operation of existing information systems. According to Gartner's prediction, the use of AI technology for automated IT operations has become a new trend. Large language models (LLMs) that have exhibited remarkable capabilities in NLP-related tasks, are showing great potential in the field of AIOps, such as in aspects of root cause analysis of failures, generation of operations and maintenance scripts, and summarizing of alert inf"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.07637","kind":"arxiv","version":5},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2023-10-11T16:33:29Z","cross_cats_sorted":["cs.NI"],"title_canon_sha256":"df284e98b36e1071928033a89a010a670d3ed765bd49e5c3d23d68a452ed528d","abstract_canon_sha256":"03402873e46989fbf41525cc887817afe7960135593c3f50cd86af6a69e5f423"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:22:28.911564Z","signature_b64":"4nInzYoSy3u/OSGWVva6339VdfXWV89nHIFKtvxS7Iu/QyfcSYw8/ompGkta+Xr/qJz+qoVbY+oMWuAJfL56CQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"0b612acf7759d2c6c5263f8f22ff759337bdcfb745141dbff634dd808caefa4e","last_reissued_at":"2026-07-05T11:22:28.911015Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:22:28.911015Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.NI"],"primary_cat":"cs.AI","authors_text":"Bohan Chen, Changhua Pei, Dan Pei, Gaogang Xie, Haiming Zhang, Jianhui Li, Kun Wang, Longlong Xu, Minghua Ma, Mingze Sun, Shenglin Zhang, Xiaohui Nie, Xidao Wen, Yongqian Sun, Yuhe Liu, Zhirui Zhang","submitted_at":"2023-10-11T16:33:29Z","abstract_excerpt":"Information Technology (IT) Operations (Ops), particularly Artificial Intelligence for IT Operations (AIOps), is the guarantee for maintaining the orderly and stable operation of existing information systems. According to Gartner's prediction, the use of AI technology for automated IT operations has become a new trend. Large language models (LLMs) that have exhibited remarkable capabilities in NLP-related tasks, are showing great potential in the field of AIOps, such as in aspects of root cause analysis of failures, generation of operations and maintenance scripts, and summarizing of alert inf"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.07637","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.07637/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.07637","created_at":"2026-07-05T11:22:28.911090+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.07637v5","created_at":"2026-07-05T11:22:28.911090+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.07637","created_at":"2026-07-05T11:22:28.911090+00:00"},{"alias_kind":"pith_short_12","alias_value":"BNQSVT3XLHJM","created_at":"2026-07-05T11:22:28.911090+00:00"},{"alias_kind":"pith_short_16","alias_value":"BNQSVT3XLHJMNRJG","created_at":"2026-07-05T11:22:28.911090+00:00"},{"alias_kind":"pith_short_8","alias_value":"BNQSVT3X","created_at":"2026-07-05T11:22:28.911090+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.29193","citing_title":"A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07161","citing_title":"SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02906","citing_title":"OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07161","citing_title":"SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios","ref_index":57,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02906","citing_title":"OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning","ref_index":25,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM","json":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM.json","graph_json":"https://pith.science/api/pith-number/BNQSVT3XLHJMNRJGH6HSF73VSM/graph.json","events_json":"https://pith.science/api/pith-number/BNQSVT3XLHJMNRJGH6HSF73VSM/events.json","paper":"https://pith.science/paper/BNQSVT3X"},"agent_actions":{"view_html":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM","download_json":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM.json","view_paper":"https://pith.science/paper/BNQSVT3X","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.07637&json=true","fetch_graph":"https://pith.science/api/pith-number/BNQSVT3XLHJMNRJGH6HSF73VSM/graph.json","fetch_events":"https://pith.science/api/pith-number/BNQSVT3XLHJMNRJGH6HSF73VSM/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM/action/timestamp_anchor","attest_storage":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM/action/storage_attestation","attest_author":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM/action/author_attestation","sign_citation":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM/action/citation_signature","submit_replication":"https://pith.science/pith/BNQSVT3XLHJMNRJGH6HSF73VSM/action/replication_record"}},"created_at":"2026-07-05T11:22:28.911090+00:00","updated_at":"2026-07-05T11:22:28.911090+00:00"}