{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:SM2TNWV4GST3LIRQX3G4KCUMDG","short_pith_number":"pith:SM2TNWV4","schema_version":"1.0","canonical_sha256":"933536dabc34a7b5a230becdc50a8c19b41e4e09acfcc3af4587b90731990cd2","source":{"kind":"arxiv","id":"2307.03109","version":9},"attestation_state":"computed","paper":{"title":"A Survey on Evaluation of Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Cunxiang Wang, Hao Chen, Jindong Wang, Kaijie Zhu, Linyi Yang, Philip S. Yu, Qiang Yang, Wei Ye, Xiaoyuan Yi, Xing Xie, Xu Wang, Yi Chang, Yidong Wang, Yuan Wu, Yue Zhang, Yupeng Chang","submitted_at":"2023-07-06T16:28:35Z","abstract_excerpt":"Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions:"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2307.03109","kind":"arxiv","version":9},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2023-07-06T16:28:35Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"ef701aee76467a8bf93611aa9ae1fcefb08f12032877534fcb4f5777d845ff4c","abstract_canon_sha256":"8d80b8b1b9f4e5bc9999b0afba0f055511feaaeca7c4157a544a597360bf7bba"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:28:44.966947Z","signature_b64":"pTSsqT3+ftl6q+MVCZxj/DO6uT6iJVv7yP9EcG7tw9Rh6cDhoTR9j1Rv3WjP/5SG6D66J2X0+lMDNkoRtIyxCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"933536dabc34a7b5a230becdc50a8c19b41e4e09acfcc3af4587b90731990cd2","last_reissued_at":"2026-07-05T07:28:44.966451Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:28:44.966451Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"A Survey on Evaluation of Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Cunxiang Wang, Hao Chen, Jindong Wang, Kaijie Zhu, Linyi Yang, Philip S. Yu, Qiang Yang, Wei Ye, Xiaoyuan Yi, Xing Xie, Xu Wang, Yi Chang, Yidong Wang, Yuan Wu, Yue Zhang, Yupeng Chang","submitted_at":"2023-07-06T16:28:35Z","abstract_excerpt":"Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions:"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2307.03109","kind":"arxiv","version":9},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2307.03109/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2307.03109","created_at":"2026-07-05T07:28:44.966517+00:00"},{"alias_kind":"arxiv_version","alias_value":"2307.03109v9","created_at":"2026-07-05T07:28:44.966517+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2307.03109","created_at":"2026-07-05T07:28:44.966517+00:00"},{"alias_kind":"pith_short_12","alias_value":"SM2TNWV4GST3","created_at":"2026-07-05T07:28:44.966517+00:00"},{"alias_kind":"pith_short_16","alias_value":"SM2TNWV4GST3LIRQ","created_at":"2026-07-05T07:28:44.966517+00:00"},{"alias_kind":"pith_short_8","alias_value":"SM2TNWV4","created_at":"2026-07-05T07:28:44.966517+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":22,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.06157","citing_title":"LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05734","citing_title":"When AI Says It Feels","ref_index":151,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04733","citing_title":"Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24298","citing_title":"An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10106","citing_title":"What makes a harness a harness: necessary and sufficient conditions for an agent harness","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2311.07911","citing_title":"Instruction-Following Evaluation for Large Language Models","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2401.02458","citing_title":"Data-Centric Foundation Models in Computational Healthcare: A Survey","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2503.01835","citing_title":"Primus: Enforcing Attention Usage for 3D Medical Image Segmentation","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15104","citing_title":"From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2505.19237","citing_title":"Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2508.05452","citing_title":"LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2603.21011","citing_title":"ALL-FEM: Agentic Large Language models Fine-tuned for Finite Element Methods","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14004","citing_title":"Conditional Attribute Estimation with Autoregressive Sequence Models","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16421","citing_title":"Measuring Representation Robustness in Large Language Models for Geometry","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24544","citing_title":"STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06423","citing_title":"Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04733","citing_title":"Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2604.07369","citing_title":"The Role of Emotional Stimuli and Intensity in Shaping Large Language Model Behavior","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2604.06666","citing_title":"A Graph-Enhanced Defense Framework for Explainable Fake News Detection with LLM","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13371","citing_title":"Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01048","citing_title":"Compared to What? Baselines and Metrics for Counterfactual Prompting","ref_index":117,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27351","citing_title":"Heterogeneous Scientific Foundation Model Collaboration","ref_index":7,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG","json":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG.json","graph_json":"https://pith.science/api/pith-number/SM2TNWV4GST3LIRQX3G4KCUMDG/graph.json","events_json":"https://pith.science/api/pith-number/SM2TNWV4GST3LIRQX3G4KCUMDG/events.json","paper":"https://pith.science/paper/SM2TNWV4"},"agent_actions":{"view_html":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG","download_json":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG.json","view_paper":"https://pith.science/paper/SM2TNWV4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2307.03109&json=true","fetch_graph":"https://pith.science/api/pith-number/SM2TNWV4GST3LIRQX3G4KCUMDG/graph.json","fetch_events":"https://pith.science/api/pith-number/SM2TNWV4GST3LIRQX3G4KCUMDG/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG/action/timestamp_anchor","attest_storage":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG/action/storage_attestation","attest_author":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG/action/author_attestation","sign_citation":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG/action/citation_signature","submit_replication":"https://pith.science/pith/SM2TNWV4GST3LIRQX3G4KCUMDG/action/replication_record"}},"created_at":"2026-07-05T07:28:44.966517+00:00","updated_at":"2026-07-05T07:28:44.966517+00:00"}