{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:MKDQXZXYQM65FXBX2JOLHBVWWG","short_pith_number":"pith:MKDQXZXY","schema_version":"1.0","canonical_sha256":"62870be6f8833dd2dc37d25cb386b6b195b386371bf4c3cf92c620d55f638080","source":{"kind":"arxiv","id":"2504.16074","version":2},"attestation_state":"computed","paper":{"title":"PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Anqi Lv, Binran Wang, Bohan Zhang, Boxuan Jing, Changkun Shao, Chencheng Tang, Chenyang Wang, Fan Cui, Feiyu Tao, Fengyuan Wang, Haoling Chang, Haoxu Zhang, Hua Xing Zhu, Jiahang Chen, Jiaming Ji, Jianxiang Li, Jiashen Wei, Jiawei Lin, Jing-Jun Zhang, Jingtian Zhang, Laifu Man, Minghao Li, Ming-Xing Luo, Muhan Zhang, Qihua Sun, Qi Liu, Qing-Hong Cao, Qiuhao Xiong, Shaoyang Guo, Shi Qiu, Shutao Zhang, Tianyu Luo, Tianyu Zhang, Weike Wang, Xianqi Yin, Xiaotian Li, Xingqi Xia, Xudong Tian, Yaodong Yang, Yi Hu, Yixuan Yin, Yuku Zhang, Yunbo Sun, Yushu Mu, Yutong Ren, Zeyu Cai, Zhangyi Liu, Zheyu Shen, Zhongxuan Li, Zhou Liang, Zhuo-Yang Song, Ziheng Zhou, Ziyang Ni, Zizhuo Fu","submitted_at":"2025-04-22T17:53:29Z","abstract_excerpt":"Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more toke"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.16074","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2025-04-22T17:53:29Z","cross_cats_sorted":[],"title_canon_sha256":"5699bcf5d7ec6de15994fd88eadec401032a96810532f6674fb3eca61ef08cbf","abstract_canon_sha256":"15cf134a205150d81bb2922e38f2273a0d66f6a1e7984eb5cb2959398716fa17"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:04:49.693612Z","signature_b64":"2h+XFO/zCBcNy3/q0+2YNLeBcH2+X6Fj8mc1IUDn+Pr+GfELBcKmCLiH1ezAOvasEswLCAE9235RG3OUhncICg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"62870be6f8833dd2dc37d25cb386b6b195b386371bf4c3cf92c620d55f638080","last_reissued_at":"2026-07-05T11:04:49.693081Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:04:49.693081Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Anqi Lv, Binran Wang, Bohan Zhang, Boxuan Jing, Changkun Shao, Chencheng Tang, Chenyang Wang, Fan Cui, Feiyu Tao, Fengyuan Wang, Haoling Chang, Haoxu Zhang, Hua Xing Zhu, Jiahang Chen, Jiaming Ji, Jianxiang Li, Jiashen Wei, Jiawei Lin, Jing-Jun Zhang, Jingtian Zhang, Laifu Man, Minghao Li, Ming-Xing Luo, Muhan Zhang, Qihua Sun, Qi Liu, Qing-Hong Cao, Qiuhao Xiong, Shaoyang Guo, Shi Qiu, Shutao Zhang, Tianyu Luo, Tianyu Zhang, Weike Wang, Xianqi Yin, Xiaotian Li, Xingqi Xia, Xudong Tian, Yaodong Yang, Yi Hu, Yixuan Yin, Yuku Zhang, Yunbo Sun, Yushu Mu, Yutong Ren, Zeyu Cai, Zhangyi Liu, Zheyu Shen, Zhongxuan Li, Zhou Liang, Zhuo-Yang Song, Ziheng Zhou, Ziyang Ni, Zizhuo Fu","submitted_at":"2025-04-22T17:53:29Z","abstract_excerpt":"Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more toke"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.16074","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.16074/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.16074","created_at":"2026-07-05T11:04:49.693147+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.16074v2","created_at":"2026-07-05T11:04:49.693147+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.16074","created_at":"2026-07-05T11:04:49.693147+00:00"},{"alias_kind":"pith_short_12","alias_value":"MKDQXZXYQM65","created_at":"2026-07-05T11:04:49.693147+00:00"},{"alias_kind":"pith_short_16","alias_value":"MKDQXZXYQM65FXBX","created_at":"2026-07-05T11:04:49.693147+00:00"},{"alias_kind":"pith_short_8","alias_value":"MKDQXZXY","created_at":"2026-07-05T11:04:49.693147+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":17,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.08034","citing_title":"Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07962","citing_title":"ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00276","citing_title":"Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00248","citing_title":"Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity","ref_index":88,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01538","citing_title":"MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2509.26574","citing_title":"Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2602.01203","citing_title":"Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2602.07064","citing_title":"OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2603.20633","citing_title":"Seed1.8 Model Card: Towards Generalized Real-World Agency","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2601.08584","citing_title":"Ministral 3","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2512.15745","citing_title":"LLaDA2.0: Scaling Up Diffusion Language Models to 100B","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02934","citing_title":"PolyReal: A Benchmark for Real-World Polymer Science Workflows","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09636","citing_title":"PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24443","citing_title":"PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23580","citing_title":"PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15411","citing_title":"PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27351","citing_title":"Heterogeneous Scientific Foundation Model Collaboration","ref_index":39,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG","json":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG.json","graph_json":"https://pith.science/api/pith-number/MKDQXZXYQM65FXBX2JOLHBVWWG/graph.json","events_json":"https://pith.science/api/pith-number/MKDQXZXYQM65FXBX2JOLHBVWWG/events.json","paper":"https://pith.science/paper/MKDQXZXY"},"agent_actions":{"view_html":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG","download_json":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG.json","view_paper":"https://pith.science/paper/MKDQXZXY","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.16074&json=true","fetch_graph":"https://pith.science/api/pith-number/MKDQXZXYQM65FXBX2JOLHBVWWG/graph.json","fetch_events":"https://pith.science/api/pith-number/MKDQXZXYQM65FXBX2JOLHBVWWG/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG/action/timestamp_anchor","attest_storage":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG/action/storage_attestation","attest_author":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG/action/author_attestation","sign_citation":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG/action/citation_signature","submit_replication":"https://pith.science/pith/MKDQXZXYQM65FXBX2JOLHBVWWG/action/replication_record"}},"created_at":"2026-07-05T11:04:49.693147+00:00","updated_at":"2026-07-05T11:04:49.693147+00:00"}