Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T10:03:58.971585Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 100 of 287 outbound references and 100 inbound Pith citation observations for arXiv:2304.06364.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T10:03:58.971585Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:30:12.075122Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 287 outbound references displayed
External citation measurements
61
pith, observed 2026-08-05T02:28:24.338817Z
Observation eceddf27-e116-4dcd-b815-c38f60c9ab3e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4827eff9-f0fb-4b93-82b5-8a54479a91d1 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 49b18a50-e937-4c9f-ba12-bafaa6da526f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , publisher =
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9cdd79b2-8143-443d-bd91-395de9f7aba4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Communications of the ACM , volume=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2e3f82e7-0e7e-452b-ab13-456dd0af2239 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a5976102-e454-45f5-bef5-0209559807c9 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79e3a7cd-34d1-41b2-9b64-cc396f12cdae · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Stoica, Ion and Xing, Eric P
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5535508e-df1d-4a78-a604-d68fa1493f2d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d06274b9-7630-4fc8-8fe0-4804566d947e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation be3ca354-5d71-494c-bc43-6a070579f636 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10aab787-1efa-4875-a11b-5002787a5e8f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d52f2b8-57f3-4bf8-b613-295730b44438 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: NAACL 2022 , pages=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7c539422-9be9-407e-8675-e7109af50f3c · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12608aff-31cb-485d-9b4e-c874ca44d5f0 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cba04094-6373-4beb-9698-012f02899b02 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69ae1072-7690-436f-ab93-049159b4b0cc · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 462520c3-05d5-427b-8e64-d395cd247bc9 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sort , volume=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b720d16e-98a0-4801-a945-97598cce481d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5fb47e10-56e5-4c44-8597-50c40d8cc416 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of AAAI , year=
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3053355e-a2ac-4da2-8062-f24b2f753da5 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , eprint=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82aaeada-ef03-40fd-8be4-3c3468aacd3a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2019 , publisher=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a1b4580-83ea-440c-a13f-f139d49f0e0b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f4235ec8-8e2f-4aa7-a7b0-782de6e941bb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 218d80ba-ae60-4f8a-8b1f-ac5ab851a87a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 44
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 913d7210-379c-4fc0-9c98-6ad335545145 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ec90262-7d23-4b5b-9803-b25d309f2201 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in Neural Information Processing Systems , volume=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1598b0a2-3d28-47b5-baa4-6946bcaf02be · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Conference on Empirical Methods in Natural Language Processing , year=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5f51fdd2-55ab-46d8-a328-85303efbe9bb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Available at SSRN , year=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b69fc3a0-1349-4b6b-81c7-acf67b5ccbf8 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2022 , eprint=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d3c9655-3c2c-4664-9efe-e4c9171ba505 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Open llm leaderboard
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f6b526fd-9ac2-4c98-9334-45eab891f8cb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 691af668-411c-430d-862d-bd1c621472ab · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Language models are few-shot learners
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 993a77f4-6bee-4aa9-9627-254650120217 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b48e4779-0801-4903-a479-770cb5d1cfa4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gonzalez, Ion Stoica, and Eric P
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f24f9ad8-ec22-4202-87ae-48da610e86b4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chatgpt goes to law school
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5f8f8b8b-7ecd-48e8-b82e-c32b3448bb82 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Scaling Instruction-Finetuned Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1b303851-ad7f-403c-8d98-99f0d9f4fc6b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf54d8b5-bc7b-4a6f-9f4b-0fb30d85ed55 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models S ent E val: An evaluation toolkit for universal sentence representations
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 33bc5809-39b6-4c1d-96f3-70b2dfa37f58 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a93b10d7-3fa4-4c49-9261-d467bef46773 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bold: Dataset and metrics for measuring biases in open-ended language generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a877411-e333-49a6-b87c-17c3ec2397e9 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Glm: General language model pretraining with autoregressive blank infilling
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 74321b76-3ab6-4a94-aa01-2c563ab54fc0 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 17fef00c-fb8a-49b0-8274-5f2caa630e97 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ef4eacd-2452-4ba7-9353-6ff6c347e65f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring mathematical problem solving with the math dataset
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bf17a0e3-d25e-4d47-a60b-66fa52e7268b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring Massive Multitask Language Understanding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11fd4998-b610-4a65-b24d-36b34abf0ca1 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7436fb72-2b7d-4a81-9515-a5c7e6526867 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Solving quantitative reasoning problems with language models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5afb741b-c594-49c9-b21f-5d1f3cd3f1de · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Program induction by rationale generation: Learning to solve and explain algebraic word problems
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 492a83d3-6c7b-4d4d-abe2-073dc79102e7 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Logiqa: a challenge dataset for machine reading comprehension with logical reasoning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8cc1fb61-b6e1-44a5-af57-97a2d314fdca · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Rebooting AI: Building artificial intelligence we can trust
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 74af0108-3cd6-40ce-b897-eaeba7dceb9f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The Natural Language Decathlon: Multitask Learning as Question Answering
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06e8bd09-ba43-41ea-a2c2-61486f56a500 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d1ac8ec6-ae06-4a8b-904c-f91c8b6cb2e4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gpt-4 technical report
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6151d5f0-9bb4-4d86-98e9-fc7ff67c8c8f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Training language models to follow instructions with human feedback
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b9a2912-9140-4006-8969-ab5871388564 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ab8fe06f-5254-42fe-9281-58575fc4fade · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Squad: 100,000+ questions for machine comprehension of text
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18d8e589-4ddc-4a2d-b85d-e354f30e5b7e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Internlm
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d2e3e2d2-394f-4046-b05f-bd1c02611323 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models FEVER: a large-scale dataset for Fact Extraction and VERification
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9bf6514a-5cf1-4f39-86fb-05d3c8f08353 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12435b9b-5f55-49a2-9727-4ae09c88d1a1 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2018
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3f086877-2f1e-44f2-91cd-ee86bbf2cd0e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Superglue: A stickier benchmark for general-purpose language understanding systems
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a9d818a3-b884-4b38-bfef-3b7826380d7d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models From lsat: The progress and challenges of complex reasoning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e7e0ed83-3f32-4c6c-b3f2-afe2dbd938ee · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bcf85aca-e9c9-490c-a91b-5ee4cc34586a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models OPT: Open Pre-trained Transformer Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4c12a26a-adaf-41eb-9701-0e88483c06e6 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Automatic Chain of Thought Prompting in Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3a2a199e-9a60-4dc3-8da9-f61c1846653d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jec-qa: A legal-domain question answering dataset
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd69c749-c915-49a7-b05f-c5c109380e87 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Analytical reasoning of text
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 03fa877b-326b-4c22-ba46-2fae0045a960 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ab3b8154-9675-4642-a5bf-c6c1aa7d393b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a4e67bf9-8441-4a5a-a2ec-0fa2fd704679 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 032ef19b-7e0b-4b5e-8f6a-c1fd2dc0406c · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models H ate BERT : Retraining BERT for Abusive Language Detection in E nglish
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6bbb9fed-d274-4707-9a09-a45bd33f599e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d4658c11-1b9c-47d1-9f33-2e05499b2b5a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 830f09bf-d889-4895-8a9c-7c83090d472f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models DALC : the D utch Abusive Language Corpus
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 90eceacf-ec56-47d7-af81-ce4bb74805ec · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Dulal, Saurab and Koirala, Diwa
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f8d4ca18-85fa-477f-8b1e-992095f15976 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 87f1d437-2827-454e-aa09-01332a2b4b52 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a438ce5e-44e9-4f87-9313-281e8e6291ff · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Improving Counterfactual Generation for Fair Hate Speech Detection
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a4de569-3b17-4519-9810-27d09fea88bf · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 96de1339-8b1d-4989-a1de-8005801b0fae · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Mitigating Biases in Toxic Language Detection through Invariant Rationalization
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8ae52ef-7779-42a0-812d-a3f4844cfe7a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 274c958d-26ad-434c-9fe7-e55115ea91aa · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bb2f7b6f-5349-4428-9352-f12a84e9a4bb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Context Sensitivity Estimation in Toxicity Detection
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b9199c39-3889-4e5a-bb1c-95885278251a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 65c9bad9-4fa3-416c-bb00-4ecc618df2f4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5cc10fd-b4ab-415c-85b1-fee0ab1328fb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and H \'e bert-Dufresne, Laurent and Roth, Allison M
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 223e06c0-2ce6-44c5-899f-22a22fd1efce · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Targets and Aspects in Social Media Hate Speech
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68ce1835-4606-4bb5-b487-31f1d36ecf11 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Abusive Language on Social Media Through the Legal Looking Glass
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fec804b5-2b12-44b5-a068-c28e9f61b2fa · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b72c17f0-1090-4f73-be34-172e5f03780a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a5ee5f4c-bd93-4234-aefa-7c4278cb278d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Racist or Sexist Meme? Classifying Memes beyond Hateful
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 470c38d8-3f6f-48ca-a061-7f05031e48aa · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e3570e05-a656-4a15-af58-58cb564b2b72 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bca6162d-039b-4d9c-8041-29ee21286cbd · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Text Simplification for Comprehension-based Question-Answering
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18b935fb-76c2-44c0-a730-b029aef26efc · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6cb98123-eb58-467c-9357-e1cb0cb67d68 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Detecting Depression in T hai Blog Posts: a Dataset and a Baseline
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 26398b59-dc55-4a42-be27-20a95b9237b4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Keyphrase Extraction with Incomplete Annotated Training Data
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d4bef67-af6d-4ce4-89ed-a86e3e0666c7 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c5c900f-130c-4450-a8d0-0ee8507ac09b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c1b92dad-805a-47d2-943c-dc989a98b5ff · inbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 96043340-8524-4b05-94fc-ebaf9b2cb6df · inbound
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ffa988c-4a1e-4c9a-9fca-c8d94dfff545 · inbound
Baichuan 2: Open Large-scale Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5134d17-bbc2-42da-b71f-ee7cedb4bc28 · inbound
Mistral 7B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9644ed94-4617-459d-9b84-777600b65e97 · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 39ee5a74-f36c-49a6-928f-b280709feedf · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c718a91d-569f-407a-aaa2-36b2e00374d4 · inbound
The Falcon Series of Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 226
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 188b5ec0-8e54-4362-9e77-31e427c34247 · inbound
GPT-4V(ision) is a Generalist Web Agent, if Grounded AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 29c50967-2771-46e4-b630-edde26399d5b · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 321e39c8-7012-4f44-9288-52993abeb961 · inbound
Mixtral of Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4283d2d9-6d40-47f3-9f11-39b38cf5d953 · inbound
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2906ed29-c1a8-46f8-91f8-639cc21f41b8 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7403b0e0-3e8f-4a45-ada0-ab469dac8475 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82f29955-ba9e-46ba-a460-84dd8ebe4ddd · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9b4d48db-aefb-4c5d-8d47-bff912511c6e · inbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5a7a2327-af5a-4761-b4cc-95eb56ddd36f · inbound
DataComp-LM: In search of the next generation of training sets for language models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 43ffd89b-4280-4515-87bd-f45c3388ec8a · inbound
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8f1175dd-89d8-48f6-8840-c3271dc946ba · inbound
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc643947-5dc4-4ac4-b4fc-a6db1271f161 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3f745850-9722-42c9-9057-9f449b0d9d04 · inbound
A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b0186bcc-c37c-48bf-856a-0451e49dfe8d · inbound
Pixtral 12B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fc4a806a-8251-421b-b9e7-8bbff3c7ca6d · inbound
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647feef1-e139-46ad-a26a-5c623538ab62 · inbound
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f964a509-040d-4db2-8c1f-3416831f6264 · inbound
Ultra-Sparse Memory Network AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a6e5e31-3d7e-4320-9336-b964bec2382f · inbound
WavChat: A Survey of Spoken Dialogue Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 256
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d22f2c5c-8282-44d1-80fa-3335a5a762e1 · inbound
AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933e5b7d-0a12-4a5c-856b-ba8c7385eca9 · inbound
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d56313-9574-4b1c-ba69-f46957e94b90 · inbound
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca57eb6a-6339-4196-ac25-3e59ea525085 · inbound
Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 62b94ffd-131f-426f-a166-edd697b32cc1 · inbound
MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7431a7ec-604a-4bc7-bd47-a25fc223e0b3 · inbound
A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6a8eaea-9549-4d9b-a2f0-5d5f2227275f · inbound
Humanity's Last Exam AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5f9bd9c4-c29c-48c5-8b43-79ba6ed27c83 · inbound
SedarEval: Automated Evaluation using Self-Adaptive Rubrics AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b974884a-c6ad-4b33-8195-5f424185ed2b · inbound
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa1fa42-6884-49d6-8aee-e3ace2c9d895 · inbound
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04cf8462-dd02-4483-b395-4a7ed7672cdc · inbound
Minerva: A Programmable Memory Test Benchmark for Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28b1695-489b-4888-b3ea-8132217b6286 · inbound
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f86f630-1d5a-440f-a596-014d8a74a2cc · inbound
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d8497d-01d9-4a05-9ebb-a9209b335a59 · inbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d83edae-8300-418e-804e-c3c73be701fb · inbound
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8ebcec-b078-4056-813b-4508d264d1df · inbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16882dc6-a0b9-46f2-89f8-3a2bbe4acf2b · inbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c1b08166-55b0-4707-a035-6f4afef9ae04 · inbound
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509a881e-5b88-4ba7-a0a5-e4b37d78091d · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 72835dbc-ee36-4493-b095-d9e9e6bb973f · inbound
Computational Reasoning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ebe5d4-d1d6-4bed-81cb-46131d2f3664 · inbound
LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37fdae11-9ece-4c3e-95b9-a05642bae739 · inbound
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c043601c-eeb7-43f9-8df0-5a99280050f5 · inbound
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b1e8266-5f36-4f91-a04c-7406a32ca961 · inbound
Learnware of Language Models: Specialized Small Language Models Can Do Big AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba8898d5-c2c4-4162-b806-bda2c55114b5 · inbound
STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d561a2b8-71af-4af8-9935-20e64dd06527 · inbound
SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e092f6-a6c0-4566-b94f-c61256041fe3 · inbound
Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad62654-f98b-45a0-bf83-193d3f8d7827 · inbound
Scalable Complexity Control Facilitates Reasoning Ability of LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf12970-07b3-43f0-bc47-85ff7762fa00 · inbound
PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a51f984-734b-48f9-8dd8-7f7e50eb66dd · inbound
MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14fa1b8e-c9e3-4b9e-bee5-e79f7b5912ee · inbound
PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075b5f16-5f97-4571-9558-5dd63cd38c43 · inbound
dots.llm1 Technical Report AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e48c119-0629-4f5a-8bfa-b655e45f50ab · inbound
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79826531-a36f-4f8f-b079-f7eb5785f4f9 · inbound
Towards Efficient and Effective Alignment of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 212
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a841bc-ceb5-48fe-aa7d-35c89d7577a5 · inbound
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 73bcba91-64fb-4d43-a1be-866f254414ef · inbound
SciDA: Scientific Dynamic Assessor of LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aadeb1d5-8c85-45ca-b4f5-08ad8ebee611 · inbound
EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb23d56-9a47-4491-989f-80b31d02cb38 · inbound
Enterprise Large Language Model Evaluation Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50807dba-7958-4801-9fed-e72e7dd5388d · inbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e3890a-8731-4577-bf06-166dddb40dd1 · inbound
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d67280-fd08-438e-b4fd-f0f2b2311890 · inbound
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 200
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7aba488e-f72e-4576-8720-e8087df009fc · inbound
Think Clearly: Improving Reasoning via Redundant Token Pruning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4642143-8cdc-4e24-b437-e5b07ca4e934 · inbound
HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db48035-c307-4261-a317-ce17b6f69174 · inbound
Language Models Improve When Pretraining Data Matches Target Tasks AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601dc364-a39a-4d5d-bd0c-df744309b6d7 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 130
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336f7a07-a1a9-4885-9ce8-204c9c234a51 · inbound
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad27a3f7-0084-4515-b40b-a8492c15271f · inbound
Kimi K2: Open Agentic Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a5e72e7-5d37-4721-9dd3-7f6e9217b879 · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ec54b7b3-8a57-4a66-9565-59282b8f0699 · inbound
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b9a7dc-8f49-4687-b24a-d7210756fdae · inbound
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dd2f60a-b4a4-4803-8b55-5ed589bf5855 · inbound
ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54363329-7bcf-438b-bfc4-dce5d3080fcb · inbound
Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7828047d-0797-4a6d-9b9a-37881801c231 · inbound
Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dde716a-4b75-4d50-b822-de3b1a0e9cca · inbound
UQ: Assessing Language Models on Unsolved Questions AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19697ec8-1d39-4cfa-9c0f-2a9f8313185c · inbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 060cab31-313c-40c3-9f4c-a3b8396c4f00 · inbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ed317f5-0155-49f9-b430-d44a63db8f32 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5883d43e-d7bb-472c-a274-a212bbdc59ff · inbound
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 88f763f2-ef0f-4d89-af4b-a892cf39b52a · inbound
Dr.LLM: Dynamic Layer Routing in LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d083d6e9-23c5-40ac-9d7e-5de221cd30fd · inbound
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ba35216e-37b0-419b-8430-b27ba4c611e8 · inbound
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a77233d4-f2aa-48e9-b4e8-4bad0640a905 · inbound
Ministral 3 AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3a0c1391-dc27-491e-a871-4bc359d5d02f · inbound
Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0910937-4558-4948-90de-e21ad3005511 · inbound
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea78a68-21b2-43a5-a1fe-e961d5eda5a1 · inbound
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9ca77fdd-8654-4f51-868d-f0caa031e0fe · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4dac4e0d-3f34-4c37-95a9-13fc45d49019 · inbound
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58ab0514-50af-41b4-88be-d962e572afc3 · inbound
"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3d130690-2278-4685-82da-9f183a0feb22 · inbound
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 572da502-e57c-4723-8c76-364e78efbe62 · inbound
Confidence Calibration in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31646264-be84-400e-a269-aefcba1406d4 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9176bcae-832f-49ed-944d-319fa2b21f7f · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23a1af6-d019-4ad3-a5ed-41e18f5a409f · inbound
The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cc2f04a1-ac84-44b1-a783-a66334d63b63 · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1aeeb5e2-0e18-4e45-92f0-ee6cf38d01b4 · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 247
Source-reported events for the cited work
Unavailable: canonical work link unavailable.