Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T10:03:58.971585Z
Paper Citation Record · LEDGER
As of 2 August 2026, this Paper Citation Record lists 100 of 287 outbound references and 53 inbound Pith citation observations for arXiv:2304.06364.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T10:03:58.971585Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T12:41:24.460278Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
100 of 287 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eceddf27-e116-4dcd-b815-c38f60c9ab3e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4827eff9-f0fb-4b93-82b5-8a54479a91d1 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 49b18a50-e937-4c9f-ba12-bafaa6da526f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , publisher =
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9cdd79b2-8143-443d-bd91-395de9f7aba4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Communications of the ACM , volume=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2e3f82e7-0e7e-452b-ab13-456dd0af2239 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a5976102-e454-45f5-bef5-0209559807c9 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 79e3a7cd-34d1-41b2-9b64-cc396f12cdae · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Stoica, Ion and Xing, Eric P
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5535508e-df1d-4a78-a604-d68fa1493f2d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d06274b9-7630-4fc8-8fe0-4804566d947e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation be3ca354-5d71-494c-bc43-6a070579f636 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 10aab787-1efa-4875-a11b-5002787a5e8f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0d52f2b8-57f3-4bf8-b613-295730b44438 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: NAACL 2022 , pages=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7c539422-9be9-407e-8675-e7109af50f3c · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 12608aff-31cb-485d-9b4e-c874ca44d5f0 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cba04094-6373-4beb-9698-012f02899b02 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 69ae1072-7690-436f-ab93-049159b4b0cc · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 462520c3-05d5-427b-8e64-d395cd247bc9 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sort , volume=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b720d16e-98a0-4801-a945-97598cce481d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5fb47e10-56e5-4c44-8597-50c40d8cc416 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of AAAI , year=
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3053355e-a2ac-4da2-8062-f24b2f753da5 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , eprint=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 82aaeada-ef03-40fd-8be4-3c3468aacd3a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2019 , publisher=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9a1b4580-83ea-440c-a13f-f139d49f0e0b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f4235ec8-8e2f-4aa7-a7b0-782de6e941bb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 218d80ba-ae60-4f8a-8b1f-ac5ab851a87a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 44
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 913d7210-379c-4fc0-9c98-6ad335545145 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7ec90262-7d23-4b5b-9803-b25d309f2201 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in Neural Information Processing Systems , volume=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1598b0a2-3d28-47b5-baa4-6946bcaf02be · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Conference on Empirical Methods in Natural Language Processing , year=
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5f51fdd2-55ab-46d8-a328-85303efbe9bb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Available at SSRN , year=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b69fc3a0-1349-4b6b-81c7-acf67b5ccbf8 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2022 , eprint=
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5d3c9655-3c2c-4664-9efe-e4c9171ba505 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Open llm leaderboard
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f6b526fd-9ac2-4c98-9334-45eab891f8cb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 691af668-411c-430d-862d-bd1c621472ab · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Language models are few-shot learners
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 993a77f4-6bee-4aa9-9627-254650120217 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b48e4779-0801-4903-a479-770cb5d1cfa4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gonzalez, Ion Stoica, and Eric P
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f24f9ad8-ec22-4202-87ae-48da610e86b4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chatgpt goes to law school
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5f8f8b8b-7ecd-48e8-b82e-c32b3448bb82 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Scaling Instruction-Finetuned Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1b303851-ad7f-403c-8d98-99f0d9f4fc6b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf54d8b5-bc7b-4a6f-9f4b-0fb30d85ed55 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models S ent E val: An evaluation toolkit for universal sentence representations
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 33bc5809-39b6-4c1d-96f3-70b2dfa37f58 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a93b10d7-3fa4-4c49-9261-d467bef46773 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bold: Dataset and metrics for measuring biases in open-ended language generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2a877411-e333-49a6-b87c-17c3ec2397e9 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Glm: General language model pretraining with autoregressive blank infilling
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 74321b76-3ab6-4a94-aa01-2c563ab54fc0 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 17fef00c-fb8a-49b0-8274-5f2caa630e97 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5ef4eacd-2452-4ba7-9353-6ff6c347e65f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring mathematical problem solving with the math dataset
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf17a0e3-d25e-4d47-a60b-66fa52e7268b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring Massive Multitask Language Understanding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 11fd4998-b610-4a65-b24d-36b34abf0ca1 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7436fb72-2b7d-4a81-9515-a5c7e6526867 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Solving quantitative reasoning problems with language models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5afb741b-c594-49c9-b21f-5d1f3cd3f1de · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Program induction by rationale generation: Learning to solve and explain algebraic word problems
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 492a83d3-6c7b-4d4d-abe2-073dc79102e7 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Logiqa: a challenge dataset for machine reading comprehension with logical reasoning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8cc1fb61-b6e1-44a5-af57-97a2d314fdca · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Rebooting AI: Building artificial intelligence we can trust
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 74af0108-3cd6-40ce-b897-eaeba7dceb9f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The Natural Language Decathlon: Multitask Learning as Question Answering
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 06e8bd09-ba43-41ea-a2c2-61486f56a500 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d1ac8ec6-ae06-4a8b-904c-f91c8b6cb2e4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gpt-4 technical report
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6151d5f0-9bb4-4d86-98e9-fc7ff67c8c8f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Training language models to follow instructions with human feedback
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0b9a2912-9140-4006-8969-ab5871388564 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ab8fe06f-5254-42fe-9281-58575fc4fade · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Squad: 100,000+ questions for machine comprehension of text
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 18d8e589-4ddc-4a2d-b85d-e354f30e5b7e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Internlm
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d2e3e2d2-394f-4046-b05f-bd1c02611323 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models FEVER: a large-scale dataset for Fact Extraction and VERification
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9bf6514a-5cf1-4f39-86fb-05d3c8f08353 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 12435b9b-5f55-49a2-9727-4ae09c88d1a1 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2018
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3f086877-2f1e-44f2-91cd-ee86bbf2cd0e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Superglue: A stickier benchmark for general-purpose language understanding systems
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a9d818a3-b884-4b38-bfef-3b7826380d7d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models From lsat: The progress and challenges of complex reasoning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e7e0ed83-3f32-4c6c-b3f2-afe2dbd938ee · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bcf85aca-e9c9-490c-a91b-5ee4cc34586a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models OPT: Open Pre-trained Transformer Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4c12a26a-adaf-41eb-9701-0e88483c06e6 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Automatic Chain of Thought Prompting in Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3a2a199e-9a60-4dc3-8da9-f61c1846653d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jec-qa: A legal-domain question answering dataset
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bd69c749-c915-49a7-b05f-c5c109380e87 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Analytical reasoning of text
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 03fa877b-326b-4c22-ba46-2fae0045a960 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ab3b8154-9675-4642-a5bf-c6c1aa7d393b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a4e67bf9-8441-4a5a-a2ec-0fa2fd704679 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 032ef19b-7e0b-4b5e-8f6a-c1fd2dc0406c · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models H ate BERT : Retraining BERT for Abusive Language Detection in E nglish
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6bbb9fed-d274-4707-9a09-a45bd33f599e · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d4658c11-1b9c-47d1-9f33-2e05499b2b5a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 830f09bf-d889-4895-8a9c-7c83090d472f · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models DALC : the D utch Abusive Language Corpus
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 90eceacf-ec56-47d7-af81-ce4bb74805ec · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Dulal, Saurab and Koirala, Diwa
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f8d4ca18-85fa-477f-8b1e-992095f15976 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 87f1d437-2827-454e-aa09-01332a2b4b52 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a438ce5e-44e9-4f87-9313-281e8e6291ff · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Improving Counterfactual Generation for Fair Hate Speech Detection
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9a4de569-3b17-4519-9810-27d09fea88bf · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 96de1339-8b1d-4989-a1de-8005801b0fae · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Mitigating Biases in Toxic Language Detection through Invariant Rationalization
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c8ae52ef-7779-42a0-812d-a3f4844cfe7a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 274c958d-26ad-434c-9fe7-e55115ea91aa · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bb2f7b6f-5349-4428-9352-f12a84e9a4bb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Context Sensitivity Estimation in Toxicity Detection
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b9199c39-3889-4e5a-bb1c-95885278251a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 65c9bad9-4fa3-416c-bb00-4ecc618df2f4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f5cc10fd-b4ab-415c-85b1-fee0ab1328fb · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and H \'e bert-Dufresne, Laurent and Roth, Allison M
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 223e06c0-2ce6-44c5-899f-22a22fd1efce · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Targets and Aspects in Social Media Hate Speech
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 68ce1835-4606-4bb5-b487-31f1d36ecf11 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Abusive Language on Social Media Through the Legal Looking Glass
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fec804b5-2b12-44b5-a068-c28e9f61b2fa · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b72c17f0-1090-4f73-be34-172e5f03780a · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a5ee5f4c-bd93-4234-aefa-7c4278cb278d · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Racist or Sexist Meme? Classifying Memes beyond Hateful
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 470c38d8-3f6f-48ca-a061-7f05031e48aa · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e3570e05-a656-4a15-af58-58cb564b2b72 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bca6162d-039b-4d9c-8041-29ee21286cbd · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Text Simplification for Comprehension-based Question-Answering
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 18b935fb-76c2-44c0-a730-b029aef26efc · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 6cb98123-eb58-467c-9357-e1cb0cb67d68 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Detecting Depression in T hai Blog Posts: a Dataset and a Baseline
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 26398b59-dc55-4a42-be27-20a95b9237b4 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Keyphrase Extraction with Incomplete Annotated Training Data
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0d4bef67-af6d-4ce4-89ed-a86e3e0666c7 · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5c5c900f-130c-4450-a8d0-0ee8507ac09b · outbound
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c1b92dad-805a-47d2-943c-dc989a98b5ff · inbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 96043340-8524-4b05-94fc-ebaf9b2cb6df · inbound
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4ffa988c-4a1e-4c9a-9fca-c8d94dfff545 · inbound
Baichuan 2: Open Large-scale Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f5134d17-bbc2-42da-b71f-ee7cedb4bc28 · inbound
Mistral 7B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9644ed94-4617-459d-9b84-777600b65e97 · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 39ee5a74-f36c-49a6-928f-b280709feedf · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c718a91d-569f-407a-aaa2-36b2e00374d4 · inbound
The Falcon Series of Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 226
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 188b5ec0-8e54-4362-9e77-31e427c34247 · inbound
GPT-4V(ision) is a Generalist Web Agent, if Grounded AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 29c50967-2771-46e4-b630-edde26399d5b · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 321e39c8-7012-4f44-9288-52993abeb961 · inbound
Mixtral of Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4283d2d9-6d40-47f3-9f11-39b38cf5d953 · inbound
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2906ed29-c1a8-46f8-91f8-639cc21f41b8 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7403b0e0-3e8f-4a45-ada0-ab469dac8475 · inbound
DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 82f29955-ba9e-46ba-a460-84dd8ebe4ddd · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9b4d48db-aefb-4c5d-8d47-bff912511c6e · inbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5a7a2327-af5a-4761-b4cc-95eb56ddd36f · inbound
DataComp-LM: In search of the next generation of training sets for language models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 220
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 43ffd89b-4280-4515-87bd-f45c3388ec8a · inbound
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8f1175dd-89d8-48f6-8840-c3271dc946ba · inbound
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bc643947-5dc4-4ac4-b4fc-a6db1271f161 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3f745850-9722-42c9-9057-9f449b0d9d04 · inbound
A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b0186bcc-c37c-48bf-856a-0451e49dfe8d · inbound
Pixtral 12B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ca57eb6a-6339-4196-ac25-3e59ea525085 · inbound
Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d6a8eaea-9549-4d9b-a2f0-5d5f2227275f · inbound
Humanity's Last Exam AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2c8ebcec-b078-4056-813b-4508d264d1df · inbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 16882dc6-a0b9-46f2-89f8-3a2bbe4acf2b · inbound
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 509a881e-5b88-4ba7-a0a5-e4b37d78091d · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 73a841bc-ceb5-48fe-aa7d-35c89d7577a5 · inbound
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 38d67280-fd08-438e-b4fd-f0f2b2311890 · inbound
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 200
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ad27a3f7-0084-4515-b40b-a8492c15271f · inbound
Kimi K2: Open Agentic Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2a5e72e7-5d37-4721-9dd3-7f6e9217b879 · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1ed317f5-0155-49f9-b430-d44a63db8f32 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5883d43e-d7bb-472c-a274-a212bbdc59ff · inbound
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 88f763f2-ef0f-4d89-af4b-a892cf39b52a · inbound
Dr.LLM: Dynamic Layer Routing in LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d083d6e9-23c5-40ac-9d7e-5de221cd30fd · inbound
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ba35216e-37b0-419b-8430-b27ba4c611e8 · inbound
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a77233d4-f2aa-48e9-b4e8-4bad0640a905 · inbound
Ministral 3 AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation fea78a68-21b2-43a5-a1fe-e961d5eda5a1 · inbound
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9ca77fdd-8654-4f51-868d-f0caa031e0fe · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4dac4e0d-3f34-4c37-95a9-13fc45d49019 · inbound
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 58ab0514-50af-41b4-88be-d962e572afc3 · inbound
"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3d130690-2278-4685-82da-9f183a0feb22 · inbound
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 572da502-e57c-4723-8c76-364e78efbe62 · inbound
Confidence Calibration in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31646264-be84-400e-a269-aefcba1406d4 · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 9176bcae-832f-49ed-944d-319fa2b21f7f · inbound
Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23a1af6-d019-4ad3-a5ed-41e18f5a409f · inbound
The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cc2f04a1-ac84-44b1-a783-a66334d63b63 · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1aeeb5e2-0e18-4e45-92f0-ee6cf38d01b4 · inbound
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 247
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ddf1b61-ddc0-44ab-9adf-c73eade29cd2 · inbound
BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bfa5fb53-4e67-4b36-8283-784db1cf8bab · inbound
BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 06aec8e5-42ad-43cd-a44f-86e5557cb9fc · inbound
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 230
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0e387e7d-72a1-4955-909c-ca676c74e572 · inbound
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 229
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b3832bec-1717-443c-92c3-1f82758dfc20 · inbound
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0299e798-3028-48b6-9547-ec7c6453b70c · inbound
Scaling Native Multimodal Pre-Training From Scratch AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.