Pith. sign in

Paper Citation Record · LEDGER

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

As of 24 August 2026, this Paper Citation Record lists 100 of 287 outbound references and 100 inbound Pith citation observations for arXiv:2304.06364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.06364 v2

Coverage vector

measured 100 of 287 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:03:58.971585Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 108 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:30:12.075122Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 287 outbound references displayed

  • verified exact49
  • verified fuzzy47
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

61
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation eceddf27-e116-4dcd-b815-c38f60c9ab3e · outbound

This paper cites Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence

Reference 1

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.073690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:226f920afe9002cdec4ef25ffdcad9c7e98e591f44934af4411182c3f63e47b2

Observation 4827eff9-f0fb-4b93-82b5-8a54479a91d1 · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.500020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:845758ac93c060d7e854392a3845ea1e1bde2d0c3a477e55f6d6b5db0f66f5b1

Observation 49b18a50-e937-4c9f-ba12-bafaa6da526f · outbound

This paper cites 2023 , publisher =.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , publisher =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.502296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:11869e2a43424ed2c3f3c7721fb4a2ee7a0edd81a1e1419a5ebe589842309092

Observation 9cdd79b2-8143-443d-bd91-395de9f7aba4 · outbound

This paper cites Communications of the ACM , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Communications of the ACM , volume=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.504555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4f8db4803eab8172ac141e8ef62e45f850d1e36577652ca824b36930969e8e9c

Observation 2e3f82e7-0e7e-452b-ab13-456dd0af2239 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.507286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c2c9474dede8a22c54f711d2dc6dfdb8c24a08f6a0a89d79913b6644a0bce52c

Observation a5976102-e454-45f5-bef5-0209559807c9 · outbound

This paper cites Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.509285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:7041af5b2a6892e7c5be8fa29c0f33a2edbac2dae467178ca84faaef9dd8cb8a

Observation 79e3a7cd-34d1-41b2-9b64-cc396f12cdae · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Stoica, Ion and Xing, Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.511624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f5dfc38058116b8780192e49ebdf9a35218a48f387e5c6cf0fbc50956dc5140a

Observation 5535508e-df1d-4a78-a604-d68fa1493f2d · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.514004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:09154197ba22d6bfce9b8fa6d483b78a00e8f43deeed3be30e211be950ced5d7

Observation d06274b9-7630-4fc8-8fe0-4804566d947e · outbound

This paper cites an unresolved cited work.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:03:59.516110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:02bc7baea6ffbf7ac40559b235d6009c9f297fb5eb1b242c24e69694ce744f59

Observation be3ca354-5d71-494c-bc43-6a070579f636 · outbound

This paper cites Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.518118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e4b734e6c2c7e5fbe4aab97fc916b4469fe0efd25142b375a0c4cfe6a0da8c5f

Observation 10aab787-1efa-4875-a11b-5002787a5e8f · outbound

This paper cites Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.520446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:72594c4ff4a4bf3da27751f3c0cc27dfb41c9cfe8e699fb21b0253f55259f4ec

Observation 0d52f2b8-57f3-4bf8-b613-295730b44438 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2022 , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: NAACL 2022 , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.522415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:b1ad0f5d8176accc7846fcf62148c52feab766fbf33a6537820461a5acacfa0a

Observation 7c539422-9be9-407e-8675-e7109af50f3c · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.524430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:fb20ca9adb8e4a352f11b4ef888167659afb6de6364f1171facb8cb15e665f4c

Observation 12608aff-31cb-485d-9b4e-c874ca44d5f0 · outbound

This paper cites Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.526825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:63b8bea4348a8d9ea014264a76e52e94520866c1effacff695933f37139fc525

Observation cba04094-6373-4beb-9698-012f02899b02 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.529167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:505146505c3b502045ca31e6b4daccb3c263c43171ef85e8486c7ae9525c09bc

Observation 69ae1072-7690-436f-ab93-049159b4b0cc · outbound

This paper cites Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.531582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0437c51fbe4b0546209decfcde835f2786325aa3af5038ef4b773065420a64d4

Observation 462520c3-05d5-427b-8e64-d395cd247bc9 · outbound

This paper cites Sort , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sort , volume=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.533715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:b6cff8e7608227befacbe17b84df36ebb74fb0b827760403ea8510b0da31c010

Observation b720d16e-98a0-4801-a945-97598cce481d · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.535952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9bcf68bbc5f046679197dc3bef4e6f5fe49e55ce4077a2f34ee537e5901064af

Observation 5fb47e10-56e5-4c44-8597-50c40d8cc416 · outbound

This paper cites Proceedings of AAAI , year=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of AAAI , year=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.538532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0c8a635d2e885df824f91615cb2518c6da8f7830f437553718b1d8d228afc556

Observation 3053355e-a2ac-4da2-8062-f24b2f753da5 · outbound

This paper cites 2023 , eprint=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , eprint=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.540859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:bb50800f607fdd0d349f58e4c336f1a90291aef973b9799ea2a06045f9670de7

Observation 82aaeada-ef03-40fd-8be4-3c3468aacd3a · outbound

This paper cites 2019 , publisher=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2019 , publisher=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.543220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0fc8d47523eab2aa20ad169d42865dbf023f0f96ac7e3e30ab1c357af1dfcc7a

Observation 9a1b4580-83ea-440c-a13f-f139d49f0e0b · outbound

This paper cites Advances in neural information processing systems , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.545617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:65b1f58994b98c4d40967091bf9de0f3ca5dd80bf084720365ed8422f10f83be

Observation f4235ec8-8e2f-4aa7-a7b0-782de6e941bb · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.550600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c80b6e2ff6d55a8d52057c10eebc5aadec435bac6ea6288df1e8a40308245010

Observation 218d80ba-ae60-4f8a-8b1f-ac5ab851a87a · outbound

This paper cites Proceedings of the 44.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 44

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.553032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:76b9579532202bf938ba0afd78d810a3dc232eed043fcea3eca7fcdee4cb7c54

Observation 913d7210-379c-4fc0-9c98-6ad335545145 · outbound

This paper cites Advances in neural information processing systems , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.555986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:19592797db87961f6ab8a11b17f44a4b7abe4647f68a8409ab498d2a2381da61

Observation 7ec90262-7d23-4b5b-9803-b25d309f2201 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.558766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:6c019c214618102a77f88fb3cd3ceb01f6c1d7696d6959793c1851449d108b65

Observation 1598b0a2-3d28-47b5-baa4-6946bcaf02be · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing , year=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Conference on Empirical Methods in Natural Language Processing , year=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.561172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2328b19c6bc693a72716714911797b6bcbcaa483efbf121da2eeb940fa8d736a

Observation 5f51fdd2-55ab-46d8-a328-85303efbe9bb · outbound

This paper cites Available at SSRN , year=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Available at SSRN , year=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.563201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9b5706fca6d52bd39839e01d7be0b640e82b1744b262f6e1751acf87a6159a18

Observation b69fc3a0-1349-4b6b-81c7-acf67b5ccbf8 · outbound

This paper cites 2022 , eprint=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2022 , eprint=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.565636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5099233e2f5f86ff361d1f6a9bcaa633217ecb7a0b9ddbb74fd0aaf680aaaf38

Observation 5d3c9655-3c2c-4664-9efe-e4c9171ba505 · outbound

This paper cites Open llm leaderboard.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Open llm leaderboard

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.567594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5b97373a952ed562d455c4f6d205bbe114f8de33d40bb3722e72163756417d3c

Observation f6b526fd-9ac2-4c98-9334-45eab891f8cb · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 44

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.079975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1ded44efd66ff6bb100d74eb64f0a91eaabb08e970f534d8a80016a10a61f59e

Observation 691af668-411c-430d-862d-bd1c621472ab · outbound

This paper cites Language models are few-shot learners.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Language models are few-shot learners

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.745209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:cdb3c82b76b183dd5ac1a3b605ddab416f6d4d9ae8825fd9c2fc1b7431661092

Observation 993a77f4-6bee-4aa9-9627-254650120217 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.375746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1894ba5a0fd976ea0f67e4f2c6e8b5a21115eedc6b729cb0b9ebf2f06023f07d

Observation b48e4779-0801-4903-a479-770cb5d1cfa4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gonzalez, Ion Stoica, and Eric P

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.738102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5fa900ae16dfc45b57be93b0db0e8ab16daffa3e48125f153bd8d2b0c0524352

Observation f24f9ad8-ec22-4202-87ae-48da610e86b4 · outbound

This paper cites Chatgpt goes to law school.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chatgpt goes to law school

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.749863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:00fa6a92194bb138e3147c6cec4c3819e56c39da40b6556ab6daa7943449e03f

Observation 5f8f8b8b-7ecd-48e8-b82e-c32b3448bb82 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Scaling Instruction-Finetuned Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.379161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4e6a5bf815be0d902750a1f949bcc0a072e6ad70b73f6dbaa1272a05989eb3a0

Observation 1b303851-ad7f-403c-8d98-99f0d9f4fc6b · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.405898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e9e6b19261b54aa7736feffcda18ba801ea5550499ef3965366cc2e4bd90be1a

Observation bf54d8b5-bc7b-4a6f-9f4b-0fb30d85ed55 · outbound

This paper cites S ent E val: An evaluation toolkit for universal sentence representations.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models S ent E val: An evaluation toolkit for universal sentence representations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.732390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:75ef26bc8e16d52a2008df82e7bb2bbdfe1f6a8278a66b06d02ed61c7f2f0708

Observation 33bc5809-39b6-4c1d-96f3-70b2dfa37f58 · outbound

This paper cites BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 52

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.083891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8890536135282824fb31f64c5c1fdac2915cf606f061be1a18aebc765708442d

Observation a93b10d7-3fa4-4c49-9261-d467bef46773 · outbound

This paper cites Bold: Dataset and metrics for measuring biases in open-ended language generation.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bold: Dataset and metrics for measuring biases in open-ended language generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.752424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a06750ca33b81e0c4bf78d1dc6067f61986dd832d5b77804840806c2f1e74fab

Observation 2a877411-e333-49a6-b87c-17c3ec2397e9 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Glm: General language model pretraining with autoregressive blank infilling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.754871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0f9ea254174b4907379efe264cbff25a4b286267e7fa938c9f6d7291729ef288

Observation 74321b76-3ab6-4a94-aa01-2c563ab54fc0 · outbound

This paper cites Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents

Reference 55

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.087485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f7ae3c809e41b159e4648a98f265d7f71028f9346c81028e8f518cac3f1990cd

Observation 17fef00c-fb8a-49b0-8274-5f2caa630e97 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.401936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f1e06f4ef2a650c61a30f06362b89ccf62188fda0bf085b1b576710a98e99c3e

Observation 5ef4eacd-2452-4ba7-9353-6ff6c347e65f · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring mathematical problem solving with the math dataset

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.709419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2cd95a921b9137cf36331ea94d59d4dbcb6d827c2f565ed68d5cb7c259b5fc6f

Observation bf17a0e3-d25e-4d47-a60b-66fa52e7268b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring Massive Multitask Language Understanding

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.361824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:b4df7acced1f60f0750ab7274dfb4925b24215516497a350edc67b2f877f8497

Observation 11fd4998-b610-4a65-b24d-36b34abf0ca1 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 59

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.090511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5c2cdf5da02a41a421d4034d5c9f2c1ca12f157d6a1939592dd345f91f265a02

Observation 7436fb72-2b7d-4a81-9515-a5c7e6526867 · outbound

This paper cites Solving quantitative reasoning problems with language models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Solving quantitative reasoning problems with language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.718709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:693644b52170f964b8621ef1724fa71efdde0a6cee9b5e6353b7fb81e4e243b8

Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · outbound

This paper cites Holistic Evaluation of Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.393937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:96d4fb3037023f9a2e0db4fac907bcc11f97c26d4d062f552482563691c7a260

Observation 5afb741b-c594-49c9-b21f-5d1f3cd3f1de · outbound

This paper cites Program induction by rationale generation: Learning to solve and explain algebraic word problems.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Program induction by rationale generation: Learning to solve and explain algebraic word problems

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.716622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1b9c08c339f31eb980be43cada9145e8c0753c692957f7e44f8fb33ef3d97145

Observation 492a83d3-6c7b-4d4d-abe2-073dc79102e7 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.742853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c2e1e94d51ad8043fef22727416640f8251f0a4d5b82c3f4858fa73e601f5781

Observation 8cc1fb61-b6e1-44a5-af57-97a2d314fdca · outbound

This paper cites Rebooting AI: Building artificial intelligence we can trust.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Rebooting AI: Building artificial intelligence we can trust

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.706679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0cd5886f63b7429819d94d2748e4da7301057698ae0d6684c21d5a800c2974e2

Observation 74af0108-3cd6-40ce-b897-eaeba7dceb9f · outbound

This paper cites The Natural Language Decathlon: Multitask Learning as Question Answering.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The Natural Language Decathlon: Multitask Learning as Question Answering

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.370251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:06a7b094797ceddd04a4d924db2c698ad42c3ceea727dd072651e7f94b8b20a1

Observation 06e8bd09-ba43-41ea-a2c2-61486f56a500 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.747617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:49f85df2ca7cf2ce9bd16a0195e29ce0998142a4903ef292cf57c71cf322272d

Observation d1ac8ec6-ae06-4a8b-904c-f91c8b6cb2e4 · outbound

This paper cites Gpt-4 technical report.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gpt-4 technical report

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.740482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:25b958111ffd2068bb6e842cc3fd5a0086a4f4af163649137d5939214c1c1463

Observation 6151d5f0-9bb4-4d86-98e9-fc7ff67c8c8f · outbound

This paper cites Training language models to follow instructions with human feedback.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Training language models to follow instructions with human feedback

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.720716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:d7c9d724cf84997f14db1e0a96819225b76d098a57798fe0c037d06b951f4dfa

Observation 0b9a2912-9140-4006-8969-ab5871388564 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.397681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c99049f94ca331841fa2f3ea3a6f6bcf06faf23392304c94259a73a361f03e99

Observation ab8fe06f-5254-42fe-9281-58575fc4fade · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Squad: 100,000+ questions for machine comprehension of text

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.724507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5315a24ce53366f9e5e2580b988ec5b18fca9cd2458ecc910a8a4bb82ec550fc

Observation 18d8e589-4ddc-4a2d-b85d-e354f30e5b7e · outbound

This paper cites Internlm.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Internlm

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.735794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8604852ed4d465d1b9bc58d0a9dd99ed7da70ba8ec5ae811c25595ae196b9ee3

Observation d2e3e2d2-394f-4046-b05f-bd1c02611323 · outbound

This paper cites FEVER: a large-scale dataset for Fact Extraction and VERification.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models FEVER: a large-scale dataset for Fact Extraction and VERification

Reference 72

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.094484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:38ac095f2f23423a37f7db5b94e8e5b1eb53caeb9c55675e77e028b45810954b

Observation 9bf6514a-5cf1-4f39-86fb-05d3c8f08353 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.366393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:71c827a2f784bc562b89ff2cc9688fa0f3e7b7f61a49c43cf250645f096f0616

Observation 12435b9b-5f55-49a2-9727-4ae09c88d1a1 · outbound

This paper cites Proceedings of the 2018.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2018

Reference 74

Resolution
metadata mismatch
doi, observed 2026-05-16T10:03:59.098487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2ec3c2f1dfbff771316928076a55be51db2846ea44d923fe204fa75ef7c30ec3

Observation 3f086877-2f1e-44f2-91cd-ee86bbf2cd0e · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.757529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:63b5a00b1039ba5d304d7bf0855a97aa8d0b0329e973ad249a09535f8410b636

Observation a9d818a3-b884-4b38-bfef-3b7826380d7d · outbound

This paper cites From lsat: The progress and challenges of complex reasoning.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models From lsat: The progress and challenges of complex reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.727175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:fac193f8dec1455e6d96ae0a144090775fec526d5620b8e69fb9646328d4ecb0

Observation e7e0ed83-3f32-4c6c-b3f2-afe2dbd938ee · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.382955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:04695ef62ec164119a3ad8dc01343559b666aacc7c84b2b7eadbefa2fe945540

Observation bcf85aca-e9c9-490c-a91b-5ee4cc34586a · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models OPT: Open Pre-trained Transformer Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.386291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:fbc51f5902c48ff74b9c4eb0d0a1133b60ae4fbb9d7b0edcf4c18c7cedd65838

Observation 4c12a26a-adaf-41eb-9701-0e88483c06e6 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Automatic Chain of Thought Prompting in Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:39:17.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8c609abba26a4cf5b6c5068b84cfa088369b08959f3a2a9372951a8ddf67b80f

Observation 3a2a199e-9a60-4dc3-8da9-f61c1846653d · outbound

This paper cites Jec-qa: A legal-domain question answering dataset.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jec-qa: A legal-domain question answering dataset

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.729592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:93168a3d1db6cd282d3a6bd3b454bc4f066a54f12192b2b3e780843491d3f7c2

Observation bd69c749-c915-49a7-b05f-c5c109380e87 · outbound

This paper cites Analytical reasoning of text.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Analytical reasoning of text

Reference 81

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.101368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:06eaebece2ab52cbb3e10e9d1120bc86457ec091d40201014b37281b43a2422d

Observation 03fa877b-326b-4c22-ba46-2fae0045a960 · outbound

This paper cites an unresolved cited work.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:10:50.714389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:85092dd097433d7dcdba2fb2e60c59be5f960fdd0dee5049e59774b07b157fd6

Observation ab3b8154-9675-4642-a5bf-c6c1aa7d393b · outbound

This paper cites Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers

Reference 83

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.104165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:b95864d5beb09dd4092ac38a570650ce46030728f1632f05ebbf86d891cb07d0

Observation a4e67bf9-8441-4a5a-a2ec-0fa2fd704679 · outbound

This paper cites Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces

Reference 84

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.107151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1373b5069db97232ef3e47aed2db876d879a0529ec79fe4620360f0ede3e7fba

Observation 032ef19b-7e0b-4b5e-8f6a-c1fd2dc0406c · outbound

This paper cites H ate BERT : Retraining BERT for Abusive Language Detection in E nglish.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models H ate BERT : Retraining BERT for Abusive Language Detection in E nglish

Reference 85

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.110275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5645c4da45a6b07668ad8eb8b9c80d09b042a5013a78bb17a2348269aa2d1fde

Observation 6bbb9fed-d274-4707-9a09-a45bd33f599e · outbound

This paper cites Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset

Reference 86

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.113286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e05f3457d894cbaff8d7bdac39d95063ab5067c0314f757a1ae79b355eb3ad97

Observation d4658c11-1b9c-47d1-9f33-2e05499b2b5a · outbound

This paper cites Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation

Reference 87

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.115917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:723d2f9b49ddfdc87fb410b2fc5fb2f5716a66e4019f6ec36788540be3bf3f78

Observation 830f09bf-d889-4895-8a9c-7c83090d472f · outbound

This paper cites DALC : the D utch Abusive Language Corpus.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models DALC : the D utch Abusive Language Corpus

Reference 88

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.118550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1616cab2e43bfcfc38c4b60dfb9ba83c1519bbfdc2e90ecdbba38193bf37a7c1

Observation 90eceacf-ec56-47d7-af81-ce4bb74805ec · outbound

This paper cites and Dulal, Saurab and Koirala, Diwa.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Dulal, Saurab and Koirala, Diwa

Reference 89

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.120857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1215a63f6ab64ca0f100ee8d65f3e59393c7303019f0cda8e0102a3d8c0eca69

Observation f8d4ca18-85fa-477f-8b1e-992095f15976 · outbound

This paper cites MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms

Reference 90

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.123500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:710215922b8c1317ab68d4c75fee06ac6897ebdcbf4019d1200fd71498b837f4

Observation 87f1d437-2827-454e-aa09-01332a2b4b52 · outbound

This paper cites Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist

Reference 91

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.126081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:d89ccef431c3cebd74f5b26fceb63318a9a2e1b894f289128e40a380b5532a23

Observation a438ce5e-44e9-4f87-9313-281e8e6291ff · outbound

This paper cites Improving Counterfactual Generation for Fair Hate Speech Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Improving Counterfactual Generation for Fair Hate Speech Detection

Reference 92

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.129146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:425c732cac86a8bf1e7bd5b3f64d80f43ac8fe4a8a401034394e81f0babfd509

Observation 9a4de569-3b17-4519-9810-27d09fea88bf · outbound

This paper cites Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon

Reference 93

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.131534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:cb5447a12a4aea1c9720082f8b7858c3fd07ebfd0f5af979d4b52feb1fe393cd

Observation 96de1339-8b1d-4989-a1de-8005801b0fae · outbound

This paper cites Mitigating Biases in Toxic Language Detection through Invariant Rationalization.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Mitigating Biases in Toxic Language Detection through Invariant Rationalization

Reference 94

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.134410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4b839480329e03ea5818634f99df24d095fb7f7f2bc0d7a3c50f7a11308ce145

Observation c8ae52ef-7779-42a0-812d-a3f4844cfe7a · outbound

This paper cites Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments

Reference 95

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.136789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a8b193665cc497639e85ec8bfd808d22bbaceeb053316da8dcdbcea6a775c27c

Observation 274c958d-26ad-434c-9fe7-e55115ea91aa · outbound

This paper cites Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse

Reference 96

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.139242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:37669e3c7197ebae8e323d30710e061bdac1188e837b700018037aafc525bffc

Observation bb2f7b6f-5349-4428-9352-f12a84e9a4bb · outbound

This paper cites Context Sensitivity Estimation in Toxicity Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Context Sensitivity Estimation in Toxicity Detection

Reference 97

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.141516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:367ed9a046d46ebfad84fa69448057bbcaa693dda5766d03e0e37a9724a05964

Observation b9199c39-3889-4e5a-bb1c-95885278251a · outbound

This paper cites A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection

Reference 98

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.144558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:32c5e938db84b529c9661c3193a3cb7d80cc56d8b538714d11f44b875a09df2e

Observation 65c9bad9-4fa3-416c-bb00-4ecc618df2f4 · outbound

This paper cites Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format

Reference 99

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.147830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1ad9d173b12d6ddda61d99caba330f4dc1aa9c9fb01dda7e159bd72318844f48

Observation f5cc10fd-b4ab-415c-85b1-fee0ab1328fb · outbound

This paper cites and H \'e bert-Dufresne, Laurent and Roth, Allison M.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and H \'e bert-Dufresne, Laurent and Roth, Allison M

Reference 100

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.150099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:b9967e2a3fb536ab2f9d41d6bf736914e9e05c77a36a1a49068e42334fbc5969

Observation 223e06c0-2ce6-44c5-899f-22a22fd1efce · outbound

This paper cites Targets and Aspects in Social Media Hate Speech.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Targets and Aspects in Social Media Hate Speech

Reference 101

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.152182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:b8a538da7b9702a8bb319fecb924204191c794eff460f3aec2e3a281eaf57edf

Observation 68ce1835-4606-4bb5-b487-31f1d36ecf11 · outbound

This paper cites Abusive Language on Social Media Through the Legal Looking Glass.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Abusive Language on Social Media Through the Legal Looking Glass

Reference 102

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.154771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a03535fb5b2f839e8d494ce3fe9f61a422268f5cb6027cf9d79e233ee02d921b

Observation fec804b5-2b12-44b5-a068-c28e9f61b2fa · outbound

This paper cites Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection

Reference 103

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.156983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4733f4eb91cbf5e224317753cbb6e80b3af253a6039ba76ff11c63ed326fac2a

Observation b72c17f0-1090-4f73-be34-172e5f03780a · outbound

This paper cites VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes

Reference 104

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.159625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8bd917a5717235aebf13d7660735acb319c2f1c8e6f061f81dc3ad32f545f4bc

Observation a5ee5f4c-bd93-4234-aefa-7c4278cb278d · outbound

This paper cites Racist or Sexist Meme? Classifying Memes beyond Hateful.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Racist or Sexist Meme? Classifying Memes beyond Hateful

Reference 105

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.161674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f8e57960d3fa9f8a6231e3bdaa97c78e5f85331e8c688a3d85fec476a260f9d0

Observation 470c38d8-3f6f-48ca-a061-7f05031e48aa · outbound

This paper cites Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes

Reference 106

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.163828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8928151e456f7aef4ca521624b444ced2ed4925d5bf633d667b5f830836693e0

Observation e3570e05-a656-4a15-af58-58cb564b2b72 · outbound

This paper cites an unresolved cited work.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work

Reference 107

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:10:50.759798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:45625fdf9dbeedec44c79daf741dd0d59ce9805e5ec4bf12cff88d9e9fb196b4

Observation bca6162d-039b-4d9c-8041-29ee21286cbd · outbound

This paper cites Text Simplification for Comprehension-based Question-Answering.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Text Simplification for Comprehension-based Question-Answering

Reference 108

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.165782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c449e0befded5dcec398a2e30df53fb870d620b72589ebf20598a8b9cb5d09e9

Observation 18b935fb-76c2-44c0-a730-b029aef26efc · outbound

This paper cites Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets

Reference 109

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.167986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:cf423c39d17945ec2a7cbc18cb41877d1968c2c316573420414d9d87e65bf204

Observation 6cb98123-eb58-467c-9357-e1cb0cb67d68 · outbound

This paper cites Detecting Depression in T hai Blog Posts: a Dataset and a Baseline.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Detecting Depression in T hai Blog Posts: a Dataset and a Baseline

Reference 110

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.170033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:765aac79f6537eb477720df87971d3bedd34f570030f195de552db0e41995eee

Observation 26398b59-dc55-4a42-be27-20a95b9237b4 · outbound

This paper cites Keyphrase Extraction with Incomplete Annotated Training Data.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Keyphrase Extraction with Incomplete Annotated Training Data

Reference 111

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.172227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:95a7b7e0f9f22af09fc0a31a2236e8fcbbaf35ed8cd9a3fab63232b255397131

Observation 0d4bef67-af6d-4ce4-89ed-a86e3e0666c7 · outbound

This paper cites Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks

Reference 112

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.174264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:49b89baed3ba62cef5b4eb664c9cc8ffdbd74aa147ee0f3027b471b438a1bab5

Observation 5c5c900f-130c-4450-a8d0-0ee8507ac09b · outbound

This paper cites Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks

Reference 113

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.176580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:87f94f2159545f0740b6311a3bd9d560e2d4e09bad7325c80c2ef0542d6e50f1

Pith citing papers

Observation c1b92dad-805a-47d2-943c-dc989a98b5ff · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:33216bccafb86d64f3452fba90c91c855f9328d3d202dcd3de0edee4fa6ae22b

Observation 96043340-8524-4b05-94fc-ebaf9b2cb6df · inbound

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning cites this paper.

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:46:39.673387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T23:46:39.330438Z digest=sha256:ace30e288eeb2d092c457e934ce875f87a86e56a1b2a606581f1788e54b21c48

Observation 4ffa988c-4a1e-4c9a-9fca-c8d94dfff545 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.494391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:45f28350427c692d033fead1c22b054a7526ffdb6e4c7740728f6e5ccae25e52

Observation f5134d17-bbc2-42da-b71f-ee7cedb4bc28 · inbound

Mistral 7B cites this paper.

Mistral 7B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:14:00.055653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T06:11:38.406350Z digest=sha256:a70b5ec460a2d8a0733c3d6d99f2d90c7cd74c555d0148d04a9c4868286019da

Observation 9644ed94-4617-459d-9b84-777600b65e97 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.798370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:b1e5e08020a5410af7fce428efe776ac6c67c7f48b4b55fd1cd41bc3040f6db9

Observation 39ee5a74-f36c-49a6-928f-b280709feedf · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:142464d43096dc9c2850b809d20eed9feae1182935b466917b98d80039711372

Observation c718a91d-569f-407a-aaa2-36b2e00374d4 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 226

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:f9f7c50cd27433c6eab634b77e358704754059e5b159393a19db9e6a146c7bf3

Observation 188b5ec0-8e54-4362-9e77-31e427c34247 · inbound

GPT-4V(ision) is a Generalist Web Agent, if Grounded cites this paper.

GPT-4V(ision) is a Generalist Web Agent, if Grounded AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T19:40:13.816153Z digest=sha256:1857b97d38a15f7a0d9f7c78ede6de71ee456757110047ff95c79ba297eef70d

Observation 29c50967-2771-46e4-b630-edde26399d5b · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:20459040c948e10c7e850242143fb3b92e5e28f9618b08ac59bd9fd3ef808082

Observation 321e39c8-7012-4f44-9288-52993abeb961 · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:13:53.841009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:f67ee5ca4d868ac565adc8b74faea46ef4e044c4aca94be97c32911f444263b5

Observation 4283d2d9-6d40-47f3-9f11-39b38cf5d953 · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:23:49.423836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:ae513013f4806d4f8aaf8719d1cba00b333e60da9989f27a410468b546c9a8f9

Observation 2906ed29-c1a8-46f8-91f8-639cc21f41b8 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:a987b8adb102cc813da2fedbeb07b7e79852e8a3b6b670bbfd6fcb74fc1c0ba9

Observation 7403b0e0-3e8f-4a45-ada0-ab469dac8475 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:0ee3715fb0b2755dc2b84faeedec7e03667b24b897ca135b62e853986b0d3f4d

Observation 82f29955-ba9e-46ba-a460-84dd8ebe4ddd · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:6597c28592087728d57f86ad97f9aedfca67d80678004b0597f4eb23b5f908e4

Observation 9b4d48db-aefb-4c5d-8d47-bff912511c6e · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:8e2913027a387ef8bb3f7438ab2e38b5068913c3b2acddcb1a5f67888ea379d4

Observation 5a7a2327-af5a-4761-b4cc-95eb56ddd36f · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 220

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:58:17.383933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:46b24cf77b146ecbd909e5846e1ccf316bce3155a40dc5d8892d25df5b5e11fe

Observation 43ffd89b-4280-4515-87bd-f45c3388ec8a · inbound

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence cites this paper.

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T01:06:07.787696Z digest=sha256:8b1707048228595878bc46398273f9e044665cea11e1a69ded45f75e96483cf6

Observation 8f1175dd-89d8-48f6-8840-c3271dc946ba · inbound

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence cites this paper.

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T01:06:07.787696Z digest=sha256:c1421ae956c2422c1452e66738233789baeb89d0f871871037b2eaf13734fbba

Observation bc643947-5dc4-4ac4-b4fc-a6db1271f161 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.324452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:b34bb46b616a46947bbe6fcd2170fb4ada3506ad2442f3ea642d001d5905c1d2

Observation 3f745850-9722-42c9-9057-9f449b0d9d04 · inbound

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio cites this paper.

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:43:25.254645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T20:42:38.782232Z digest=sha256:51e37d541c86e829d20cb133e3a3e45dc54963b13c53792d3fc07708aa51a5c5

Observation b0186bcc-c37c-48bf-856a-0451e49dfe8d · inbound

Pixtral 12B cites this paper.

Pixtral 12B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T23:53:29.862702Z digest=sha256:497515b3e1cfac8f5efdb88f9568c6d199303bdfa3ed06646020bdb8e17395dd

Observation fc4a806a-8251-421b-b9e7-8bbff3c7ca6d · inbound

MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis cites this paper.

MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T18:49:29.309672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:49:29.309672Z digest=sha256:ed19c9eceb73209a0ee8b3921ed0ef5350186aa5d2d82b31be978e7d77dd77d0

Observation 647feef1-e139-46ad-a26a-5c623538ab62 · inbound

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs cites this paper.

VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T18:51:32.072557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:51:32.072557Z digest=sha256:f9b495947125d66189ff813b60b78e0f004f4418db95a0ed103cb4d42cb875c8

Observation f964a509-040d-4db2-8c1f-3416831f6264 · inbound

Ultra-Sparse Memory Network cites this paper.

Ultra-Sparse Memory Network AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:40:11.873164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:40:11.873164Z digest=sha256:b4320079338f6511aa3da7765db23dee242e22787a5cc681e6fe1ae7573d703f

Observation 3a6e5e31-3d7e-4320-9336-b964bec2382f · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 256

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:58.335131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:58.335131Z digest=sha256:5f57169d7d0cb1ed92d2d18064790cab86856ed2fd5a16a7df2425705a6bf0f3

Observation d22f2c5c-8282-44d1-80fa-3335a5a762e1 · inbound

AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity cites this paper.

AI Tailoring: Evaluating Influence of Image Features on Fashion Product Popularity AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:43.709102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:43.709102Z digest=sha256:0558439541a22002075c04fff84aacd43e65d6815deb4fafb161f4a908076dc4

Observation 933e5b7d-0a12-4a5c-856b-ba8c7385eca9 · inbound

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation cites this paper.

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T22:37:56.544734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:37:56.544734Z digest=sha256:7d4cf1c5643b8f402722096a2de410707aab956e82247f28aff48cea7666aa6f

Observation 99d56313-9574-4b1c-ba69-f46957e94b90 · inbound

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement cites this paper.

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T21:56:12.423708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:56:12.423708Z digest=sha256:9f058908f9bec2abc9fb7604d4d6bfc395d574439f0f962a49ba3181bf21a4e5

Observation ca57eb6a-6339-4196-ac25-3e59ea525085 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:52:37.581335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:52284a0f3e8f8c5ece3c97ccae14b759d15583cb45520c263c3516d37e122504

Observation 62b94ffd-131f-426f-a166-edd697b32cc1 · inbound

MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models cites this paper.

MoE$^2$: Optimizing Collaborative Inference for Edge Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:14.893961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:14.893961Z digest=sha256:81b61474ede8fdbbdf3720f310164c89f1256932f9a210473b229e0cd7361315

Observation 7431a7ec-604a-4bc7-bd47-a25fc223e0b3 · inbound

A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks cites this paper.

A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T19:27:23.551756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:27:23.551756Z digest=sha256:ccb721f99d57243d1c9c7754084d248164f140e9b491f692083b923a8fdcb3df

Observation d6a8eaea-9549-4d9b-a2f0-5d5f2227275f · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:3255570b6aa6c7f440602c0807a2fc3c22ce9ab20dff42473b03aed47104548f

Observation 5f9bd9c4-c29c-48c5-8b43-79ba6ed27c83 · inbound

SedarEval: Automated Evaluation using Self-Adaptive Rubrics cites this paper.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.012661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.012661Z digest=sha256:43379aaabf57dc5f9287f3073c37186eb8f19bae83e387004eccb80a39fd978e

Observation b974884a-c6ad-4b33-8195-5f424185ed2b · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.919382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.919382Z digest=sha256:18ae07de54ab7c2a70c71dd2261dca4afaa26f47ca597cc71865a7020edb61b3

Observation daa1fa42-6884-49d6-8aee-e3ace2c9d895 · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:07.043089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:07.043089Z digest=sha256:3816dc9441f87188e3ac0d0d400c74da05ccab8950a49e28bda0fb9b8e9acb9e

Observation 04cf8462-dd02-4483-b395-4a7ed7672cdc · inbound

Minerva: A Programmable Memory Test Benchmark for Language Models cites this paper.

Minerva: A Programmable Memory Test Benchmark for Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.839991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.839991Z digest=sha256:a8625655a6f07fcc0ed7b063febe03c1c139824935c0a5246bce5e475b3dc7cd

Observation d28b1695-489b-4888-b3ea-8132217b6286 · inbound

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis cites this paper.

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T00:35:28.557246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:35:28.557246Z digest=sha256:318bbac4704d86d76c87d9c6ab9244ba1c463a555b395396435d010f194e1c8b

Observation 0f86f630-1d5a-440f-a596-014d8a74a2cc · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.432963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.432963Z digest=sha256:e25c0087a1d29db573552d0035d6dd39cb65c5f4a3ff1273b8ed4fa494980e22

Observation a5d8497d-01d9-4a05-9ebb-a9209b335a59 · inbound

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring cites this paper.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:56.123451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:56.123451Z digest=sha256:8ebeb25d9312aefe52099d51f68821c27f28a07b0c0691e7dff230cdd5e29d7f

Observation 6d83edae-8300-418e-804e-c3c73be701fb · inbound

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models cites this paper.

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:03:44.710819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:03:44.710819Z digest=sha256:e032ffcc1d68f53d2411514e37eda6ba6f4fdc88b32adff4175eb236cbb7fd77

Observation 2c8ebcec-b078-4056-813b-4508d264d1df · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 135

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.537102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:f5e40c1eca5d52783986d4ab0700917e81bcd30068ff47eddb62488d26fb7e05

Observation 16882dc6-a0b9-46f2-89f8-3a2bbe4acf2b · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 25

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T00:19:22.173928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:153b60423db28fd6b4c9631a90c715dd1dc9ef47c4ce73f6eb25e7885e2e5e1b

Observation c1b08166-55b0-4707-a035-6f4afef9ae04 · inbound

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark cites this paper.

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:30:12.075122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:30:12.075122Z digest=sha256:12c047aa18db06ea2699ca1dc9a129cebb05fbba405d519f4378e2877e1a3678

Observation 509a881e-5b88-4ba7-a0a5-e4b37d78091d · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:36:58.695985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:b6befd0d6110e35d880896261c7784ed52a225d77c76bec578303c205bdf6239

Observation 72835dbc-ee36-4493-b095-d9e9e6bb973f · inbound

Computational Reasoning of Large Language Models cites this paper.

Computational Reasoning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:24:24.022914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:24:24.022914Z digest=sha256:19ebb83cc840467631ee54b7f1f96187016558747b914a89a5a419d51d23054f

Observation f1ebe5d4-d1d6-4bed-81cb-46131d2f3664 · inbound

LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models cites this paper.

LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T04:37:48.384477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:37:48.384477Z digest=sha256:0e60de9ec9a36c6f65d5eb20466fd9f842ed1b5f3f83e8347d4cbd8979f7c0f2

Observation 37fdae11-9ece-4c3e-95b9-a05642bae739 · inbound

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection cites this paper.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.616305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.616305Z digest=sha256:43f8c99299ebcd5b30b58a12e6c723413c47b66e47d50287cbccb8be34836967

Observation c043601c-eeb7-43f9-8df0-5a99280050f5 · inbound

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches cites this paper.

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:58.017216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:58.017216Z digest=sha256:139a1df6525bb1f125eba56a007e6e18ab0a5b1b84eecec48ef09f666c02cad2

Observation 7b1e8266-5f36-4f91-a04c-7406a32ca961 · inbound

Learnware of Language Models: Specialized Small Language Models Can Do Big cites this paper.

Learnware of Language Models: Specialized Small Language Models Can Do Big AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:55.893235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:55.893235Z digest=sha256:eec00bdef5f739367d5e23869ee34fa6e77ac2ba60f8f00a288efdc74a0874c1

Observation ba8898d5-c2c4-4162-b806-bda2c55114b5 · inbound

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs cites this paper.

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:42.297149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:42.297149Z digest=sha256:baa5141900fb655c18c927f093053a89b4c614bc46f69e349e3e94c1693d69aa

Observation d561a2b8-71af-4af8-9935-20e64dd06527 · inbound

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis cites this paper.

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:15.511809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:15.511809Z digest=sha256:a011ca68d1107ec5fd4a5acfb86c6304d1b0cb5cd91ae600bae3dc269fa14400

Observation c6e092f6-a6c0-4566-b94f-c61256041fe3 · inbound

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese cites this paper.

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 65

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:08:42.832179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:42.832179Z digest=sha256:1fdc8aecff9728227f633f5fa3e3b70b3f56e96fb4aeeffb1a7b3aeebb47469a

Observation aad62654-f98b-45a0-bf83-193d3f8d7827 · inbound

Scalable Complexity Control Facilitates Reasoning Ability of LLMs cites this paper.

Scalable Complexity Control Facilitates Reasoning Ability of LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:16.366510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:16.366510Z digest=sha256:fbf917db0dcc717a1dd34498e077f7412b41bfb1be0f562014b84b8f4b16a272

Observation cdf12970-07b3-43f0-bc47-85ff7762fa00 · inbound

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models cites this paper.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.129251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.129251Z digest=sha256:53026862c825b087df7fd7e83c5bbe5c97c3a51b49da6a37c2e337aa40925aed

Observation 3a51f984-734b-48f9-8dd8-7f7e50eb66dd · inbound

MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation cites this paper.

MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:38:49.905192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:38:49.905192Z digest=sha256:21a10375e791729ef7e6b25d4d83e28a9b9f775135048536447f41b5bf8a90a1

Observation 14fa1b8e-c9e3-4b9e-bee5-e79f7b5912ee · inbound

PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs cites this paper.

PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:10.040307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:10.040307Z digest=sha256:6435690602b0b3960a7dde988302f771c742e584d68582e67465e36503136715

Observation 075b5f16-5f97-4571-9558-5dd63cd38c43 · inbound

dots.llm1 Technical Report cites this paper.

dots.llm1 Technical Report AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:10.836071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:10.836071Z digest=sha256:d7c0a2e337a3ffa88e46065929fa75b07e49ec8ca81fb3870fc0901c5f7dde46

Observation 5e48c119-0629-4f5a-8bfa-b655e45f50ab · inbound

Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models cites this paper.

Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:09:18.900758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:09:18.900758Z digest=sha256:214f65cfb91196bc7311390fe30e61bee70854715dd1e4fdc17da786542bbf5e

Observation 79826531-a36f-4f8f-b079-f7eb5785f4f9 · inbound

Towards Efficient and Effective Alignment of Large Language Models cites this paper.

Towards Efficient and Effective Alignment of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 212

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:43.200783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:43.200783Z digest=sha256:7f020b3e86568f74e80512c9f0042455d04bdc7d0e98d781b71792bf8343e558

Observation 73a841bc-ceb5-48fe-aa7d-35c89d7577a5 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:05:47.607647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:cba00cc52a7f713bf03b3742abda6c4f4751c3e000e58e4bfe56f8d96d5f9683

Observation 73bcba91-64fb-4d43-a1be-866f254414ef · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:19.463215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:19.463215Z digest=sha256:10a4dab39362778a2bef583c740919ff6497223c04fcabb46f2b320b127a57ec

Observation aadeb1d5-8c85-45ca-b4f5-08ad8ebee611 · inbound

EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration cites this paper.

EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:09:04.281671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:09:04.281671Z digest=sha256:520af559a4ffc1ad7156931bbbba7187bf3a0313410f32b0521cd011c37dddb2

Observation dcb23d56-9a47-4491-989f-80b31d02cb38 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:32.689119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:32.689119Z digest=sha256:4b21b0d0d60573d50de7ea670ba74cf2ccd56a6085cfbd0e82bb3281a0b14bdc

Observation 50807dba-7958-4801-9fed-e72e7dd5388d · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:43.160042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:43.160042Z digest=sha256:6f78c520e0432d478b6e657e4647a4fd740e5d18dc2b43f9fa91d1c0d7da7be1

Observation 26e3890a-8731-4577-bf06-166dddb40dd1 · inbound

BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining cites this paper.

BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:29:45.412584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:29:45.412584Z digest=sha256:fbaa9648a4a8004d68a76b86237ea853650e7b436c0eab182ad5f4027a487c61

Observation 38d67280-fd08-438e-b4fd-f0f2b2311890 · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 200

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:01:10.053215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:5753905413f0bdc059fe7548154fb239bad284f57f09e96ffbba8720d80ebdb2

Observation 7aba488e-f72e-4576-8720-e8087df009fc · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.756978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.756978Z digest=sha256:17ec829279a316fd72eef75ac1d9497ced3b50e64106c2198d59fb30ffbb9733

Observation d4642143-8cdc-4e24-b437-e5b07ca4e934 · inbound

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong cites this paper.

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:18.704950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:18.704950Z digest=sha256:fa38cb63d96f8c181a740906d2fe5ce2ead8fad5a26e7df53bbb1ef277822cf5

Observation 9db48035-c307-4261-a317-ce17b6f69174 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:16.623970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:16.623970Z digest=sha256:432722c1f7b90001b7048da101abf32286f56ff710080761257d61a209cb24a5

Observation 601dc364-a39a-4d5d-bd0c-df744309b6d7 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:08.633367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:08.633367Z digest=sha256:7c3303ce0fd60d28ecc9f2a239057aa9d9001de9fb4c85dada5278830abf0f0b

Observation 336f7a07-a1a9-4885-9ce8-204c9c234a51 · inbound

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) cites this paper.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:49.735291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:49.735291Z digest=sha256:01167f7732227bbdfe5500328fa7a823cf6fab974625ccd49aef51abf222e443

Observation ad27a3f7-0084-4515-b40b-a8492c15271f · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:746b3706bd935a63def4c17d38a1c5e035a6f398be3ed9ff2d461e032c140135

Observation 2a5e72e7-5d37-4721-9dd3-7f6e9217b879 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.336650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:47c5965750f0aabadbec56a7213e9f56d1f86d93776674c4c125bf6da532e200

Observation ec54b7b3-8a57-4a66-9565-59282b8f0699 · inbound

Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts cites this paper.

Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:12.954012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:12.954012Z digest=sha256:664c9186c84888e17e3d78517336888101c3bfbecd7b5fd8c9f3772b456bdea1

Observation 86b9a7dc-8f49-4687-b24a-d7210756fdae · inbound

Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models cites this paper.

Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:47.930440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:30:47.930440Z digest=sha256:966ad58d2198bfd25057bb0ec2cb26556ce761aada32c3c84049e404d257effb

Observation 6dd2f60a-b4a4-4803-8b55-5ed589bf5855 · inbound

ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models cites this paper.

ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:10.358136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:30:10.358136Z digest=sha256:e2e7d7d76853a2ec20d4a7b3a96b58ec0ea17bd4528ab5e428903eeb15168c6d

Observation 54363329-7bcf-438b-bfc4-dce5d3080fcb · inbound

Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models cites this paper.

Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T19:30:46.125327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:30:46.125327Z digest=sha256:75c06522d3eaa8b57ef18d87e118d7f39ede9a25569f99f4f816f7689a9a3465

Observation 7828047d-0797-4a6d-9b9a-37881801c231 · inbound

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation cites this paper.

Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-15T17:21:06.866133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:21:06.866133Z digest=sha256:ea8157c93d0f6022cfdd9073d60ba04714639856d64859abd17f547ecea5bf08

Observation 5dde716a-4b75-4d50-b822-de3b1a0e9cca · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.885288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.885288Z digest=sha256:f8d6e805f3c90966f1cf894c657358ee874523c0e6f3e8aea0dc99fadb010dcd

Observation 19697ec8-1d39-4cfa-9c0f-2a9f8313185c · inbound

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding cites this paper.

How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:36.191653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:36.191653Z digest=sha256:9288b4e96e8c300131ca87589208d82751226d60d9e311c249a28015d4fe1ba8

Observation 060cab31-313c-40c3-9f4c-a3b8396c4f00 · inbound

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools cites this paper.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.740219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.740219Z digest=sha256:45bacf95eafcb5670e2f0d82d3f76da55ab6a953d22f23391359896a01efb9b8

Observation 1ed317f5-0155-49f9-b430-d44a63db8f32 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:00:41.543036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:79dfe7d7cc1e5be13cfa93fdc0abbddbf48b843da5d2f0a690cd3b38692968db

Observation 5883d43e-d7bb-472c-a274-a212bbdc59ff · inbound

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models cites this paper.

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T21:45:40.615196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T21:44:36.351517Z digest=sha256:82b537d0161f69948dea50bc102eccf2d508920bf98694a7de7d57a2d9ada31a

Observation 88f763f2-ef0f-4d89-af4b-a892cf39b52a · inbound

Dr.LLM: Dynamic Layer Routing in LLMs cites this paper.

Dr.LLM: Dynamic Layer Routing in LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:04:20.318861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T20:01:48.806709Z digest=sha256:7dfc99d7ecd3b1fe0d141f61e189c150b771a6b2e9c8f5b74fa912f72378c5ed

Observation d083d6e9-23c5-40ac-9d7e-5de221cd30fd · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T11:01:17.080772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:8e416f0eb06e7a0c6806ce7234d4a81e8f088d74581052d03546370cf0face13

Observation ba35216e-37b0-419b-8430-b27ba4c611e8 · inbound

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models cites this paper.

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:38:34.178169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T21:36:24.376401Z digest=sha256:3193b5582e2fa16768f642e94b9e322096908798ed20afe266afcee489a24782

Observation a77233d4-f2aa-48e9-b4e8-4bad0640a905 · inbound

Ministral 3 cites this paper.

Ministral 3 AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:6ef211690edcddf66cfa11c71f1b64dc4fff8fa958df5101b1238b379e49018e

Observation 3a0c1391-dc27-491e-a871-4bc359d5d02f · inbound

Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study cites this paper.

Bilingual Bias in Large Language Models: A Taiwan Sovereignty Benchmark Study AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T03:59:21.087147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:59:21.087147Z digest=sha256:39084c814690b84292065f2e79404ff9b8374925402d0aade0dd7f29efef0e88

Observation d0910937-4558-4948-90de-e21ad3005511 · inbound

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy cites this paper.

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:06.954126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:06.954126Z digest=sha256:74a1f1ac15f4afbe5897098e6341b6fe991948199b5e08037e700e075bd85ba3

Observation fea78a68-21b2-43a5-a1fe-e961d5eda5a1 · inbound

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment cites this paper.

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:36:44.401045Z digest=sha256:303b97571983fcf53035019722212944d6434a189916038d1fa560ed62a70698

Observation 9ca77fdd-8654-4f51-868d-f0caa031e0fe · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:7fd7f67f6085d2dd942e967ded532db3d6520908f2d5a1bb9b8e389ba7409851

Observation 4dac4e0d-3f34-4c37-95a9-13fc45d49019 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:68b2158a2da2cb71e118fdc211be4aee1304400362e8690c4d17090f64a2e1a8

Observation 58ab0514-50af-41b4-88be-d962e572afc3 · inbound

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation cites this paper.

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T18:30:03.367870Z digest=sha256:0b33c4e16ec35188b2f99551d5c6cd642da339f72599047a2fab04b7ac95088a

Observation 3d130690-2278-4685-82da-9f183a0feb22 · inbound

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability cites this paper.

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T01:12:44.476894Z digest=sha256:5dd5aba3a4a1c3c52232def6935ba8124fbddf764951ea6b881ee3c1684e6d19

Observation 572da502-e57c-4723-8c76-364e78efbe62 · inbound

Confidence Calibration in Large Language Models cites this paper.

Confidence Calibration in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:2d6a2168c1ef12bad210a0097a8c99ae003f82a6110161815321863035e5e250

Observation 31646264-be84-400e-a269-aefcba1406d4 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:23.517497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:1e79744afc7df2b077ec25adb716aa2e995fd676880006f59023684be7e3ceb5

Observation 9176bcae-832f-49ed-944d-319fa2b21f7f · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:24.460278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:24.460278Z digest=sha256:851cc4cb3be5975f7f37c014577b11743145ec6324f8ce9344266fb5d059aa71

Observation a23a1af6-d019-4ad3-a5ed-41e18f5a409f · inbound

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers cites this paper.

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:13:30.266780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T14:07:26.366172Z digest=sha256:215341dbf9357066be761853faff45e667c2bade4b5bf2f3743fd6355ad84751

Observation cc2f04a1-ac84-44b1-a783-a66334d63b63 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 245

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.972501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:d0b737af9eee8f9bd377b83e6f5a480db8820c43c9d808f4018b645193118582

Observation 1aeeb5e2-0e18-4e45-92f0-ee6cf38d01b4 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 247

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.407055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.407055Z digest=sha256:46511f1cab0df57307e05616b6de8fe3c969cdade1c903365e8bd401ccf43d13