Pith. sign in

Paper Citation Record · LEDGER

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

As of 2 August 2026, this Paper Citation Record lists 100 of 287 outbound references and 53 inbound Pith citation observations for arXiv:2304.06364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.06364 v2

Coverage vector

measured 100 of 287 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:03:58.971585Z

measured 153 of 153 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T12:41:24.460278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

100 of 287 outbound references displayed

  • verified exact49
  • verified fuzzy47
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eceddf27-e116-4dcd-b815-c38f60c9ab3e · outbound

This paper cites Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence

Reference 1

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.073690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e4e87b3809282db7f2a8aa12b64e9c70c2f026c23fa09a8fd68dc8dff418d25b

Observation 4827eff9-f0fb-4b93-82b5-8a54479a91d1 · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.500020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9afd81d6438198ae94898fb032fc1928afc5da5b19fa13db2f50727204197e77

Observation 49b18a50-e937-4c9f-ba12-bafaa6da526f · outbound

This paper cites 2023 , publisher =.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , publisher =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.502296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e851a3dc285b434c36b0e7544c2465e32e87b9531cf8e19cac3d7426a2133568

Observation 9cdd79b2-8143-443d-bd91-395de9f7aba4 · outbound

This paper cites Communications of the ACM , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Communications of the ACM , volume=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.504555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f0ac0ffe70ac7cba4ae3f3bd3be6d07d4511b46209996d3cee7455750b695a2f

Observation 2e3f82e7-0e7e-452b-ab13-456dd0af2239 · outbound

This paper cites Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.507286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:bb1a3733bb94bccf59697bd54aa037723e4654366afcdf066d01c4116e6fbd6d

Observation a5976102-e454-45f5-bef5-0209559807c9 · outbound

This paper cites Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.509285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a4c7c6ccd3fc9dd9de91c73278a56c4c472ee137198d42368c7e8a9f0251f8dc

Observation 79e3a7cd-34d1-41b2-9b64-cc396f12cdae · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Stoica, Ion and Xing, Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.511624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2a91d71423468b5e83e7d409261856460629fd931adcc7e005a9e647015d1f73

Observation 5535508e-df1d-4a78-a604-d68fa1493f2d · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.514004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:27b583767e1149f1528a36c4b9e177f21f8bd237c2b164f82ab1ef805d9ab405

Observation d06274b9-7630-4fc8-8fe0-4804566d947e · outbound

This paper cites an unresolved cited work.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:03:59.516110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a391aab93e329fafba6623985955047a473d71fb200f67fc4bdaa91fedf48430

Observation be3ca354-5d71-494c-bc43-6a070579f636 · outbound

This paper cites Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.518118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0b0d4f18d1ae4d729eda836aab2a00d3167f4bafab0d7a0ce7b638be933571b6

Observation 10aab787-1efa-4875-a11b-5002787a5e8f · outbound

This paper cites Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.520446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:dbe4fccd15a1903c30aecec5392780ce76b7872c88a282e77b1a3398452bc1b4

Observation 0d52f2b8-57f3-4bf8-b613-295730b44438 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2022 , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: NAACL 2022 , pages=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.522415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9ca3a25e5caad07dcd2947a00a8ef9106cfd5c9624cbb5bdda40fcc7a5d9ae3a

Observation 7c539422-9be9-407e-8675-e7109af50f3c · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.524430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0b22b5750702ec692f7d0c1f65ff83bb89a2821c3a577b994742f294c51c54c7

Observation 12608aff-31cb-485d-9b4e-c874ca44d5f0 · outbound

This paper cites Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Natural Language Processing and Chinese Computing: 8th CCF International Conference, NLPCC 2019, Dunhuang, China, October 9--14, 2019, Proceedings, Part I , pages=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.526825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:df54cfc1a5543e181b49359fdae5414c5c009a6bab9c1fb27f8a99617c8cc958

Observation cba04094-6373-4beb-9698-012f02899b02 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.529167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0cc4339bbaf2302d5e0e516b519ea3a2a7f0dedffeeb5f367703e9ce1a364a03

Observation 69ae1072-7690-436f-ab93-049159b4b0cc · outbound

This paper cites Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.531582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c41e4a1a9dda443e92eadc2ae5ac9598d63101d0a6a5efe403ffa3243e786ed0

Observation 462520c3-05d5-427b-8e64-d395cd247bc9 · outbound

This paper cites Sort , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sort , volume=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.533715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:7216345b37c9c2691f52f640d54910534089016daa18ff0c2ea417e3aeafd1e5

Observation b720d16e-98a0-4801-a945-97598cce481d · outbound

This paper cites Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.535952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:846a7420bfc6a7242b4505b240bdd68ecf8dcd49925680768dec22b2dab3bdb5

Observation 5fb47e10-56e5-4c44-8597-50c40d8cc416 · outbound

This paper cites Proceedings of AAAI , year=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of AAAI , year=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.538532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e26c21d8d738cf8664efbd2cc3923ebf1192ff1a926586cfb5f9a1bc7682f0d8

Observation 3053355e-a2ac-4da2-8062-f24b2f753da5 · outbound

This paper cites 2023 , eprint=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2023 , eprint=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.540859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:402b37fb9dca0d1a16c22fccdde264b04647dc1b1b92f2aa63d3d45201bca3f9

Observation 82aaeada-ef03-40fd-8be4-3c3468aacd3a · outbound

This paper cites 2019 , publisher=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2019 , publisher=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.543220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:da618f58cf6b00fa91164bce1d42d48ef102a85b32159f03fb6a5f72823974fc

Observation 9a1b4580-83ea-440c-a13f-f139d49f0e0b · outbound

This paper cites Advances in neural information processing systems , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.545617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:bca4146ab3ceb7ef22d71cffa0d9cce6ece052ebd2f4500caabd3ac415a9673c

Observation f4235ec8-8e2f-4aa7-a7b0-782de6e941bb · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.550600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:fcee4a047ba4d4246a69fa5bce611c153e79ed0fc25e7524eef903f2598f7d7f

Observation 218d80ba-ae60-4f8a-8b1f-ac5ab851a87a · outbound

This paper cites Proceedings of the 44.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 44

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.553032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:7ec5d1e02a00f0e6321693b8080534bfe2d1d2abd38e5dfa71b7f9b9846387f0

Observation 913d7210-379c-4fc0-9c98-6ad335545145 · outbound

This paper cites Advances in neural information processing systems , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in neural information processing systems , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.555986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c281d999562eea3a427aa0e8946b717e34451ae51dcbe361da2ecd00495104ce

Observation 7ec90262-7d23-4b5b-9803-b25d309f2201 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.558766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:aa3b06d1d9742e90a6dea6f6644c5d0d022d9d84df1dec593884516677551a34

Observation 1598b0a2-3d28-47b5-baa4-6946bcaf02be · outbound

This paper cites Conference on Empirical Methods in Natural Language Processing , year=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Conference on Empirical Methods in Natural Language Processing , year=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.561172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a7281f5d5bdcd7c180dd455e335ec5ba5bfd2ad9e8002f6b97580561170a76c2

Observation 5f51fdd2-55ab-46d8-a328-85303efbe9bb · outbound

This paper cites Available at SSRN , year=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Available at SSRN , year=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.563201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0158285b5ed319d42ecddaf23f4c6d4d31099b0fbea5b5a17c6a7da52b5ddc04

Observation b69fc3a0-1349-4b6b-81c7-acf67b5ccbf8 · outbound

This paper cites 2022 , eprint=.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models 2022 , eprint=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.565636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:43a41b31256eea652d1a24accc937b3fe3f696acefa0c8b58542d474be29f7b4

Observation 5d3c9655-3c2c-4664-9efe-e4c9171ba505 · outbound

This paper cites Open llm leaderboard.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Open llm leaderboard

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:03:59.567594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2b5cc494979a02e6b79b7b16497110ba2276e5bb5d41ac0e6fbddf4dcc962083

Observation f6b526fd-9ac2-4c98-9334-45eab891f8cb · outbound

This paper cites Bowman, Gabor Angeli, Christopher Potts, and Christopher D.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bowman, Gabor Angeli, Christopher Potts, and Christopher D

Reference 44

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.079975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9635548a4ae45a5ff3f466d7d3902e12c9a122c403ff7f549eac3a0348982579

Observation 691af668-411c-430d-862d-bd1c621472ab · outbound

This paper cites Language models are few-shot learners.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Language models are few-shot learners

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.745209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c8176e654717e46dd5126d1276f8aef698236180dd9f59a41746452392f5a254

Observation 993a77f4-6bee-4aa9-9627-254650120217 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.375746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:dd5a5f2703690fb7eb2a6a01dac7ab0555ed8677ca299ae1f24e526fd02aca8b

Observation b48e4779-0801-4903-a479-770cb5d1cfa4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gonzalez, Ion Stoica, and Eric P

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.738102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2a07623a0c6d8f22743117bc09e986f1cc8b83415aa46c88ff6c21aa277782a8

Observation f24f9ad8-ec22-4202-87ae-48da610e86b4 · outbound

This paper cites Chatgpt goes to law school.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chatgpt goes to law school

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.749863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:7c6b2c1ed1e7883460886e578263d67da60941e4f17fdbf360f7f394110ee172

Observation 5f8f8b8b-7ecd-48e8-b82e-c32b3448bb82 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Scaling Instruction-Finetuned Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.379161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e7b35ca45e49f3f6f600e2ce77684c9ce6acc99d9ad0c032a6b00fb18a537708

Observation 1b303851-ad7f-403c-8d98-99f0d9f4fc6b · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.405898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5ba1b571ad8b35ea42fc735784bab2d950ec3f9c98db3f2b922db6d7df772892

Observation bf54d8b5-bc7b-4a6f-9f4b-0fb30d85ed55 · outbound

This paper cites S ent E val: An evaluation toolkit for universal sentence representations.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models S ent E val: An evaluation toolkit for universal sentence representations

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.732390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:d0cb9bc62b28c24b4c7fd76055f7fb89d6759792eec0e1765320d357c96ead4d

Observation 33bc5809-39b6-4c1d-96f3-70b2dfa37f58 · outbound

This paper cites BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 52

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.083891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:691a9474c4c1c4a34303b64c67dac595d04526d8e86b4dcc3803ca999874212a

Observation a93b10d7-3fa4-4c49-9261-d467bef46773 · outbound

This paper cites Bold: Dataset and metrics for measuring biases in open-ended language generation.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Bold: Dataset and metrics for measuring biases in open-ended language generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.752424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:7a81fea99e1f587fc48092112fb5a6f7219c827c0c78b00fe6ee5a3130b6aff2

Observation 2a877411-e333-49a6-b87c-17c3ec2397e9 · outbound

This paper cites Glm: General language model pretraining with autoregressive blank infilling.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Glm: General language model pretraining with autoregressive blank infilling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.754871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:0e4203a2f38c2ad2e18eb7e9c66fed77174ba325f7dd613874a22a3ada1d02ad

Observation 74321b76-3ab6-4a94-aa01-2c563ab54fc0 · outbound

This paper cites Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Introducing the HIPE 2022 Shared Task:Named Entity Recognition and Linking in Multilingual Historical Documents

Reference 55

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.087485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9a9692d517e1e4bd527a8c36b0ffa72d05b94796d7faaa01dca113ca3999731c

Observation 17fef00c-fb8a-49b0-8274-5f2caa630e97 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.401936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a0d39b2e2e78a48a17ce32b3bb92cccdfd062203c2e1994f38f695edbe8ec52e

Observation 5ef4eacd-2452-4ba7-9353-6ff6c347e65f · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring mathematical problem solving with the math dataset

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.709419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a38b7eca54cc0b1720b7e3c35b249f6f0ee0446fd6a84300df40a806901ed718

Observation bf17a0e3-d25e-4d47-a60b-66fa52e7268b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring Massive Multitask Language Understanding

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.361824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:25661cea6ea0937d41c298987a9b73e6ef6fba2123a8a3e531d691ce2720aceb

Observation 11fd4998-b610-4a65-b24d-36b34abf0ca1 · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 59

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.090511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:6036e37c4103359f19d77d5c56c65eb9b943c6ae7ff610a1459ed94c24b804a1

Observation 7436fb72-2b7d-4a81-9515-a5c7e6526867 · outbound

This paper cites Solving quantitative reasoning problems with language models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Solving quantitative reasoning problems with language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.718709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:359a3106bdac1ad8aaeb43a7526dd690b0c3c3f07109049511e3d0a8d3d4079d

Observation 71fd7756-0d77-4e79-98c8-079e786055e0 · outbound

This paper cites Holistic Evaluation of Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.393937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4b79df65b611df4363136257faf68b70387a1a4731fd6d360c7bbb0dd350a90a

Observation 5afb741b-c594-49c9-b21f-5d1f3cd3f1de · outbound

This paper cites Program induction by rationale generation: Learning to solve and explain algebraic word problems.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Program induction by rationale generation: Learning to solve and explain algebraic word problems

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.716622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:3718636095d42d4f50b92555f5632e8d524ca09f21976b14e76ac57250ad32fa

Observation 492a83d3-6c7b-4d4d-abe2-073dc79102e7 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.742853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:35ae9a5a7b36f23bf92dc6e154ba6664bd28a478650c7eeeb1a2e879a9e9b2ef

Observation 8cc1fb61-b6e1-44a5-af57-97a2d314fdca · outbound

This paper cites Rebooting AI: Building artificial intelligence we can trust.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Rebooting AI: Building artificial intelligence we can trust

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.706679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:035a009d55516fac6a69cc3c99c5bdf13b4a6e615473d8d1184baa314a411743

Observation 74af0108-3cd6-40ce-b897-eaeba7dceb9f · outbound

This paper cites The Natural Language Decathlon: Multitask Learning as Question Answering.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The Natural Language Decathlon: Multitask Learning as Question Answering

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.370251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:14ce9e1d11ae4849b1468c1a96e4c4af136c021af3a1d776510a181b5b491b60

Observation 06e8bd09-ba43-41ea-a2c2-61486f56a500 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.747617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:bd34d99ffe5ea933d7ecdf25b024a81b529d1ae2d5418fca0db96cfaa12e0825

Observation d1ac8ec6-ae06-4a8b-904c-f91c8b6cb2e4 · outbound

This paper cites Gpt-4 technical report.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Gpt-4 technical report

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.740482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:92b895eef640972a47893c4557279d7443e9aa80c1914e31bc66dec1cda837cd

Observation 6151d5f0-9bb4-4d86-98e9-fc7ff67c8c8f · outbound

This paper cites Training language models to follow instructions with human feedback.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Training language models to follow instructions with human feedback

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.720716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:d7adabaa6f5327c3dc7126b810b8b2f6679a73a040646d2be2cf3da336567cd1

Observation 0b9a2912-9140-4006-8969-ab5871388564 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.397681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a85642d07cdfb3a0fab5c5c4a10f68db7a5b9281e855be1030fe96ebe7eb238f

Observation ab8fe06f-5254-42fe-9281-58575fc4fade · outbound

This paper cites Squad: 100,000+ questions for machine comprehension of text.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Squad: 100,000+ questions for machine comprehension of text

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.724507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5b218f282e6f89b6b642ab053ce7c3d77657ceb1cb8600f5e1667a15041d8f3d

Observation 18d8e589-4ddc-4a2d-b85d-e354f30e5b7e · outbound

This paper cites Internlm.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Internlm

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.735794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:736ede3bd7006cad11673c0dba6e9c2c924a2fabbc5097988879bda7f2a5a2c3

Observation d2e3e2d2-394f-4046-b05f-bd1c02611323 · outbound

This paper cites FEVER: a large-scale dataset for Fact Extraction and VERification.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models FEVER: a large-scale dataset for Fact Extraction and VERification

Reference 72

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.094484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:505603f66553ff1099df9ee9bfc0a7201053cda2883587b5ee0d5c7cd0cb4b58

Observation 9bf6514a-5cf1-4f39-86fb-05d3c8f08353 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.366393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2b34cf05c31d5604c3787a9322a0aa140e58652c0997bf68af82172c571f8bdd

Observation 12435b9b-5f55-49a2-9727-4ae09c88d1a1 · outbound

This paper cites Proceedings of the 2018.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Proceedings of the 2018

Reference 74

Resolution
metadata mismatch
doi, observed 2026-05-16T10:03:59.098487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1f7870f6b1151ab1abd9d5f720bd3530a0b4ae7bf6a89cc6e5b92e7d333c45c3

Observation 3f086877-2f1e-44f2-91cd-ee86bbf2cd0e · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.757529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1cd612e82b972e57388b6c56f1bfa82fd27479f6be1f4df8340bf1460d47391e

Observation a9d818a3-b884-4b38-bfef-3b7826380d7d · outbound

This paper cites From lsat: The progress and challenges of complex reasoning.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models From lsat: The progress and challenges of complex reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.727175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:44910b97edbfe5a00fd2e99473302f36454f8d4f033231c76dc7133219c8dfe5

Observation e7e0ed83-3f32-4c6c-b3f2-afe2dbd938ee · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.382955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4e2b7c2c26d2bfd2011809fd019e08d2dc9024fa2954ebd1ee81db7e3ba1ab98

Observation bcf85aca-e9c9-490c-a91b-5ee4cc34586a · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models OPT: Open Pre-trained Transformer Language Models

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:03:59.386291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:79a658419e2961eecbbe31b79d5b64b3792c98689e8bb830a7f9bdfd3e479364

Observation 4c12a26a-adaf-41eb-9701-0e88483c06e6 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Automatic Chain of Thought Prompting in Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:39:17.100641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:aee4e3b32a203d9ab7e4fd1e8a8791cd04d450f348768a27906bc97bda4ea1af

Observation 3a2a199e-9a60-4dc3-8da9-f61c1846653d · outbound

This paper cites Jec-qa: A legal-domain question answering dataset.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jec-qa: A legal-domain question answering dataset

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:10:50.729592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:4bfb358fb684a0b1820d05fedfb4e810cf9426ced4a8f589b20e4263ad0b3b43

Observation bd69c749-c915-49a7-b05f-c5c109380e87 · outbound

This paper cites Analytical reasoning of text.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Analytical reasoning of text

Reference 81

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.101368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:d145f0575844cdc56fb89374cefcc1dab1bab79ccc235d4c95ac6486054e4a0c

Observation 03fa877b-326b-4c22-ba46-2fae0045a960 · outbound

This paper cites an unresolved cited work.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:10:50.714389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:6977e5d534a3354def5e77e0f4ece8599b3dc7c6dd0a3b4a7b4240542d78ae10

Observation ab3b8154-9675-4642-a5bf-c6c1aa7d393b · outbound

This paper cites Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Exploiting Auxiliary Data for Offensive Language Detection with Bidirectional Transformers

Reference 83

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.104165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c74e23c904084cf22d6694d95cf6f1dd0a5088329135b0407efdd48027515ee1

Observation a4e67bf9-8441-4a5a-a2ec-0fa2fd704679 · outbound

This paper cites Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Modeling Profanity and Hate Speech in Social Media with Semantic Subspaces

Reference 84

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.107151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:27e0a2d129fbc23ef422dd8ecf1301fa340c620967c2dbc3f3db5be734c898a7

Observation 032ef19b-7e0b-4b5e-8f6a-c1fd2dc0406c · outbound

This paper cites H ate BERT : Retraining BERT for Abusive Language Detection in E nglish.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models H ate BERT : Retraining BERT for Abusive Language Detection in E nglish

Reference 85

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.110275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9cda7cf49af29bd4fdd7622ba0aab50ce671c61ff4f113b991e787606ab09d61

Observation 6bbb9fed-d274-4707-9a09-a45bd33f599e · outbound

This paper cites Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset

Reference 86

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.113286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e3b6937023a6fecab31f6dd7a60e3c6a76eaec7a6028b865cff77d8c341b4ec3

Observation d4658c11-1b9c-47d1-9f33-2e05499b2b5a · outbound

This paper cites Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Measuring and Improving Model-Moderator Collaboration using Uncertainty Estimation

Reference 87

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.115917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:9619c0f23e7b9998dd1e59c07f39bbee0f5852dd6b22d5fcf363e6ea445d92ba

Observation 830f09bf-d889-4895-8a9c-7c83090d472f · outbound

This paper cites DALC : the D utch Abusive Language Corpus.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models DALC : the D utch Abusive Language Corpus

Reference 88

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.118550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:a7c84f4305f1e9c315c9a4b4f6fa2444228b31f0032e57cf942c01aba69ba44b

Observation 90eceacf-ec56-47d7-af81-ce4bb74805ec · outbound

This paper cites and Dulal, Saurab and Koirala, Diwa.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and Dulal, Saurab and Koirala, Diwa

Reference 89

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.120857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8327f385575528bc3a41ca80acb9319d90fa53d19b37620223177f98a7c31c1b

Observation f8d4ca18-85fa-477f-8b1e-992095f15976 · outbound

This paper cites MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models MIN \_ PT : An E uropean P ortuguese Lexicon for Minorities Related Terms

Reference 90

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.123500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e71b748a15ca621359e101174ebc4671ec8a3a76f3cfedc92c1627c242ae726d

Observation 87f1d437-2827-454e-aa09-01332a2b4b52 · outbound

This paper cites Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-Grained Fairness Analysis of Abusive Language Detection Systems with C heck L ist

Reference 91

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.126081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8e2feb13cc44623648fca0912b28f815e49ec522076b4aacbbaf163acc704f2d

Observation a438ce5e-44e9-4f87-9313-281e8e6291ff · outbound

This paper cites Improving Counterfactual Generation for Fair Hate Speech Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Improving Counterfactual Generation for Fair Hate Speech Detection

Reference 92

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.129146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:26e80b5665ccdc1d7b0d9adc75999940486c001c8096a4155aa619203eac7f2d

Observation 9a4de569-3b17-4519-9810-27d09fea88bf · outbound

This paper cites Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Hell Hath No Fury? Correcting Bias in the NRC Emotion Lexicon

Reference 93

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.131534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:5e0ad243bcafdb7336d9dc17d159488b85bf490662c8794d927cc7b43977b4cc

Observation 96de1339-8b1d-4989-a1de-8005801b0fae · outbound

This paper cites Mitigating Biases in Toxic Language Detection through Invariant Rationalization.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Mitigating Biases in Toxic Language Detection through Invariant Rationalization

Reference 94

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.134410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:2675fdaf4d493b3c2f1df981bba293ccc1afdf4afb22698c7da9b13b73ce00a8

Observation c8ae52ef-7779-42a0-812d-a3f4844cfe7a · outbound

This paper cites Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Classification of Political Bias in G erman News: A Data Set and Initial Experiments

Reference 95

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.136789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:7e71f69090c7b1318210bdad2fd50acd9cf697e5e874cd33aa789dfdc172d315

Observation 274c958d-26ad-434c-9fe7-e55115ea91aa · outbound

This paper cites Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Jibes & Delights: A Dataset of Targeted Insults and Compliments to Tackle Online Abuse

Reference 96

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.139242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:6d09366f7fc2c5b7849d7ae308793cead94798eac36bad0755bb8be9a8a5b5e6

Observation bb2f7b6f-5349-4428-9352-f12a84e9a4bb · outbound

This paper cites Context Sensitivity Estimation in Toxicity Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Context Sensitivity Estimation in Toxicity Detection

Reference 97

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.141516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:035cde186b5f0f372375008da318a70f5270af1e13f2e2a737c64a8f31823841

Observation b9199c39-3889-4e5a-bb1c-95885278251a · outbound

This paper cites A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models A Large-Scale E nglish Multi-Label T witter Dataset for Cyberbullying and Online Abuse Detection

Reference 98

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.144558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:06a965a78fa87b01ebaa497c054d173a8158cf5a385ad711b1e020df4949140c

Observation 65c9bad9-4fa3-416c-bb00-4ecc618df2f4 · outbound

This paper cites Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Data Integration for Toxic Comment Classification: Making More Than 40 Datasets Easily Accessible in One Unified Format

Reference 99

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.147830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:8f0e17d414863a3d773f6cffdf4d8f96856d9164f57bf9cb895961bc471cda84

Observation f5cc10fd-b4ab-415c-85b1-fee0ab1328fb · outbound

This paper cites and H \'e bert-Dufresne, Laurent and Roth, Allison M.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models and H \'e bert-Dufresne, Laurent and Roth, Allison M

Reference 100

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.150099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:ed6a58d2e8127ae4cbba6143af1e2c5f6d5a26ceb206951afd32438c0667d96a

Observation 223e06c0-2ce6-44c5-899f-22a22fd1efce · outbound

This paper cites Targets and Aspects in Social Media Hate Speech.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Targets and Aspects in Social Media Hate Speech

Reference 101

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.152182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:c06b72d1aaddd46d903ed19057553a1070041b5dfa3c314d1f5526d80dd3cbf5

Observation 68ce1835-4606-4bb5-b487-31f1d36ecf11 · outbound

This paper cites Abusive Language on Social Media Through the Legal Looking Glass.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Abusive Language on Social Media Through the Legal Looking Glass

Reference 102

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.154771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:277b94b0875aacb25317dde4a1e498f0e41f6dd9c989caf32d2528b6fa8ebeb1

Observation fec804b5-2b12-44b5-a068-c28e9f61b2fa · outbound

This paper cites Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Findings of the WOAH 5 Shared Task on Fine Grained Hateful Memes Detection

Reference 103

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.156983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1d5c6489a0c5393001ced3ec16310ed39f177a8dc966e0124f616c144776947b

Observation b72c17f0-1090-4f73-be34-172e5f03780a · outbound

This paper cites VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models VL - BERT +: Detecting Protected Groups in Hateful Multimodal Memes

Reference 104

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.159625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:685395faef5bf364f75b88b13577d9a445b776f79ec129aa34e51d2b24b02b60

Observation a5ee5f4c-bd93-4234-aefa-7c4278cb278d · outbound

This paper cites Racist or Sexist Meme? Classifying Memes beyond Hateful.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Racist or Sexist Meme? Classifying Memes beyond Hateful

Reference 105

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.161674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:55a5f938b91b26a9e89c2f7d92499652830cbb5fdbbb1a98ac5b24e77abc7d89

Observation 470c38d8-3f6f-48ca-a061-7f05031e48aa · outbound

This paper cites Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Multimodal or Text? Retrieval or BERT ? Benchmarking Classifiers for the Shared Task on Hateful Memes

Reference 106

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.163828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e267d7c2ace8f99d0cf16072cd4b9e6705f604960241c954988175c974452e83

Observation e3570e05-a656-4a15-af58-58cb564b2b72 · outbound

This paper cites an unresolved cited work.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Unresolved cited work

Reference 107

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:10:50.759798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:1c26f6d2d4440622ca8e1ae3e27b04db6e479b343d71209bd619c4f979433061

Observation bca6162d-039b-4d9c-8041-29ee21286cbd · outbound

This paper cites Text Simplification for Comprehension-based Question-Answering.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Text Simplification for Comprehension-based Question-Answering

Reference 108

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.165782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:d93cd680af92c58ced94b831004451698933ac9bcfcf3880bfc8092ad0063b71

Observation 18b935fb-76c2-44c0-a730-b029aef26efc · outbound

This paper cites Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Finding the needle in a haystack: Extraction of Informative COVID -19 D anish Tweets

Reference 109

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.167986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f3a6e847dadb18f2bc0be3b4f4267c2cde2f09bc9d7645395579a2be4271a438

Observation 6cb98123-eb58-467c-9357-e1cb0cb67d68 · outbound

This paper cites Detecting Depression in T hai Blog Posts: a Dataset and a Baseline.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Detecting Depression in T hai Blog Posts: a Dataset and a Baseline

Reference 110

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.170033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:65330554091a68d2d55249315dd014774921740e57ca497f7b988eef1b6e8cc5

Observation 26398b59-dc55-4a42-be27-20a95b9237b4 · outbound

This paper cites Keyphrase Extraction with Incomplete Annotated Training Data.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Keyphrase Extraction with Incomplete Annotated Training Data

Reference 111

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.172227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:49be097f88a542e8fcb0603f6800e2b86714ff883c30ab24eb5237c589439edb

Observation 0d4bef67-af6d-4ce4-89ed-a86e3e0666c7 · outbound

This paper cites Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Fine-grained Temporal Relation Extraction with Ordered-Neuron LSTM and Graph Convolutional Networks

Reference 112

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.174264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:f0edd0d691f1baf56e95f9e96a33045e5e7b44a6879dbb4ba279b1bc9d6a7a06

Observation 5c5c900f-130c-4450-a8d0-0ee8507ac09b · outbound

This paper cites Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks.

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models Does It Happen? Multi-hop Path Structures for Event Factuality Prediction with Graph Transformer Networks

Reference 113

Resolution
verified exact
doi, observed 2026-05-16T10:03:59.176580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T10:03:58.971585Z digest=sha256:e3ac2bcc3a2de132eb0407f2d3762eb789af0b51c40d89c913e6347a49c10927

Pith citing papers

Observation c1b92dad-805a-47d2-943c-dc989a98b5ff · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:70b2352098e29c45fc75c35165122e56119cfc6c6aff776cef1a132063f216ed

Observation 96043340-8524-4b05-94fc-ebaf9b2cb6df · inbound

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning cites this paper.

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:46:39.673387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-17T23:46:39.330438Z digest=sha256:341427f6b493473ae9dd84a4163c78044720522a3915d17e6805592f0ab92e24

Observation 4ffa988c-4a1e-4c9a-9fca-c8d94dfff545 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.494391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:0f4e15485dc7a59223229f2a3e98ecf5db9f134cd86092b4f75b961f04d3aaf5

Observation f5134d17-bbc2-42da-b71f-ee7cedb4bc28 · inbound

Mistral 7B cites this paper.

Mistral 7B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:14:00.055653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-24T06:11:38.406350Z digest=sha256:cebe46906d11c49971207121514bef501f1c04db64160345d2766a9bdb334e32

Observation 9644ed94-4617-459d-9b84-777600b65e97 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.798370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:a1e45a7c07d703e560fb5aa686044f666bc3b0f8399f0d0a0ec05a53d6b5c92f

Observation 39ee5a74-f36c-49a6-928f-b280709feedf · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:9996ea17363039dee336b5a29b48ec88c3a140e1b2c890e211f009f632a2d830

Observation c718a91d-569f-407a-aaa2-36b2e00374d4 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 226

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:7fe9aba0fc40f70ba79d99861a4fa9bcfa66b898095d597811f2d2e8a0a05d04

Observation 188b5ec0-8e54-4362-9e77-31e427c34247 · inbound

GPT-4V(ision) is a Generalist Web Agent, if Grounded cites this paper.

GPT-4V(ision) is a Generalist Web Agent, if Grounded AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T19:40:13.816153Z digest=sha256:e4aaa996d31f30fbe7da7c0100db769472f868090fb7632e892f4c7406b8790a

Observation 29c50967-2771-46e4-b630-edde26399d5b · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:74405519567ad3cd0630cd2dfd74ccf7363180fe6874ca38af037ab1f0186158

Observation 321e39c8-7012-4f44-9288-52993abeb961 · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:13:53.841009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:7ae558f2b696931a9c822215f963d01110e58792ad03b527b06845a9cecc0c12

Observation 4283d2d9-6d40-47f3-9f11-39b38cf5d953 · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:23:49.423836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:a6f8402bc8b7fde82db993d0bcc60c7de7bc2aa816f0c2a9eb225e41ea9ff85a

Observation 2906ed29-c1a8-46f8-91f8-639cc21f41b8 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:eb5403facc24b40ab71d67b5251d4477ab2943b65a13773781b7dc70beee5b24

Observation 7403b0e0-3e8f-4a45-ada0-ab469dac8475 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:6e29f32d43b2868a250aa3c1b7ded2fa2e0865eb823141292ec70a76a850a587

Observation 82f29955-ba9e-46ba-a460-84dd8ebe4ddd · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:653d14fa4de1535ecb999358d8dc754ff8cc08b3bfce986b57f119b4efb6dd1a

Observation 9b4d48db-aefb-4c5d-8d47-bff912511c6e · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:734de7cbc60e2093f78d8747577a38f353c7fc844326b278b47e172932e90bcb

Observation 5a7a2327-af5a-4761-b4cc-95eb56ddd36f · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 220

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:58:17.383933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:a92b2b55694a3d9fb47bad7ddaf4bd2316cfe2af323651f808bd20bbe3a8e9d9

Observation 43ffd89b-4280-4515-87bd-f45c3388ec8a · inbound

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence cites this paper.

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-16T01:06:07.787696Z digest=sha256:644233b75acc706c0ed7558483c9e7bf4021b8ac443c295de9dfad2b56ab775e

Observation 8f1175dd-89d8-48f6-8840-c3271dc946ba · inbound

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence cites this paper.

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-16T01:06:07.787696Z digest=sha256:0339d733dd9998e91d836c3004fde4a9d24c090740528f106bf95fece4f79324

Observation bc643947-5dc4-4ac4-b4fc-a6db1271f161 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.324452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:e858f7903b74b19058fafee718a386a51a45a63ad7807b57c6aa35026162308c

Observation 3f745850-9722-42c9-9057-9f449b0d9d04 · inbound

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio cites this paper.

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:43:25.254645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-23T20:42:38.782232Z digest=sha256:812c6b0e7d9802ffe74a3854b7995699722ba45524fb1fc411b56d80fca8b823

Observation b0186bcc-c37c-48bf-856a-0451e49dfe8d · inbound

Pixtral 12B cites this paper.

Pixtral 12B AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T23:53:29.862702Z digest=sha256:412f8fec5d9b3529c4bbac4efe406ae947a04d6a84a704563ea79a9d9d8ef950

Observation ca57eb6a-6339-4196-ac25-3e59ea525085 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:52:37.581335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:5761059e253b88bc43ac3502c5264bbe3aa0963de15b5e1e038c77a228f6537e

Observation d6a8eaea-9549-4d9b-a2f0-5d5f2227275f · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:26af33d32a5581a7716a439fd9f93471cead356b16a73b67364d866746cb9918

Observation 2c8ebcec-b078-4056-813b-4508d264d1df · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 135

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.537102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:a02e5746fa19fa1f1e21024c0d5164bab3f6540600553a365f51ec412412e58e

Observation 16882dc6-a0b9-46f2-89f8-3a2bbe4acf2b · inbound

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning cites this paper.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 25

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T00:19:22.173928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:255430ed59fe0dfde8075f89de96477ce777fab7549c2dbab1148459faa48277

Observation 509a881e-5b88-4ba7-a0a5-e4b37d78091d · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:36:58.695985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:b06d14bef234156dfb86edadee4545f8e5d228ecb76baab268f92a4cb3385f47

Observation 73a841bc-ceb5-48fe-aa7d-35c89d7577a5 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:05:47.607647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:8d29336bdb24f563f36e66206086bc2cc67e8bfc845dda934eed86333ddff860

Observation 38d67280-fd08-438e-b4fd-f0f2b2311890 · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 200

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:01:10.053215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:9e28844d42d4a9710e09f878f6cde6965b91645e7325e50cf981468d3ce47489

Observation ad27a3f7-0084-4515-b40b-a8492c15271f · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:f23a7a8be35c4b2b8c88be06ddeb5f0d4c01b4201358d76a90f1bbde6dd281bd

Observation 2a5e72e7-5d37-4721-9dd3-7f6e9217b879 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.336650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:dd5d2e840dff8a4cb1dbaa15abe6c30bfcf04bba15b367dfccb0c72b7f5b46c5

Observation 1ed317f5-0155-49f9-b430-d44a63db8f32 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:00:41.543036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:055dd28f25c5e9ea3106bb358a51a51d02167156ba2d577340ff1358facbd337

Observation 5883d43e-d7bb-472c-a274-a212bbdc59ff · inbound

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models cites this paper.

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T21:45:40.615196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-21T21:44:36.351517Z digest=sha256:321eeed4c8fba257fedf2b45bd16c1f1d116f6703353d3e86a977ffeca63e627

Observation 88f763f2-ef0f-4d89-af4b-a892cf39b52a · inbound

Dr.LLM: Dynamic Layer Routing in LLMs cites this paper.

Dr.LLM: Dynamic Layer Routing in LLMs AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:04:20.318861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-21T20:01:48.806709Z digest=sha256:f89c403338a7f2fec74637aa75ee2846f14ed02ca95d70a8d31d4a5e733d89cd

Observation d083d6e9-23c5-40ac-9d7e-5de221cd30fd · inbound

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners cites this paper.

Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T11:01:17.080772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:59:16.139525Z digest=sha256:298700b212caedbf61a33d5df9a1f08e338226537de0142a87d83e1af78a22a5

Observation ba35216e-37b0-419b-8430-b27ba4c611e8 · inbound

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models cites this paper.

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:38:34.178169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-16T21:36:24.376401Z digest=sha256:db1536986b506ec4ef65d210159501074fd2f13510d025eb69dce58b1c3f1272

Observation a77233d4-f2aa-48e9-b4e8-4bad0640a905 · inbound

Ministral 3 cites this paper.

Ministral 3 AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:32022bd9797e82402c7f453b2951e6a493c768a8f449350ce7a37737adc26030

Observation fea78a68-21b2-43a5-a1fe-e961d5eda5a1 · inbound

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment cites this paper.

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T18:36:44.401045Z digest=sha256:162872e4910c603f5acd0e952435cea15b080dfc9f1371b94e5e60985b2305aa

Observation 9ca77fdd-8654-4f51-868d-f0caa031e0fe · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:9ba87ab9c3e7be58b73b07c982b92627d64a5b5e2fef531e73573762554dc4d4

Observation 4dac4e0d-3f34-4c37-95a9-13fc45d49019 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:5843de027005f3b0ac33ea75a8df7a520ffd1b0fa293c624f45ce3fc71b48ba1

Observation 58ab0514-50af-41b4-88be-d962e572afc3 · inbound

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation cites this paper.

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T18:30:03.367870Z digest=sha256:d8499ad75162e2b76c25f7049b429d96aecdddef966adbae50572ae5c152ce62

Observation 3d130690-2278-4685-82da-9f183a0feb22 · inbound

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability cites this paper.

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-13T01:12:44.476894Z digest=sha256:68c6232fb385fc6292905458857dc0d957ae4e8d262310cdeb969d3fb378d790

Observation 572da502-e57c-4723-8c76-364e78efbe62 · inbound

Confidence Calibration in Large Language Models cites this paper.

Confidence Calibration in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:a799e810cfcee796527e231b6566afcb9775f4c2f3d90aba9e099a40b2a7fe03

Observation 31646264-be84-400e-a269-aefcba1406d4 · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:16:23.517497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-06-28T14:35:48.292081Z digest=sha256:3cdaad3d6fb4a0cff7342045efd8e348af50a8aa01ebae686eca85fe8c7399a0

Observation 9176bcae-832f-49ed-944d-319fa2b21f7f · inbound

Do Value Vectors in Deep Layers Need Context from the Residual Stream? cites this paper.

Do Value Vectors in Deep Layers Need Context from the Residual Stream? AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T12:41:24.460278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:41:24.460278Z digest=sha256:685fdafd43060bb6972f7cb32420d0f9c6e94bc00dc9e4b4b0cf1f2d3060e344

Observation a23a1af6-d019-4ad3-a5ed-41e18f5a409f · inbound

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers cites this paper.

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:13:30.266780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-29T14:07:26.366172Z digest=sha256:58c7640eb3d8a89d4976425898ffb759d6427a475cc0cbb3ed6cedb5b711c889

Observation cc2f04a1-ac84-44b1-a783-a66334d63b63 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 245

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.972501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:b16ea0adce7ebf41d0a3061098176a48b145b41d1da20307d1ec5297ce1444a1

Observation 1aeeb5e2-0e18-4e45-92f0-ee6cf38d01b4 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 247

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.407055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.407055Z digest=sha256:d590593a362ce54b93eeb48b62e5e322302b7d6f8dcce94db6562e26660ccb88

Observation 3ddf1b61-ddc0-44ab-9adf-c73eade29cd2 · inbound

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams cites this paper.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:29:43.607054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-26T09:59:24.703388Z digest=sha256:60b9ffdb759bf39b22d160e8f876247d16778b029f342c1a6584395215a9089a

Observation bfa5fb53-4e67-4b36-8283-784db1cf8bab · inbound

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams cites this paper.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:05:28.242446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:acf7838bf801e5f47bd6a76542dfabdcd418271e0f7f15af499585b3c4433397

Observation 06aec8e5-42ad-43cd-a44f-86e5557cb9fc · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 230

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:59:44.874337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:48cc5deee77ec628e46f55390cc86041ea4cf84887ce554f482688ebfc97712a

Observation 0e387e7d-72a1-4955-909c-ca676c74e572 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 229

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:06:02.470889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:456c7fedf07949cfe4838eec2fea795f2138902825c1604091db2327e4425bb8

Observation b3832bec-1717-443c-92c3-1f82758dfc20 · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:0ded595a1f1d8b940d6c22ba705b33a16fc4792a29fac1254028f0718074da72

Observation 0299e798-3028-48b6-9547-ec7c6453b70c · inbound

Scaling Native Multimodal Pre-Training From Scratch cites this paper.

Scaling Native Multimodal Pre-Training From Scratch AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:05:51.646144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:05:51.646144Z digest=sha256:b6bc367bfe82cc60cde918ff30c930d8c0b8d8a7eeed839c90bc6a989161525e