Pith. sign in

Paper Citation Record · LEDGER

The Science of Evaluating Foundation Models

As of 8 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 3 inbound Pith citation observations for arXiv:2502.09670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09670 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.932998Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:59:00.294806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:35:26.411399Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42c9f172-4f43-45bb-a54b-f79576fea6e9 · outbound

This paper cites GPT-4 Technical Report.

The Science of Evaluating Foundation Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.614892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.614892Z digest=sha256:ef8eb01c4c1ee88489cdbf5fed2fd4530c9d873abcc03e4a7968ce857828b582

Observation 7a60ccf5-011b-40cb-a6bf-a6017ed74b60 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.619032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.619032Z digest=sha256:5d4b03a69758c414f2660d28417987d8d7c87bea70e94020db904387b51bdb65

Observation 6f7afbaa-a360-4dd2-ad22-312df632e1e6 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

The Science of Evaluating Foundation Models AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.621868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.621868Z digest=sha256:230e6987dce2c197ac8251239f5bf71fc0bb90c472dbed8ab7c638f57c291adc

Observation 8f75078f-c123-401a-a525-63783e7c7e83 · outbound

This paper cites Qwen Technical Report.

The Science of Evaluating Foundation Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.625010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.625010Z digest=sha256:b2eff3c7e7d32813d5163c0e5f83ed8ac0bf7320d43514a05dc846aa2d16d88f

Observation 1e7b9739-cbf3-4c4b-9f8b-764e250d0aae · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

The Science of Evaluating Foundation Models MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.628193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.628193Z digest=sha256:bdd78cb8786e27aa5e75159c5894b3cad820d41296bf51cc5c63639951893729

Observation 54d987c0-95d4-4044-9d51-e53b33a6e710 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.631123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.631123Z digest=sha256:9e2f63b61316babce039a1ef510b0bf7fd7953b5345c76333014419346d1cde2

Observation 41511bae-4253-4804-b53a-48a73a38451c · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.634066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.634066Z digest=sha256:bfdd6b358914ef8cf0db2d656b80a5acbdf204d70e297041e3f6e10508b1b05f

Observation b85ccb95-4a73-48d2-b1e6-a9281c26e127 · outbound

This paper cites A large annotated corpus for learning natural language inference.

The Science of Evaluating Foundation Models A large annotated corpus for learning natural language inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.637213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.637213Z digest=sha256:7c4d36f55ed7be894eb2dc0f9b1867a3f9ca92b85a678965a3c25a26dff6d0e2

Observation 5967fedb-279c-4f5d-b6e8-245121c56191 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.640482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.640482Z digest=sha256:a48d2b0f2cb5e865255db17c7ec0d12b17bfc9c6f7399fa4c3f2b3fa4cd53671

Observation b9a4d036-ebc9-4ecd-9edd-e3393beb80e4 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.643255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.643255Z digest=sha256:eecea6679072ac48e13503c6ccd5b9428c090521b4280b5d18ec0c1ba5f0c767

Observation 29df79b0-d079-4839-bd57-ba240b292bfa · outbound

This paper cites Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations.

The Science of Evaluating Foundation Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.646158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.646158Z digest=sha256:d4811c24caa490f4d4c6e65c3167164e6f6b36dc4e915b74a70994ccd433fd4b

Observation 7e87e0c8-0110-4424-b767-beb1980c8365 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

The Science of Evaluating Foundation Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.649410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.649410Z digest=sha256:38bc35bca5601409bc6fa5ba5e12183c0a00fad630fb8340bb00503da9881527

Observation 8ab6763d-8366-497c-8cf7-6fd8c8c54b69 · outbound

This paper cites Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification.

The Science of Evaluating Foundation Models Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.652396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.652396Z digest=sha256:6c3fe6aff480e48caba7e600411f7c9543e7dc4253df22baa3851b591315849a

Observation fe5d6907-4908-4224-bf26-be7280651b14 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.655458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.655458Z digest=sha256:3224852e1d25595dc4b4b09a95ab05e77b71e7a0fafc32b7e194064441ab7a61

Observation 1596f007-4aa1-4066-9a88-8c96bc8e9316 · outbound

This paper cites RobustBench: a standardized adversarial robustness benchmark.

The Science of Evaluating Foundation Models RobustBench: a standardized adversarial robustness benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.658408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.658408Z digest=sha256:ce801c2ed1ed78132a5683250a1235ab33264fce044057dd59d1be674078e732

Observation 80be2df7-f961-48de-a6bf-b87100cf0ffb · outbound

This paper cites Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems.

The Science of Evaluating Foundation Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.661836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.661836Z digest=sha256:72a4528bcd17685c73aae4cb3c3436b26038f4ae751803d139bd6fc2deca2f44

Observation 3537a145-0e41-4fe3-80e4-182d8abd2937 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.664945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.664945Z digest=sha256:aa0be746c0b478ba175bc46bd1f0b8c865c2f2d6ef2056f9fe387686df592a0d

Observation 45b93dc5-4227-475a-9f92-ee852864aeb3 · outbound

This paper cites ERASER: A Benchmark to Evaluate Rationalized NLP Models.

The Science of Evaluating Foundation Models ERASER: A Benchmark to Evaluate Rationalized NLP Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.667856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.667856Z digest=sha256:59c72e04344cbd9682e8f413d6281ef60f8f8ec4a9c85146ee8f76d1f1000e73

Observation e5204879-ae56-4307-a2e4-b57e582ad2d3 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.671333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.671333Z digest=sha256:1c32a3f738791af65fc4f8d1a462907c4f2682c219338106b1417bf7866e9533

Observation 5a6fbcf9-4b9d-4bf2-856e-f4cb37c65970 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.674183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.674183Z digest=sha256:75399717f61b2bae3162c8b3c1eb4832a32ac696ff11927482957d917f00dff9

Observation 706442f5-54ec-4af2-82af-ac355f421310 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.677278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.677278Z digest=sha256:0d72d9f25e9763d04a2964f9fbf2817f014bfd8809f21cf1a28f16065f65aac8

Observation b512d5ff-cb96-4239-a582-af833ebb7998 · outbound

This paper cites An Intersectional Definition of Fairness.

The Science of Evaluating Foundation Models An Intersectional Definition of Fairness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.680314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.680314Z digest=sha256:1b8ef5104a5a8613e7d53dc8f1ad249c2f1e0887b658cecf01d3898fe4de5777

Observation f66b685a-41b0-454b-b26f-954ce8367b2e · outbound

This paper cites Selective Classification for Deep Neural Networks.

The Science of Evaluating Foundation Models Selective Classification for Deep Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.683229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.683229Z digest=sha256:e2a35f329b19269464a6a629ee475556f84c517ff39c740fe3b3d9bd7ddbb3be

Observation 90609972-6f82-472e-8db8-3423a1fe6404 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.686588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.686588Z digest=sha256:8895343f9a671a6828fb7523a30cd3cfc3886309f736450ef23eec46f7bab6da

Observation 57fca449-93b7-4e35-9ca3-947cb0e00a98 · outbound

This paper cites On Calibration of Modern Neural Networks.

The Science of Evaluating Foundation Models On Calibration of Modern Neural Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.689465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.689465Z digest=sha256:c3fec9219cacffa6294e4ad740a7b1cfeb3b446eac2788db4bed8f7f4f81ee9c

Observation ed300267-aa21-4039-8308-acb5efe10507 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

The Science of Evaluating Foundation Models Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.692496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.692496Z digest=sha256:7b7bd6f6482a9cd19a3d162b8412386df9bf06d82a04788808c5766f82b71bce

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:96a4617e016965e283bcb7378f63fce9dba7afaff713ee38e842be946b723089

Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · outbound

This paper cites The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models.

The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.698050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.698050Z digest=sha256:18b26d1af8e580f81a3080b77caea6715995e3442fdb3f6665d6900d08a87009

Observation 12b41117-de41-4783-9142-bb499db5a010 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.700573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.700573Z digest=sha256:994af9e8665f98ece658eafe2fed2127514f7f1734b327c035d07063c43fdfe4

Observation 0a328410-f232-4698-99d7-280aaf548897 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.702763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.702763Z digest=sha256:577f5aeb27fae92b9218186adceace62313ac126983dda995d3409f3b54379c6

Observation 8f579ee1-1005-4ee9-8589-065189ae3929 · outbound

This paper cites Mistral 7B.

The Science of Evaluating Foundation Models Mistral 7B

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.704954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.704954Z digest=sha256:73ee9b9ef53fd5f7f4dda794f28430466e7da50f33f90e8e3c30bcd315f21926

Observation d995213e-6350-442b-842a-42a32a09be59 · outbound

This paper cites Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment.

The Science of Evaluating Foundation Models Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.707286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.707286Z digest=sha256:f7ab7d48e1c5defb32c8d9359d81a16d027b042049a0464cd246cf3eb1e8f390

Observation 53901a15-baae-434b-a914-f309e49e8e35 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.709720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.709720Z digest=sha256:c5bac3fc45c39f53ebb7157d91611534915fa748fda947c604caf92ab259830f

Observation 972216ea-dae4-4d33-b25f-1cb12a26cb26 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 34

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.715639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.715639Z digest=sha256:08072542eddf9b0005359ec0024eb9e949c0786720fe07e10b82fcaf6489ff3a

Observation 552ef713-bd7d-483a-9c03-a35e71a76be5 · outbound

This paper cites WILDS: A Benchmark of in-the-Wild Distribution Shifts.

The Science of Evaluating Foundation Models WILDS: A Benchmark of in-the-Wild Distribution Shifts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.718445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.718445Z digest=sha256:d845d0ddeb52976274e80dfab78acf59c782be083a8921c16e7a0d22d6ad7f35

Observation cad99ac7-c79a-42bf-90e9-c92e65216222 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.721987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.721987Z digest=sha256:2f5a49e5e1aa8a7a7876866ce42597c40b3265192e28ca114921fc0eee8aae47

Observation a59ec1e8-3887-4873-9d5f-81906eb6e7b0 · outbound

This paper cites Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov.

The Science of Evaluating Foundation Models Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.724900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.724900Z digest=sha256:e3161fe1b702a2b8aa17f389a4664dcadef0dede23af2ef8d48577a729edbcd5

Observation 075f5f71-150d-420a-a730-4a4277b33180 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

The Science of Evaluating Foundation Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.727917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.727917Z digest=sha256:a8c26d00333513ad6c1eb5e83db3a5f9630f2e8283fad0d58f1466a1a8ad312d

Observation a619e8b3-ff6f-4ee2-8d17-20b15e5522da · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.731195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.731195Z digest=sha256:fe321baaf9003091f010037fde32d112a8d8cb3a69ca8100c8715d8c8f147c82

Observation 65441e79-c127-4d2f-b3c1-c41653c992cc · outbound

This paper cites Evaluating Human-Language Model Interaction.

The Science of Evaluating Foundation Models Evaluating Human-Language Model Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.734153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.734153Z digest=sha256:6d7fae28ee72a29f153f8b9d3f46876ec033a755f8fdbd41bff81be982515f5b

Observation 6be26943-f8c2-421c-884b-dba5cd61f4c9 · outbound

This paper cites Can Large Language Models Capture Dissenting Human Voices?.

The Science of Evaluating Foundation Models Can Large Language Models Capture Dissenting Human Voices?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.737130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.737130Z digest=sha256:80995e6200c7ecad4d9a3990cf8efb7e15ea99d16970367cdd6e95503e789fd9

Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.740185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.740185Z digest=sha256:66aac13685b2064a826a4a18c966de2ccd00d23e56aa488100e0ff991c7595c1

Observation 2a33283f-1f8c-4b73-8fa4-c8bc1bec5053 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.743247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.743247Z digest=sha256:17cda03818a582d862c9a1e68297153700496fd3a6746734c43ba39183a21298

Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · outbound

This paper cites Holistic Evaluation of Language Models.

The Science of Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.752311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.752311Z digest=sha256:ac88178e7577f6d6dab08b604d18a9b0a972e2ac4219874e3bf4433ebda45350

Observation f4539750-84aa-460b-904f-863ae408ad74 · outbound

This paper cites AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents.

The Science of Evaluating Foundation Models AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.385959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.755226Z digest=sha256:3707620e47cfeb4431a28f3123d5dc56d202764fc534ba7d5389f5a670efa08b

Observation f9d79ee3-df9a-4330-8066-1f8c920c6163 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.758625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.758625Z digest=sha256:26758b267b8a963f4d3bef74f8a0622ea6fe77f800dcb9c6cc95b7ba05e97c84

Observation 6338a6ff-677a-4838-86d6-58c4a007de11 · outbound

This paper cites Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models.

The Science of Evaluating Foundation Models Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.761353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.761353Z digest=sha256:6b33e08436335f1bca3a359ff125532b83db3834d56ee20f17e3d47cc083460d

Observation 358695cd-1638-4b11-81ea-3e5ca8383f1b · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.764319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.764319Z digest=sha256:71e78da392b2ae0bcdaf1b2db5f44db5a7541f21bc970becff05d6c21854a455

Observation 3620ce70-def8-4509-b833-ea60b5484b25 · outbound

This paper cites Social Bias Probing: Fairness Benchmarking for Language Models.

The Science of Evaluating Foundation Models Social Bias Probing: Fairness Benchmarking for Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.767138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.767138Z digest=sha256:1ff820c0bb2a47e454d51f0838c09181dd30f1cab4acfdeed53c1d4b56b23a22

Observation 9ad605d3-0278-4cc8-9617-9aee498174d5 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.770006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.770006Z digest=sha256:29b9baa8f8ecd2dbcd0855dd9ddb26281a4067ba45a2e5171545e8d952e4084c

Observation c116ac70-da36-4744-993d-1caae578a27c · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

The Science of Evaluating Foundation Models StereoSet: Measuring stereotypical bias in pretrained language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.774783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.774783Z digest=sha256:6a1975934c69feb9d83d3c4c47109b71e387dc8efa76f9fd05ac63a5f139a70a

Observation e110ec71-f37a-4eee-bcae-08209c44724d · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.777377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.777377Z digest=sha256:ca44ee3281c18b831f87fabb7316cc4e5bd277e908a511da9e6ff72b12b52436

Observation d9a87917-5e9f-4fd5-9f10-1e37eab5d1a8 · outbound

This paper cites Pointer Sentinel Mixture Models.

The Science of Evaluating Foundation Models Pointer Sentinel Mixture Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.772371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.772371Z digest=sha256:c19e604ba5c895e9cac822560aea2230469d1fb2d7dd5f0c8d680af253eb597f

Observation b030956a-0cfb-4015-96ce-8a20373fa5e8 · outbound

This paper cites Cohen, and Mirella Lapata.

The Science of Evaluating Foundation Models Cohen, and Mirella Lapata

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.787601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.787601Z digest=sha256:173a7881e79cedef2a38d8256eaaf3f492f6a2f8f581d2c5ab45cf8f206fe836

Observation 1f715672-1eff-412e-8e4c-c29850f55252 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.790559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.790559Z digest=sha256:9cd6257895a5b35e1829b4e02a6024aaa41e387806fc5e0fe2e9524e9c163d34

Observation fe1fa383-b0af-4bd7-befa-bbc665696e72 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

The Science of Evaluating Foundation Models Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.779786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.779786Z digest=sha256:2fa73e296b9d1ca5cf921491ffa5b2cb29c2bddcca3d4d0ded7ee2144a1591b0

Observation 81653e57-c00c-4a4a-9b16-663d87670804 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.782282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.782282Z digest=sha256:f341beaa834385f5d5f1b125555adf0c94131b46799cc9065a1255cbe15c5b55

Observation 90319e29-dc75-449c-87b7-3c6aa7c23604 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.802455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.802455Z digest=sha256:c8dab540020dad979e2edd8472e0c343e05d8a8e175d1e3e406e39a0c02fc998

Observation 98b069ff-6c98-47d4-8c32-b0b209f1ef16 · outbound

This paper cites Is ChatGPT a General-Purpose Natural Language Processing Task Solver?.

The Science of Evaluating Foundation Models Is ChatGPT a General-Purpose Natural Language Processing Task Solver?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.805153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.805153Z digest=sha256:2b57b004039b429ef3ce47c502ec47fba1f56dc55fd2adeedd8264e1ce754043

Observation bb7c04e8-9804-4129-8c8b-a96455d5984a · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.808366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.808366Z digest=sha256:a92db4feac379128e6db3c79ab1b07688a913cb0f290699841ba97c634c219a3

Observation 2d65a7e5-c86c-47ff-9d8a-40cf58e7980a · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.793354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.793354Z digest=sha256:7c4dfb16f88bf89a8b909d464306c9aed5dd46ce24ff1788d4486cee77513d41

Observation 23278487-8ae8-4dd8-b04e-786112a023a7 · outbound

This paper cites Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

The Science of Evaluating Foundation Models Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.817579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.817579Z digest=sha256:f0ae19c46928e7997b1d085e6f1e4b43a06ce8dbafcd418ca30114faa4798aed

Observation b750a0c5-1e45-4cf6-bb69-966979345786 · outbound

This paper cites A Survey of Useful LLM Evaluation.

The Science of Evaluating Foundation Models A Survey of Useful LLM Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.799324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.799324Z digest=sha256:c07a878bb87c334810a9f18d0e93dbaff56050d62dc50faa2903de8f4153aef6

Observation 2c0b10e6-86cd-480f-9519-3439a79e9530 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.823750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.823750Z digest=sha256:2d857e23cb6ff2dda43e5fe88dc0b4a4391d7aa8a09450b32a5be7e0bc54b33e

Observation a1f5741e-f9c0-45b3-a459-22c647d2a224 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 66

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.829679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.829679Z digest=sha256:382068dd432a533f0233733e3d117c46fada720dd826c0231ca7b2989cbc07df

Observation dec6e7cb-b5c9-48f1-8070-dfe70d40a234 · outbound

This paper cites Rush, Sumit Chopra, and Jason Weston.

The Science of Evaluating Foundation Models Rush, Sumit Chopra, and Jason Weston

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.832607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.832607Z digest=sha256:2fd0cf80c063adb69fc670fc3d246e988ee7e366522f062c2e2406a35b13fb73

Observation 13445a94-7ada-472e-be00-90bbbe398e95 · outbound

This paper cites Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition.

The Science of Evaluating Foundation Models Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.835466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.835466Z digest=sha256:7f1de60578fdc3eef5fe722105ecf348bbf492c19318cd22d1389c52a99ca9de

Observation 01e16b8e-c3f9-4db7-b028-afd108b7d7fe · outbound

This paper cites LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models.

The Science of Evaluating Foundation Models LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.814464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.814464Z digest=sha256:a4eb54ed815cef07fea53a2360c3b167ccc3563c8a5f4462cc5ef7dae532e355

Observation 5f764ef2-c647-432d-b8c1-e4e3a3f9e7d0 · outbound

This paper cites An Interpretability Evaluation Benchmark for Pre-trained Language Models.

The Science of Evaluating Foundation Models An Interpretability Evaluation Benchmark for Pre-trained Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.259889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.841478Z digest=sha256:cbc6454f1a398541c83026c616dd76ec4ae14d5d5590d1881bb1a13ae670ecf1

Observation 5e1c44a6-b37c-4e66-a242-719aa6949594 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

The Science of Evaluating Foundation Models SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.820557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.820557Z digest=sha256:f3ef27d68ba4223389df98b3e30c53e618779283675ad7f843612b10e7e694a0

Observation 781e5b33-a901-4679-8be3-d59c8d9d1e48 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.847446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.847446Z digest=sha256:b57b819631b59e969e200d44dd675b1dabe0da2bff107d38e7450c5a0e02453c

Observation 4bd856cc-13c9-4883-9f6b-0f6a181f49bb · outbound

This paper cites In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.).

The Science of Evaluating Foundation Models In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.)

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.826508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.826508Z digest=sha256:845054dd6a4a220b48c6d76f14cf4cc994242a3294dcb5b61c52e56d7e7bf4a5

Observation 3a20dc31-a722-4aef-8497-83847d01c642 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.852413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.852413Z digest=sha256:78321e3ee355a965ea95ba6b99dee0bd03fca7ba7b1dccb82cf3b0d4ca8d036c

Observation 70765af2-18bd-473d-ab69-26c73a29e4ad · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.854802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.854802Z digest=sha256:c43d09216c3c067cde238e31e042668f6b904d40e4aaa3b5da5fdbba0df72108

Observation d7c248df-e0b5-46e8-a888-d41c432aec1b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

The Science of Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.856940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.856940Z digest=sha256:357ba42ad50a5aeb2a9954a4e7d406d5f2991bddbed0d98083dcd4b38880a628

Observation eb1aeec2-bb7d-4074-a21e-c13535ff7ecc · outbound

This paper cites Evaluating Large Language Models with fmeval.

The Science of Evaluating Foundation Models Evaluating Large Language Models with fmeval

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.272176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.838370Z digest=sha256:05ab7cf69afe7fa0ca7a70b16336ef85c98ab9c6996715817ee1ac347c8d25ed

Observation e5469b24-71e6-4831-9b20-84825d9d63c7 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

The Science of Evaluating Foundation Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.863898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.863898Z digest=sha256:56346af62f073b5dbfd22199af7f06e2150056708ac061cd455e8bd7a7fdc8ea

Observation 6bd40893-61f8-4de3-ae64-01645250aaba · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.844395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.844395Z digest=sha256:42c5baa1e218f8d502099dbd64e0de01c2972e92a2bc2ff2c4b5bf5524fe9572

Observation 0955883b-c37d-4f8a-a42f-19e68cc01e40 · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

The Science of Evaluating Foundation Models DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.869827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.869827Z digest=sha256:2e755de5c924b001eb3a30ff4b1a5b7a65557e6635632c4401cc7652eb56f75d

Observation 9dab2ad0-c7f6-4b39-8ac3-479ce16c67f5 · outbound

This paper cites LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models.

The Science of Evaluating Foundation Models LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.246503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.849901Z digest=sha256:8dc8a2270918d337a1707b9476e1df498dc364b5788d8a20c62f511da2cc6f9f

Observation 3e92f87e-2686-455f-83f1-4bca005cd247 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

The Science of Evaluating Foundation Models AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.876049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.876049Z digest=sha256:7259e9cef09986b8cbfa280396e8c6c0c8fa673822e5970e71d1e56c50404f3c

Observation b31ecd82-ff6f-4bce-8643-30eec429bcf5 · outbound

This paper cites Document-Level Machine Translation with Large Language Models.

The Science of Evaluating Foundation Models Document-Level Machine Translation with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.879026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.879026Z digest=sha256:50011ccc222c334825cd7ded05a02dd07c28a83119bd875b24663cdbe2c2f124

Observation 58fd2708-310c-406b-a025-44a2c6c04a52 · outbound

This paper cites Smith, and Teruko Mitamura.

The Science of Evaluating Foundation Models Smith, and Teruko Mitamura

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.882205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.882205Z digest=sha256:62de045933818397c86be91f8141d2f24c118abd34a8bfd736cd8fed5fe9b560

Observation 5df30076-7d26-4ce6-8d0c-9d2f199cec50 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.859441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.859441Z digest=sha256:d4c1566f146467aa881d897c9ac2f2af8609c0de6500e41a0146918c01af483d

Observation 6ce5459a-8839-4d48-9da7-b3cf16f2b55f · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

The Science of Evaluating Foundation Models Advances in Neural Information Processing Systems 36 (2024)

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.861678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.861678Z digest=sha256:40c0b3958110ee55b7e08bbee5cf51253dc9a42a0528472f273e2aa17b62a1cd

Observation ba2eda73-8fe5-4972-a8bd-8add8a617f16 · outbound

This paper cites DHP Benchmark: Are LLMs Good NLG Evaluators?.

The Science of Evaluating Foundation Models DHP Benchmark: Are LLMs Good NLG Evaluators?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.891112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.891112Z digest=sha256:1a7257887b531ed0d39ea9bf3901f453280ce2cfbe3eb32c08a95a61be34ef7e

Observation 43dc27af-db77-471d-9052-f793109c05f7 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.867035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.867035Z digest=sha256:720adad7621fc88ff9ecaa012412207533415f817070f2e40c81cddc116f9d14

Observation 6fd23ec5-6389-4d9a-92ee-cd3f27240be0 · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference.

The Science of Evaluating Foundation Models A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.897234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.897234Z digest=sha256:3fd57d2c1ca2e9bcd249e05079701477a84ec82ef45a8df335212acf35dff9ca

Observation 51a6cad7-7eab-4dbe-9466-5731b9e54fb6 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

The Science of Evaluating Foundation Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.872968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.872968Z digest=sha256:054284d6559dd6b96f9248d1489ade304ecb11b13132df54574d5187a502e24f

Observation 6b7b782d-e3d9-49aa-8842-b8816314797e · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.598649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.903121Z digest=sha256:ed7af22747aaac7f63f482a7942d99fb4da8510917950526dd600473c5b6030e

Observation 10a31815-2774-45dd-9871-f3e508c4e89f · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.589935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.909203Z digest=sha256:fe3d6d8f37022f87236b2b6a903b50a09b1cdcd4cbecfc3e829276109ae2d8a0

Observation 2639926f-7a1a-44a4-bd2a-1327389db7fa · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

The Science of Evaluating Foundation Models R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.912130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.912130Z digest=sha256:09bd491458ada3291ad6f3dec423faa5cb68d111ea3caaee30d79d915903a466

Observation a40c18e0-4c67-4b32-bd52-1fb541fbcbcf · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.885080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.885080Z digest=sha256:b323ab94c64eda1ff6579e40300e82f4bd4a85fba072aab0ada566a43c889675

Observation ffce3c16-92b4-40d9-a5a9-e9114ee65dc5 · outbound

This paper cites A Theoretical Analysis of NDCG Type Ranking Measures.

The Science of Evaluating Foundation Models A Theoretical Analysis of NDCG Type Ranking Measures

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.887915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.887915Z digest=sha256:26a7b59f56ab1420991b13516703c6b36d67fc781585c9146f942e4920020e6f

Observation 8a072e5a-e6ea-4d15-8dcd-93e24a914f35 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.581098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.922646Z digest=sha256:645e6208d4125da2872bcdf81f543be135c9621760b6e8e64d3ea66464c4f0c7

Observation 9ea7a298-1b5d-4e66-81fa-4c2d9a97de88 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

The Science of Evaluating Foundation Models Aligning Large Language Models with Human: A Survey

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.894068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.894068Z digest=sha256:71183d47e56fdf6459135511082e0bbe4d7be2f851ce2bb90bd12c668c0360df

Observation defb6ea1-6aed-4bd9-b8e7-c2bba422d85b · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

The Science of Evaluating Foundation Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.927960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.927960Z digest=sha256:bb823df35060a8ccaefa77a28076bcbdee6b74e2537dbd180d99f71c04154950

Observation 3a059a11-00b6-45aa-9180-3e42089465e4 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.606460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.900246Z digest=sha256:b4a67243e3ef9b88297783ab2f8c571914b3e1eafab1ae33319d168dda893799

Observation 4c8835d8-bdec-40c5-981d-85d628bf4a3b · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

The Science of Evaluating Foundation Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.932998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.932998Z digest=sha256:b349f2d496a1cea6701ae2c6755a84a346d42fd0fe6a612ec414458bb157a13c

Observation 32c74742-06d8-4d11-b2cf-0bad6dd002e1 · outbound

This paper cites LLM as a System Service on Mobile Devices.

The Science of Evaluating Foundation Models LLM as a System Service on Mobile Devices

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.905965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.905965Z digest=sha256:32a576990f58ff02c81dc5ba5f71ed10878ee79773b7e634c62618a2ff7923c4

Pith citing papers

Observation 2fabdbc0-8037-4596-bd31-d827de9ce9cf · inbound

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch cites this paper.

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch The Science of Evaluating Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:00.294806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:59:00.294806Z digest=sha256:4586a632dce2858219b951a08938753b8cbe6e89d43340e1d653570325c7609b

Observation df984a59-0850-4fc8-88b0-ca88a2f458de · inbound

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models cites this paper.

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models The Science of Evaluating Foundation Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.412843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:31:46.940449Z digest=sha256:ae69cc45185dbeff9826f334b8bc7f5039496c04e6a46a3536aa86c361485f03

Observation 817c9b0d-9628-423d-9255-866ac54a79ff · inbound

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis cites this paper.

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis The Science of Evaluating Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T20:50:53.567415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:50:53.567415Z digest=sha256:8e0fab99237884c37e2493b3538392ac20ae18e2d956dfaf641df9f351dfb70d