Pith. sign in

Paper Citation Record · LEDGER

The Science of Evaluating Foundation Models

As of 9 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 3 inbound Pith citation observations for arXiv:2502.09670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09670 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.932998Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:59:00.294806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:35:26.411399Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42c9f172-4f43-45bb-a54b-f79576fea6e9 · outbound

This paper cites GPT-4 Technical Report.

The Science of Evaluating Foundation Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.614892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.614892Z digest=sha256:57c7456c76eeac5d762b2fc84070f5d2ff4ab098d9c331d56e6cd789fe9196ed

Observation 7a60ccf5-011b-40cb-a6bf-a6017ed74b60 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.619032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.619032Z digest=sha256:b9126dd5ddcff9c94bf2d335aaf5956ad8f12cf52a10e32caf671c162a4b0738

Observation 6f7afbaa-a360-4dd2-ad22-312df632e1e6 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

The Science of Evaluating Foundation Models AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.621868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.621868Z digest=sha256:9a04f189b07d666bd4afc01056670c0bd469faebd6ef5459bdd8b2bb6a434433

Observation 8f75078f-c123-401a-a525-63783e7c7e83 · outbound

This paper cites Qwen Technical Report.

The Science of Evaluating Foundation Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.625010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.625010Z digest=sha256:d8131799df6439aa9f95d327931a3c94ae4358761fb13ea5da06e8e2ddcc0477

Observation 1e7b9739-cbf3-4c4b-9f8b-764e250d0aae · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

The Science of Evaluating Foundation Models MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.628193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.628193Z digest=sha256:0c57142dbd3eb5b664a6b8f6df7af09e61b99dfb4b5730393a3b04b499a5c331

Observation 54d987c0-95d4-4044-9d51-e53b33a6e710 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.631123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.631123Z digest=sha256:0fb63c753da2048f9a23b822c6b3697217a6194245f949e6bb5c2d9084090eea

Observation 41511bae-4253-4804-b53a-48a73a38451c · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.634066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.634066Z digest=sha256:1dcff158f7d6f5b4db604cb025c6596f136e014f0783e38880d96bd0aa990780

Observation b85ccb95-4a73-48d2-b1e6-a9281c26e127 · outbound

This paper cites A large annotated corpus for learning natural language inference.

The Science of Evaluating Foundation Models A large annotated corpus for learning natural language inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.637213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.637213Z digest=sha256:86cf42641970150a34341969f67c027ad1000ce88c589b0a1c9f5b81c9cb3f7d

Observation 5967fedb-279c-4f5d-b6e8-245121c56191 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.640482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.640482Z digest=sha256:fe6172572dc5fae984b4bcd2da079bfc0ea3ebd891a8558b08dbc87dad8a5f9c

Observation b9a4d036-ebc9-4ecd-9edd-e3393beb80e4 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.643255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.643255Z digest=sha256:03e6421089aee3c4751f6ff48459882f90ad6ccf5ec8d7e228d9be84e84c4452

Observation 29df79b0-d079-4839-bd57-ba240b292bfa · outbound

This paper cites Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations.

The Science of Evaluating Foundation Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.646158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.646158Z digest=sha256:f7279aba95c19ea4e4e5d99e252786ae181841d7d17313b08b5b072ff69590f1

Observation 7e87e0c8-0110-4424-b767-beb1980c8365 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

The Science of Evaluating Foundation Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.649410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.649410Z digest=sha256:cfdba691415674937dbd75594611d97173fb1a736472cd00888dd0cf467c1748

Observation 8ab6763d-8366-497c-8cf7-6fd8c8c54b69 · outbound

This paper cites Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification.

The Science of Evaluating Foundation Models Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.652396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.652396Z digest=sha256:32d5dbf8d5a6b881370f78cd0d9518ee7b2d4cc67abd689f593617e1c0ed973d

Observation fe5d6907-4908-4224-bf26-be7280651b14 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.655458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.655458Z digest=sha256:691bda582e6de285a40ee01cbe15b2e9f2e4941c44835796acdf488e960b1a51

Observation 1596f007-4aa1-4066-9a88-8c96bc8e9316 · outbound

This paper cites RobustBench: a standardized adversarial robustness benchmark.

The Science of Evaluating Foundation Models RobustBench: a standardized adversarial robustness benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.658408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.658408Z digest=sha256:66440bdd0622309fe8b24bdb2e3d5125e4ead3e0f7809dec63fa9e25e47fc6c4

Observation 80be2df7-f961-48de-a6bf-b87100cf0ffb · outbound

This paper cites Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems.

The Science of Evaluating Foundation Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.661836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.661836Z digest=sha256:961ed1bbaf459cfac453141118ea2836172edcdec9cb1c08757aa71a9e02bc17

Observation 3537a145-0e41-4fe3-80e4-182d8abd2937 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.664945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.664945Z digest=sha256:a80d00f06c457cdb825e8ad0b6544e8534436e2952f6ae27f59bc576d2e5699d

Observation 45b93dc5-4227-475a-9f92-ee852864aeb3 · outbound

This paper cites ERASER: A Benchmark to Evaluate Rationalized NLP Models.

The Science of Evaluating Foundation Models ERASER: A Benchmark to Evaluate Rationalized NLP Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.667856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.667856Z digest=sha256:a22d31be6718dc6aeea22a48ed48c7bfb010a55b9137387c0b2f06118279da5e

Observation e5204879-ae56-4307-a2e4-b57e582ad2d3 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.671333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.671333Z digest=sha256:03e3b4fca6be2691a346003987b1c060052b39f5bad34c8b8d1d8d3fb156217f

Observation 5a6fbcf9-4b9d-4bf2-856e-f4cb37c65970 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.674183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.674183Z digest=sha256:9e4d9fe9e4bc3cf6200c36d4ad42fe001a4fe0f4e4dcf783abed87b96600cd5d

Observation 706442f5-54ec-4af2-82af-ac355f421310 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.677278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.677278Z digest=sha256:b39ea01afa1e8b2ad5efcb8f72a721b9498832e38f30876b08102d832ac46c29

Observation b512d5ff-cb96-4239-a582-af833ebb7998 · outbound

This paper cites An Intersectional Definition of Fairness.

The Science of Evaluating Foundation Models An Intersectional Definition of Fairness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.680314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.680314Z digest=sha256:e6744aafb07f6e2c0fdf00783ba728b3e11ec32b70dc248b99ecdc92cc78bdf6

Observation f66b685a-41b0-454b-b26f-954ce8367b2e · outbound

This paper cites Selective Classification for Deep Neural Networks.

The Science of Evaluating Foundation Models Selective Classification for Deep Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.683229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.683229Z digest=sha256:c3c80c775e5e5185ad4b03246ae558d557e8ec2ce3ae04826b96089319d9fa31

Observation 90609972-6f82-472e-8db8-3423a1fe6404 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.686588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.686588Z digest=sha256:c01a617228f019d5c84bd7d99ab15bccb102510f1f7c91962a7bebca66029a1c

Observation 57fca449-93b7-4e35-9ca3-947cb0e00a98 · outbound

This paper cites On Calibration of Modern Neural Networks.

The Science of Evaluating Foundation Models On Calibration of Modern Neural Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.689465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.689465Z digest=sha256:4be5b7e38b951d667b36b7046ecc96b6f87415753fbe2f702958fbe72534e97b

Observation ed300267-aa21-4039-8308-acb5efe10507 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

The Science of Evaluating Foundation Models Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.692496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.692496Z digest=sha256:60d3fbe93786fe2250a5c6220c1e48a733cbad7a37ea402d37f72387de51d58d

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:668774763e10a26a35407e4c29ad0d4fb10cf0d68d526b52bd5b4c2128475016

Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · outbound

This paper cites The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models.

The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.698050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.698050Z digest=sha256:dd308ecdf251826a6c0b2917f3515075b107dcadb442b62b16864f6d0142d8e8

Observation 12b41117-de41-4783-9142-bb499db5a010 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.700573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.700573Z digest=sha256:a934a7f4d00d3b19a3f592f837b7971230c6f41fa5e0f48be19ae98f340819fc

Observation 0a328410-f232-4698-99d7-280aaf548897 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.702763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.702763Z digest=sha256:777c1db997c00dc07b688ed7b54e93ff28de0c459245f01b70869ac2072c06f3

Observation 8f579ee1-1005-4ee9-8589-065189ae3929 · outbound

This paper cites Mistral 7B.

The Science of Evaluating Foundation Models Mistral 7B

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.704954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.704954Z digest=sha256:e4db00819007331a9648eb4a7080e12d707586c6d630cb0d1a7f4a2275cc077b

Observation d995213e-6350-442b-842a-42a32a09be59 · outbound

This paper cites Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment.

The Science of Evaluating Foundation Models Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.707286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.707286Z digest=sha256:a543cac947f6a3b69efef3dab87c3d7b722fa22ead4077202928feb8494e991f

Observation 53901a15-baae-434b-a914-f309e49e8e35 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.709720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.709720Z digest=sha256:667f728ad927bc0a64e06336c3080fe40464b8a570816e8dfd85d6f07efda320

Observation 972216ea-dae4-4d33-b25f-1cb12a26cb26 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 34

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.715639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.715639Z digest=sha256:0643298dd7025f5cd5a03e1a0ecc5e0fb546c045a1c8b18425d93c6fa95680c1

Observation 552ef713-bd7d-483a-9c03-a35e71a76be5 · outbound

This paper cites WILDS: A Benchmark of in-the-Wild Distribution Shifts.

The Science of Evaluating Foundation Models WILDS: A Benchmark of in-the-Wild Distribution Shifts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.718445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.718445Z digest=sha256:ec356eecb906f48c61bcec328766967421c69108e054999549929348e10e6251

Observation cad99ac7-c79a-42bf-90e9-c92e65216222 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.721987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.721987Z digest=sha256:fcf0f810e42765e0833d07ca03dfd034cc88d9ea43571572b14c81be59e4971d

Observation a59ec1e8-3887-4873-9d5f-81906eb6e7b0 · outbound

This paper cites Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov.

The Science of Evaluating Foundation Models Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.724900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.724900Z digest=sha256:720614b2e31da9b8b9d81648e396a9c48547a6460540740be182defd232a612d

Observation 075f5f71-150d-420a-a730-4a4277b33180 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

The Science of Evaluating Foundation Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.727917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.727917Z digest=sha256:e6e815b74b3f3e7349b5932f28a8d8e3cc08c6b28f903947f4f62dfadf3a15dc

Observation a619e8b3-ff6f-4ee2-8d17-20b15e5522da · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.731195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.731195Z digest=sha256:148397d0b1b7075372a27f86e69c28bb87db8aeed234805adffc4366af7da5ec

Observation 65441e79-c127-4d2f-b3c1-c41653c992cc · outbound

This paper cites Evaluating Human-Language Model Interaction.

The Science of Evaluating Foundation Models Evaluating Human-Language Model Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.734153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.734153Z digest=sha256:0cd4d198c492c8c1d94c8cd608ea348f2d47295878b1ce03c37666ec3246fac6

Observation 6be26943-f8c2-421c-884b-dba5cd61f4c9 · outbound

This paper cites Can Large Language Models Capture Dissenting Human Voices?.

The Science of Evaluating Foundation Models Can Large Language Models Capture Dissenting Human Voices?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.737130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.737130Z digest=sha256:f8a408ab256b3b68c0c3547300762d8ba08dc19e4708f80163c7cb32f9be573e

Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.740185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.740185Z digest=sha256:0fb5357067e86716768bbb33bb8283b50405a21e56115fde92a55c1a6e4e28b3

Observation 2a33283f-1f8c-4b73-8fa4-c8bc1bec5053 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.743247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.743247Z digest=sha256:76ecdceb4987e7758ef4346c40b75be93e2a3565deeb2182c663907c0b487788

Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · outbound

This paper cites Holistic Evaluation of Language Models.

The Science of Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.752311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.752311Z digest=sha256:4f6d44eedbd2be03ef07da05b17b461db1271f92a6a778a0d906260cf02584d5

Observation f4539750-84aa-460b-904f-863ae408ad74 · outbound

This paper cites AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents.

The Science of Evaluating Foundation Models AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.385959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.755226Z digest=sha256:820719aef79c7358db15f15e79094ae3c7a7f09f783f2b798e68f7ee4773f37c

Observation f9d79ee3-df9a-4330-8066-1f8c920c6163 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.758625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.758625Z digest=sha256:a425c80181986c8fee2adb976596e72cdf8b1a73571d1d06030f5d66b89554af

Observation 6338a6ff-677a-4838-86d6-58c4a007de11 · outbound

This paper cites Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models.

The Science of Evaluating Foundation Models Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.761353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.761353Z digest=sha256:848157adaf62f5dd7704e75fc86984d055218826d7748966986c8ec59daa4fb0

Observation 358695cd-1638-4b11-81ea-3e5ca8383f1b · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.764319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.764319Z digest=sha256:6a12071f721c6b6ce5b1ff8b97aee7d8d748c6b2e609b2b8ee128505de41cf91

Observation 3620ce70-def8-4509-b833-ea60b5484b25 · outbound

This paper cites Social Bias Probing: Fairness Benchmarking for Language Models.

The Science of Evaluating Foundation Models Social Bias Probing: Fairness Benchmarking for Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.767138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.767138Z digest=sha256:a40d12e493bc06c7963e48e7a14dcdda46d29600bd012a4ef8984dd5b299a1c4

Observation 9ad605d3-0278-4cc8-9617-9aee498174d5 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.770006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.770006Z digest=sha256:288d87bebdd44a7ac92ed4ab1cbce6a0f335d1753f1ee9888ab2e3ee9034da4b

Observation c116ac70-da36-4744-993d-1caae578a27c · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

The Science of Evaluating Foundation Models StereoSet: Measuring stereotypical bias in pretrained language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.774783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.774783Z digest=sha256:5616a2a6ba8a2c33b97922d28e331711c30a51e7d27e627e9d35de228c72d22c

Observation e110ec71-f37a-4eee-bcae-08209c44724d · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.777377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.777377Z digest=sha256:b32785f2e4082430f2d28f45ca7da84e0174166268338993409438d91059f0b5

Observation d9a87917-5e9f-4fd5-9f10-1e37eab5d1a8 · outbound

This paper cites Pointer Sentinel Mixture Models.

The Science of Evaluating Foundation Models Pointer Sentinel Mixture Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.772371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.772371Z digest=sha256:f0a41110bc0b68f3cfa3fd26ee19bb81cd31e5d96a5521a0710e3f6d1bd888f5

Observation b030956a-0cfb-4015-96ce-8a20373fa5e8 · outbound

This paper cites Cohen, and Mirella Lapata.

The Science of Evaluating Foundation Models Cohen, and Mirella Lapata

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.787601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.787601Z digest=sha256:0e0ed7d679c196a55a3741a20cf9fc84362116db637f9e78aaec4979b9734d24

Observation 1f715672-1eff-412e-8e4c-c29850f55252 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.790559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.790559Z digest=sha256:32802437aac5c27a7b974350d4e5e6b153ec2036e376f7d4dd0cd103a35a956e

Observation fe1fa383-b0af-4bd7-befa-bbc665696e72 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

The Science of Evaluating Foundation Models Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.779786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.779786Z digest=sha256:72c7abeda9a8ef2d769d8061e01c2c173a12178e5f55f6396531cbe9068a34b4

Observation 81653e57-c00c-4a4a-9b16-663d87670804 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.782282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.782282Z digest=sha256:58f8b2889488c1f9e101436c47d3959fd392f33928d440e3bc333ef7ad832561

Observation 90319e29-dc75-449c-87b7-3c6aa7c23604 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.802455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.802455Z digest=sha256:f653bdaf78f70601c2c5a002b605d71722fcf8c54b735465636275b4e20fa91f

Observation 98b069ff-6c98-47d4-8c32-b0b209f1ef16 · outbound

This paper cites Is ChatGPT a General-Purpose Natural Language Processing Task Solver?.

The Science of Evaluating Foundation Models Is ChatGPT a General-Purpose Natural Language Processing Task Solver?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.805153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.805153Z digest=sha256:e77bae0b932dc7301c8e47cbdea5eef1121a150572cb5a25e3ad5725a3294591

Observation bb7c04e8-9804-4129-8c8b-a96455d5984a · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.808366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.808366Z digest=sha256:f950f79977234f14076d1d2502e0d02e0be13e386a27a11d4194ff213e5689b1

Observation 2d65a7e5-c86c-47ff-9d8a-40cf58e7980a · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.793354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.793354Z digest=sha256:6320b3a565b2d9b1a3e222d53aeafb2225ded3d6ee747c2d7931278c6c086626

Observation 23278487-8ae8-4dd8-b04e-786112a023a7 · outbound

This paper cites Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

The Science of Evaluating Foundation Models Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.817579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.817579Z digest=sha256:80fd7c21c4c1a3aecbd271ed1ad5e42070f04fa7b25a16139aef5f06d17602b6

Observation b750a0c5-1e45-4cf6-bb69-966979345786 · outbound

This paper cites A Survey of Useful LLM Evaluation.

The Science of Evaluating Foundation Models A Survey of Useful LLM Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.799324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.799324Z digest=sha256:0ee8763b3b7e4ed552bf2af591ba1974729841cf6b2c69aab3167184a53f6531

Observation 2c0b10e6-86cd-480f-9519-3439a79e9530 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.823750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.823750Z digest=sha256:64e910a4b7bf4b1c1e8c7470b96f109e4950576598127b5ec72f8acd8dd54ed8

Observation a1f5741e-f9c0-45b3-a459-22c647d2a224 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 66

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.829679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.829679Z digest=sha256:19de2f16be0da62520699d818230651b9579e6ae20ba18e2ef5edb84a79ef2c0

Observation dec6e7cb-b5c9-48f1-8070-dfe70d40a234 · outbound

This paper cites Rush, Sumit Chopra, and Jason Weston.

The Science of Evaluating Foundation Models Rush, Sumit Chopra, and Jason Weston

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.832607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.832607Z digest=sha256:8627ba9f33eb838244c5bf3b913e7cf77778e3eb42b08fd90180309de1666766

Observation 13445a94-7ada-472e-be00-90bbbe398e95 · outbound

This paper cites Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition.

The Science of Evaluating Foundation Models Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.835466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.835466Z digest=sha256:3ce09de5e7eacdf2ede27df219bcb5b3bcf19c0b019355f2f0c412be31a73fbd

Observation 01e16b8e-c3f9-4db7-b028-afd108b7d7fe · outbound

This paper cites LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models.

The Science of Evaluating Foundation Models LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.814464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.814464Z digest=sha256:f452bbf482322b0e3a239c2c3d3e92f8d39f67d1f5eacf0c910fad3ccd70a90c

Observation 5f764ef2-c647-432d-b8c1-e4e3a3f9e7d0 · outbound

This paper cites An Interpretability Evaluation Benchmark for Pre-trained Language Models.

The Science of Evaluating Foundation Models An Interpretability Evaluation Benchmark for Pre-trained Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.259889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.841478Z digest=sha256:3a5e9fd67611dd9aa2998ba895f6f15cee9809de0e07104b26352bbe92989344

Observation 5e1c44a6-b37c-4e66-a242-719aa6949594 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

The Science of Evaluating Foundation Models SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.820557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.820557Z digest=sha256:d6c00660a9a94eee5986be89801ab7ab44b78ce026ab4d640a43d9a1010ad2af

Observation 781e5b33-a901-4679-8be3-d59c8d9d1e48 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.847446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.847446Z digest=sha256:cb0445c417b7288a94604cfb3c0dcec55c7fa04b44585b75ce8a7f487f086fc5

Observation 4bd856cc-13c9-4883-9f6b-0f6a181f49bb · outbound

This paper cites In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.).

The Science of Evaluating Foundation Models In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.)

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.826508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.826508Z digest=sha256:13db2b1b2a6064494fbcad318b1d5be119ab6d725ab9751a473a1971a2b71ed2

Observation 3a20dc31-a722-4aef-8497-83847d01c642 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.852413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.852413Z digest=sha256:d2591590406bb9ea3ff33685a8dd7bd1bedc83b8a49390c095f2413de4411a98

Observation 70765af2-18bd-473d-ab69-26c73a29e4ad · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.854802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.854802Z digest=sha256:6217df8f6446adb8dadba722d998bc37e84712b360cd92ee74c27c78246d6d72

Observation d7c248df-e0b5-46e8-a888-d41c432aec1b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

The Science of Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.856940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.856940Z digest=sha256:3b3b94b3eefabfd0c220770b4b4f16e05de2aed9547ac64781c80812840b3706

Observation eb1aeec2-bb7d-4074-a21e-c13535ff7ecc · outbound

This paper cites Evaluating Large Language Models with fmeval.

The Science of Evaluating Foundation Models Evaluating Large Language Models with fmeval

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.272176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.838370Z digest=sha256:44c9fe7e75554c4a7fff20c0eb4e93fb47e09b02e5c183189a47afd1131fa5c0

Observation e5469b24-71e6-4831-9b20-84825d9d63c7 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

The Science of Evaluating Foundation Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.863898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.863898Z digest=sha256:e482c4d0ff0b5809b511265ef8a187e0281bdd7551ab283ed1f2e0602c7911dd

Observation 6bd40893-61f8-4de3-ae64-01645250aaba · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.844395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.844395Z digest=sha256:c3ad0f0b2f96199d0dd870b87d42141e3fca7a42e271675525686d8138cd7b3a

Observation 0955883b-c37d-4f8a-a42f-19e68cc01e40 · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

The Science of Evaluating Foundation Models DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.869827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.869827Z digest=sha256:d399793517691aad6b391c0877be5d0df2b40748a24ba42f1102896b2bd9af3c

Observation 9dab2ad0-c7f6-4b39-8ac3-479ce16c67f5 · outbound

This paper cites LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models.

The Science of Evaluating Foundation Models LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.246503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.849901Z digest=sha256:16f37195199346b2c08af79010c602d6303893740db2ecaa4126b06aebae0560

Observation 3e92f87e-2686-455f-83f1-4bca005cd247 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

The Science of Evaluating Foundation Models AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.876049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.876049Z digest=sha256:d37eb2f6d683775482af05fc42a574e3a22ef2919688c3ac1e3420c7f52bd8ce

Observation b31ecd82-ff6f-4bce-8643-30eec429bcf5 · outbound

This paper cites Document-Level Machine Translation with Large Language Models.

The Science of Evaluating Foundation Models Document-Level Machine Translation with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.879026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.879026Z digest=sha256:f9ba647913f07ae942c347df0470092c9340374a40275af8d57d05fe978b58ec

Observation 58fd2708-310c-406b-a025-44a2c6c04a52 · outbound

This paper cites Smith, and Teruko Mitamura.

The Science of Evaluating Foundation Models Smith, and Teruko Mitamura

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.882205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.882205Z digest=sha256:b9d8f1ce36dfa3f5cfa9b10df08bc2d54b31ccd16e9ed6d7e47e2a9702158e95

Observation 5df30076-7d26-4ce6-8d0c-9d2f199cec50 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.859441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.859441Z digest=sha256:d10a39fc7ec079d8ffaf9e34bc92b1f8a30ddbb5715a7416aafedf0063fe21ae

Observation 6ce5459a-8839-4d48-9da7-b3cf16f2b55f · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

The Science of Evaluating Foundation Models Advances in Neural Information Processing Systems 36 (2024)

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.861678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.861678Z digest=sha256:96c6fa703d979bc79090c416c245cb454814dbb99c45b426d4c89b70d16da6da

Observation ba2eda73-8fe5-4972-a8bd-8add8a617f16 · outbound

This paper cites DHP Benchmark: Are LLMs Good NLG Evaluators?.

The Science of Evaluating Foundation Models DHP Benchmark: Are LLMs Good NLG Evaluators?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.891112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.891112Z digest=sha256:2a7e720fa791010c83a714d5eae16a9324f257fc993235567b5edaa848845ac6

Observation 43dc27af-db77-471d-9052-f793109c05f7 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.867035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.867035Z digest=sha256:a25b24570a7817bb817a7ee4350aca3d2598bf2414966993f80a12e2627a8fee

Observation 6fd23ec5-6389-4d9a-92ee-cd3f27240be0 · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference.

The Science of Evaluating Foundation Models A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.897234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.897234Z digest=sha256:9e5e7d1c5dbadcb2100376fef33baab00d2263d678b107b34259bc3100aff8a5

Observation 51a6cad7-7eab-4dbe-9466-5731b9e54fb6 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

The Science of Evaluating Foundation Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.872968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.872968Z digest=sha256:0a236d1d91ba79e58ab6cc0076178ad7e0e23e92b89f723cffa038892be42239

Observation 6b7b782d-e3d9-49aa-8842-b8816314797e · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.598649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.903121Z digest=sha256:e4ecc7924f08c2ac19028c7b71e0a907ea346806d16ce93aea6e773585b11bdb

Observation 10a31815-2774-45dd-9871-f3e508c4e89f · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.589935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.909203Z digest=sha256:4a2cabe345d3943cdefd5a0d8973e095c260d45068907746ab58da951b771060

Observation 2639926f-7a1a-44a4-bd2a-1327389db7fa · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

The Science of Evaluating Foundation Models R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.912130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.912130Z digest=sha256:04c012519ad1c7f462887284e537aa344e1ccc8e582c1a65f20c7ed55d04e14e

Observation a40c18e0-4c67-4b32-bd52-1fb541fbcbcf · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.885080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.885080Z digest=sha256:fe056b2b7097d5fc8a5b472258af249322b797294a9f644675659adf65a7304a

Observation ffce3c16-92b4-40d9-a5a9-e9114ee65dc5 · outbound

This paper cites A Theoretical Analysis of NDCG Type Ranking Measures.

The Science of Evaluating Foundation Models A Theoretical Analysis of NDCG Type Ranking Measures

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.887915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.887915Z digest=sha256:b1688c8a5cff4faf4a146899158a9c5ddfb1659fd7d0623d27aaecc147865b7e

Observation 8a072e5a-e6ea-4d15-8dcd-93e24a914f35 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.581098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.922646Z digest=sha256:10072d12fc24cc1ad557324390a75d5918529f50592d015ddc353edbfdb35fe9

Observation 9ea7a298-1b5d-4e66-81fa-4c2d9a97de88 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

The Science of Evaluating Foundation Models Aligning Large Language Models with Human: A Survey

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.894068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.894068Z digest=sha256:2401fa9d31fb39c36645b04ebb1287ac68a645dff9e24b36610e1c2795fd77ad

Observation defb6ea1-6aed-4bd9-b8e7-c2bba422d85b · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

The Science of Evaluating Foundation Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.927960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.927960Z digest=sha256:d4c052fa724ca47200b131a6997d80b42ba436f97f347dedbefd1f4f0696340e

Observation 3a059a11-00b6-45aa-9180-3e42089465e4 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.606460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:35:42.900246Z digest=sha256:b8b4033a1929df0d49d1aef986e77955a2fc809c62bf459f4405817a79e48ec9

Observation 4c8835d8-bdec-40c5-981d-85d628bf4a3b · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

The Science of Evaluating Foundation Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.932998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.932998Z digest=sha256:a49a12b7a088a022cfcee59e51a6e74f915abcb476154411ed73e77e82aa9496

Observation 32c74742-06d8-4d11-b2cf-0bad6dd002e1 · outbound

This paper cites LLM as a System Service on Mobile Devices.

The Science of Evaluating Foundation Models LLM as a System Service on Mobile Devices

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.905965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.905965Z digest=sha256:4912b3b208aa10d814953cde2d73f9a513fbe1dcba91b56debbc60f2bb8ad321

Pith citing papers

Observation 2fabdbc0-8037-4596-bd31-d827de9ce9cf · inbound

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch cites this paper.

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch The Science of Evaluating Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:00.294806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:59:00.294806Z digest=sha256:79c60c4568334fa912d3781bb674ea101d9135599ddfaca9a79163db1da0696a

Observation df984a59-0850-4fc8-88b0-ca88a2f458de · inbound

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models cites this paper.

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models The Science of Evaluating Foundation Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.412843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:31:46.940449Z digest=sha256:ae312337506ae901ae2bfe24c47510c2158714f06981c00f396eece00e9d0439

Observation 817c9b0d-9628-423d-9255-866ac54a79ff · inbound

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis cites this paper.

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis The Science of Evaluating Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T20:50:53.567415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:50:53.567415Z digest=sha256:d201efadd718c1e1e24a577c686dbb5ba7149b937a4c1092846ef096e7978944