Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.932998Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 3 inbound Pith citation observations for arXiv:2502.09670.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.932998Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:59:00.294806Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T13:35:26.411399Z
100 of 109 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 42c9f172-4f43-45bb-a54b-f79576fea6e9 · outbound
The Science of Evaluating Foundation Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a60ccf5-011b-40cb-a6bf-a6017ed74b60 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7afbaa-a360-4dd2-ad22-312df632e1e6 · outbound
The Science of Evaluating Foundation Models AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f75078f-c123-401a-a525-63783e7c7e83 · outbound
The Science of Evaluating Foundation Models Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7b9739-cbf3-4c4b-9f8b-764e250d0aae · outbound
The Science of Evaluating Foundation Models MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d987c0-95d4-4044-9d51-e53b33a6e710 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41511bae-4253-4804-b53a-48a73a38451c · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85ccb95-4a73-48d2-b1e6-a9281c26e127 · outbound
The Science of Evaluating Foundation Models A large annotated corpus for learning natural language inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5967fedb-279c-4f5d-b6e8-245121c56191 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a4d036-ebc9-4ecd-9edd-e3393beb80e4 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29df79b0-d079-4839-bd57-ba240b292bfa · outbound
The Science of Evaluating Foundation Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e87e0c8-0110-4424-b767-beb1980c8365 · outbound
The Science of Evaluating Foundation Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab6763d-8366-497c-8cf7-6fd8c8c54b69 · outbound
The Science of Evaluating Foundation Models Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5d6907-4908-4224-bf26-be7280651b14 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1596f007-4aa1-4066-9a88-8c96bc8e9316 · outbound
The Science of Evaluating Foundation Models RobustBench: a standardized adversarial robustness benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80be2df7-f961-48de-a6bf-b87100cf0ffb · outbound
The Science of Evaluating Foundation Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3537a145-0e41-4fe3-80e4-182d8abd2937 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b93dc5-4227-475a-9f92-ee852864aeb3 · outbound
The Science of Evaluating Foundation Models ERASER: A Benchmark to Evaluate Rationalized NLP Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5204879-ae56-4307-a2e4-b57e582ad2d3 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6fbcf9-4b9d-4bf2-856e-f4cb37c65970 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706442f5-54ec-4af2-82af-ac355f421310 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b512d5ff-cb96-4239-a582-af833ebb7998 · outbound
The Science of Evaluating Foundation Models An Intersectional Definition of Fairness
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66b685a-41b0-454b-b26f-954ce8367b2e · outbound
The Science of Evaluating Foundation Models Selective Classification for Deep Neural Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90609972-6f82-472e-8db8-3423a1fe6404 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fca449-93b7-4e35-9ca3-947cb0e00a98 · outbound
The Science of Evaluating Foundation Models On Calibration of Modern Neural Networks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed300267-aa21-4039-8308-acb5efe10507 · outbound
The Science of Evaluating Foundation Models Large Language Model based Multi-Agents: A Survey of Progress and Challenges
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · outbound
The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · outbound
The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b41117-de41-4783-9142-bb499db5a010 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a328410-f232-4698-99d7-280aaf548897 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f579ee1-1005-4ee9-8589-065189ae3929 · outbound
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d995213e-6350-442b-842a-42a32a09be59 · outbound
The Science of Evaluating Foundation Models Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53901a15-baae-434b-a914-f309e49e8e35 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972216ea-dae4-4d33-b25f-1cb12a26cb26 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552ef713-bd7d-483a-9c03-a35e71a76be5 · outbound
The Science of Evaluating Foundation Models WILDS: A Benchmark of in-the-Wild Distribution Shifts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad99ac7-c79a-42bf-90e9-c92e65216222 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59ec1e8-3887-4873-9d5f-81906eb6e7b0 · outbound
The Science of Evaluating Foundation Models Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075f5f71-150d-420a-a730-4a4277b33180 · outbound
The Science of Evaluating Foundation Models Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a619e8b3-ff6f-4ee2-8d17-20b15e5522da · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65441e79-c127-4d2f-b3c1-c41653c992cc · outbound
The Science of Evaluating Foundation Models Evaluating Human-Language Model Interaction
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6be26943-f8c2-421c-884b-dba5cd61f4c9 · outbound
The Science of Evaluating Foundation Models Can Large Language Models Capture Dissenting Human Voices?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · outbound
The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a33283f-1f8c-4b73-8fa4-c8bc1bec5053 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · outbound
The Science of Evaluating Foundation Models Holistic Evaluation of Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4539750-84aa-460b-904f-863ae408ad74 · outbound
The Science of Evaluating Foundation Models AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9d79ee3-df9a-4330-8066-1f8c920c6163 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6338a6ff-677a-4838-86d6-58c4a007de11 · outbound
The Science of Evaluating Foundation Models Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 358695cd-1638-4b11-81ea-3e5ca8383f1b · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3620ce70-def8-4509-b833-ea60b5484b25 · outbound
The Science of Evaluating Foundation Models Social Bias Probing: Fairness Benchmarking for Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad605d3-0278-4cc8-9617-9aee498174d5 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c116ac70-da36-4744-993d-1caae578a27c · outbound
The Science of Evaluating Foundation Models StereoSet: Measuring stereotypical bias in pretrained language models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e110ec71-f37a-4eee-bcae-08209c44724d · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a87917-5e9f-4fd5-9f10-1e37eab5d1a8 · outbound
The Science of Evaluating Foundation Models Pointer Sentinel Mixture Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b030956a-0cfb-4015-96ce-8a20373fa5e8 · outbound
The Science of Evaluating Foundation Models Cohen, and Mirella Lapata
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f715672-1eff-412e-8e4c-c29850f55252 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1fa383-b0af-4bd7-befa-bbc665696e72 · outbound
The Science of Evaluating Foundation Models Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81653e57-c00c-4a4a-9b16-663d87670804 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90319e29-dc75-449c-87b7-3c6aa7c23604 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b069ff-6c98-47d4-8c32-b0b209f1ef16 · outbound
The Science of Evaluating Foundation Models Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7c04e8-9804-4129-8c8b-a96455d5984a · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d65a7e5-c86c-47ff-9d8a-40cf58e7980a · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23278487-8ae8-4dd8-b04e-786112a023a7 · outbound
The Science of Evaluating Foundation Models Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b750a0c5-1e45-4cf6-bb69-966979345786 · outbound
The Science of Evaluating Foundation Models A Survey of Useful LLM Evaluation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0b10e6-86cd-480f-9519-3439a79e9530 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f5741e-f9c0-45b3-a459-22c647d2a224 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec6e7cb-b5c9-48f1-8070-dfe70d40a234 · outbound
The Science of Evaluating Foundation Models Rush, Sumit Chopra, and Jason Weston
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13445a94-7ada-472e-be00-90bbbe398e95 · outbound
The Science of Evaluating Foundation Models Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e16b8e-c3f9-4db7-b028-afd108b7d7fe · outbound
The Science of Evaluating Foundation Models LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f764ef2-c647-432d-b8c1-e4e3a3f9e7d0 · outbound
The Science of Evaluating Foundation Models An Interpretability Evaluation Benchmark for Pre-trained Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e1c44a6-b37c-4e66-a242-719aa6949594 · outbound
The Science of Evaluating Foundation Models SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781e5b33-a901-4679-8be3-d59c8d9d1e48 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd856cc-13c9-4883-9f6b-0f6a181f49bb · outbound
The Science of Evaluating Foundation Models In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.)
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a20dc31-a722-4aef-8497-83847d01c642 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70765af2-18bd-473d-ab69-26c73a29e4ad · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c248df-e0b5-46e8-a888-d41c432aec1b · outbound
The Science of Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb1aeec2-bb7d-4074-a21e-c13535ff7ecc · outbound
The Science of Evaluating Foundation Models Evaluating Large Language Models with fmeval
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5469b24-71e6-4831-9b20-84825d9d63c7 · outbound
The Science of Evaluating Foundation Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd40893-61f8-4de3-ae64-01645250aaba · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0955883b-c37d-4f8a-a42f-19e68cc01e40 · outbound
The Science of Evaluating Foundation Models DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dab2ad0-c7f6-4b39-8ac3-479ce16c67f5 · outbound
The Science of Evaluating Foundation Models LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e92f87e-2686-455f-83f1-4bca005cd247 · outbound
The Science of Evaluating Foundation Models AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31ecd82-ff6f-4bce-8643-30eec429bcf5 · outbound
The Science of Evaluating Foundation Models Document-Level Machine Translation with Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fd2708-310c-406b-a025-44a2c6c04a52 · outbound
The Science of Evaluating Foundation Models Smith, and Teruko Mitamura
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df30076-7d26-4ce6-8d0c-9d2f199cec50 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce5459a-8839-4d48-9da7-b3cf16f2b55f · outbound
The Science of Evaluating Foundation Models Advances in Neural Information Processing Systems 36 (2024)
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba2eda73-8fe5-4972-a8bd-8add8a617f16 · outbound
The Science of Evaluating Foundation Models DHP Benchmark: Are LLMs Good NLG Evaluators?
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43dc27af-db77-471d-9052-f793109c05f7 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd23ec5-6389-4d9a-92ee-cd3f27240be0 · outbound
The Science of Evaluating Foundation Models A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a6cad7-7eab-4dbe-9466-5731b9e54fb6 · outbound
The Science of Evaluating Foundation Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7b782d-e3d9-49aa-8842-b8816314797e · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10a31815-2774-45dd-9871-f3e508c4e89f · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2639926f-7a1a-44a4-bd2a-1327389db7fa · outbound
The Science of Evaluating Foundation Models R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40c18e0-4c67-4b32-bd52-1fb541fbcbcf · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffce3c16-92b4-40d9-a5a9-e9114ee65dc5 · outbound
The Science of Evaluating Foundation Models A Theoretical Analysis of NDCG Type Ranking Measures
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a072e5a-e6ea-4d15-8dcd-93e24a914f35 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ea7a298-1b5d-4e66-81fa-4c2d9a97de88 · outbound
The Science of Evaluating Foundation Models Aligning Large Language Models with Human: A Survey
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defb6ea1-6aed-4bd9-b8e7-c2bba422d85b · outbound
The Science of Evaluating Foundation Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a059a11-00b6-45aa-9180-3e42089465e4 · outbound
The Science of Evaluating Foundation Models Unresolved cited work
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c8835d8-bdec-40c5-981d-85d628bf4a3b · outbound
The Science of Evaluating Foundation Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c74742-06d8-4d11-b2cf-0bad6dd002e1 · outbound
The Science of Evaluating Foundation Models LLM as a System Service on Mobile Devices
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fabdbc0-8037-4596-bd31-d827de9ce9cf · inbound
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch The Science of Evaluating Foundation Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df984a59-0850-4fc8-88b0-ca88a2f458de · inbound
Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models The Science of Evaluating Foundation Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 817c9b0d-9628-423d-9255-866ac54a79ff · inbound
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis The Science of Evaluating Foundation Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.