Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models: A Comprehensive Survey

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2310.19736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19736 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:35:47.201977Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.498619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f25dd113-06c5-482c-85d7-af39db2e81b7 · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.777329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:08f5f457b1f390028680df50c14fd3597b1aa46a57c2e83c85409fdc764d2c10

Observation fbb7778b-83db-4141-b825-2edcf7f93bfb · inbound

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning cites this paper.

Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:03.151597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:03.151597Z digest=sha256:ced2b41379d662ac4bc59742f26aa54158780adcbf0595c5ac69cd4e2c5ee17c

Observation bab42ba8-c35d-4cc3-8fd7-51d0647ecae7 · inbound

Evaluating Language Models as Synthetic Data Generators cites this paper.

Evaluating Language Models as Synthetic Data Generators Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:17:42.625582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:17:42.625582Z digest=sha256:fb26fd059a1cc951629ff12d010ff0e14778c21ff17731a7e65b602d1c50cc40

Observation 740a16bc-4999-489c-a844-8ecd51b1b38a · inbound

Explingo: Explaining AI Predictions using Large Language Models cites this paper.

Explingo: Explaining AI Predictions using Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:52:55.207821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:52:55.207821Z digest=sha256:0a7219c1e5ea713ef2af9f429f64664e48ead234814a528af52a7ef8872c5548

Observation 9e34b787-346b-4e4c-b9b9-2ffb293d6fa7 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Evaluating Large Language Models: A Comprehensive Survey

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.614193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:3462556aef3b33b0edcded73130be04ff83c8b8a9cdfcb5b895cf63bba1fae08

Observation eb47dc57-1fd4-4f9c-b3c6-e33711b03902 · inbound

How to Choose a Threshold for an Evaluation Metric for Large Language Models cites this paper.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.594727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.594727Z digest=sha256:45f5e5d56963d9a5cacf38bdbd2ff1dc2bc2d459f31b1f274cf418a7c676d358

Observation cebd7e80-5b9d-468f-adfc-b00a36abe7e8 · inbound

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems cites this paper.

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems Evaluating Large Language Models: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:22:56.682754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:22:56.682754Z digest=sha256:bf443b214607d43b06931e48aff6d19900cb8188b43b6fd008ef3c83b048b2b1

Observation 8ee7822b-61b3-4ba6-9fac-f8aef6d6887f · inbound

Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts cites this paper.

Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts Evaluating Large Language Models: A Comprehensive Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:40:01.628248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:40:01.628248Z digest=sha256:0bf879bcb06e6268d32f5b24a59b9da3f4b2cde19f069d5358083df42658ade6

Observation 459a184b-ab83-4925-b692-784313e395a0 · inbound

Validation of GPU Computation in Decentralized, Trustless Networks cites this paper.

Validation of GPU Computation in Decentralized, Trustless Networks Evaluating Large Language Models: A Comprehensive Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:40.201341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:17:40.201341Z digest=sha256:b6f2e0a34a62b32f2a237629157cd3b6a7c8fd1495aeb95034e031b530003b85

Observation c5cee3cf-8b6c-408a-bb3a-5ad66676cb23 · inbound

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong cites this paper.

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:37:36.032998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T05:37:22.955895Z digest=sha256:c0be905b3deec433a926c92e3376b006c3266be0ba256d2444fc51e80f23407c

Observation e78a7fc0-5b4e-4db2-b008-fa580cf2cbbb · inbound

Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework cites this paper.

Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:39.679430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:39.679430Z digest=sha256:14ec57d0a0e69d4184f31406883f4326a12d7c1c32a002994a11acf5a0ef47be

Observation a24a0a9f-4ac6-4fea-9d7a-f8c312a689e8 · inbound

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact cites this paper.

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact Evaluating Large Language Models: A Comprehensive Survey

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-08T05:35:36.355786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:35:36.355786Z digest=sha256:269c053a404e73d8f858dd82df282c8d4e30e569e857a0d1b20c8d014379aef2

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:e4eaab1ba27281eb032c0c1b19d4e4c615e384133d61153a5adbfc7d9079e38f

Observation 510c1041-1a1c-4bd7-8645-dfa5d220573b · inbound

Benchmarking LLM-based Relevance Judgment Methods cites this paper.

Benchmarking LLM-based Relevance Judgment Methods Evaluating Large Language Models: A Comprehensive Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:47.201977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:47.201977Z digest=sha256:7991bfe2a9f3c233c8180c562c81005fcdeeef6f62bdf1a344a2709638c1944b

Observation 50a6b8f2-afa1-4c46-ad3b-5d24bccc4d6a · inbound

The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks cites this paper.

The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks Evaluating Large Language Models: A Comprehensive Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:52.857915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:52.857915Z digest=sha256:17677beab67289b2bdda81c69a68098f89fa0dff0cde3bcda5ffdfdf0e522f30

Observation 320b609c-1579-4ee3-afd3-ad55c33d9230 · inbound

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments cites this paper.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Evaluating Large Language Models: A Comprehensive Survey

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.776501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.776501Z digest=sha256:53cbfe23b9e2268a5e39be2a12943f68b78e8ea32cfbbab0e0392a39fa420b20

Observation b7765eaf-25a3-4ae2-971b-30f6480a8f1a · inbound

Reasoning Capabilities and Invariability of Large Language Models cites this paper.

Reasoning Capabilities and Invariability of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:57.711984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:57.711984Z digest=sha256:1acd8f32333b39c6e0f3b60d3c4d93d2882178c5860f19b205e0af66e45114d7

Observation d243a491-2a7e-4761-872b-768ccad37ed5 · inbound

Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets cites this paper.

Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets Evaluating Large Language Models: A Comprehensive Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T01:09:03.149645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T01:09:03.149645Z digest=sha256:5add36b4c9bc7599e2e2454faeab0c79c5cd41a8d59013cc178fc48ff05c54d6

Observation 601f26b7-3277-4b41-89f3-f6ef3b87d70c · inbound

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks cites this paper.

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks Evaluating Large Language Models: A Comprehensive Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:55:11.811419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:55:11.811419Z digest=sha256:abaa0f9c151b8a64c1e6fb9ec07045e7d3b8310a258358f3c7854aeca7ef7c43

Observation b8be67ff-7b22-4515-8609-98a2f1cfd4aa · inbound

A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges cites this paper.

A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges Evaluating Large Language Models: A Comprehensive Survey

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:24.541525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:24.541525Z digest=sha256:19ef15e3047161a44679203974f21dec700bf8fb5f3c38d183b8f846a9f1baa3

Observation c030bab5-6e7e-41c6-a4c0-6476e242bd4f · inbound

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models cites this paper.

Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:20:08.506938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:20:08.506938Z digest=sha256:39859efcb40bb2b0c12319c837c76aac59a2d611cd49995fba31c2d8127ae135

Observation d0045ef0-0397-4995-8aa2-f362b4d487ed · inbound

Semantic Retention and Extreme Compression in LLMs: Can We Have Both? cites this paper.

Semantic Retention and Extreme Compression in LLMs: Can We Have Both? Evaluating Large Language Models: A Comprehensive Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:25:56.079760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:25:56.079760Z digest=sha256:8c8f4dc4665601c2c7b506880e803c35a4f57911e90b83cb6e4b4c196c805eae

Observation 06f07d09-4a10-4598-9100-985f02458367 · inbound

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data cites this paper.

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data Evaluating Large Language Models: A Comprehensive Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:03.725406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:03.725406Z digest=sha256:d391477715aa7752a89a37108e603f775c7b41ccc85e437eaffd3d2964126014

Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · inbound

Research-Oriented Human-Centric Evaluation for Foundation Models cites this paper.

Research-Oriented Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.124209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.124209Z digest=sha256:ff6674b155cc692d0efdd0a8e5cf9f30caa23669d7efc1d06b1ea124aeb4a503

Observation ae2a5157-fa04-47d9-af21-e7a9e5c61a1a · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Evaluating Large Language Models: A Comprehensive Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.263848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.263848Z digest=sha256:b49e810566847e611f977a3f97fbc758a18d5db9134c447db70a0500cd56b6d3

Observation b6665b51-1be9-4974-b389-cd2db10e5faf · inbound

TeleEval-OS: Performance evaluations of large language models for operations scheduling cites this paper.

TeleEval-OS: Performance evaluations of large language models for operations scheduling Evaluating Large Language Models: A Comprehensive Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:49.893683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:49.893683Z digest=sha256:a5718875f0d7558dcae1bf0719022717fe7be9be38c67dfe2f615dbcd8de12bd

Observation d8a0c089-0661-4617-a4b5-db7137f29cea · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.624960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.624960Z digest=sha256:f4bd2419b1932392f76082474d20946db2a6c1c5c8a75f8807ebd6995fe327b5

Observation a6a852ba-7372-4622-9b0c-2306905712da · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.603597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.603597Z digest=sha256:883c61234f4d87fb3077ef4cdfe9f8ae50195e3103c7c8645646cf73d5795386

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · inbound

Benchmarking the Pedagogical Knowledge of Large Language Models cites this paper.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:9270f64d56648c3b271fe7a19c8a6c87ee65cdc488adee5ed9ce70db2fe1278d

Observation f3488155-84cf-43bb-8959-c2b068c75164 · inbound

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans cites this paper.

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:39.388006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:39.388006Z digest=sha256:668114afb9f8593c1a6163f61b9e02efb00f542787d2b7e1a71c84802a731e3e

Observation 360b283c-c81c-4884-a279-da2d7e445a4c · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Evaluating Large Language Models: A Comprehensive Survey

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:31.666043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:31.666043Z digest=sha256:deea9eb2542abcdc38d8dcd355b4c5649540bea812e9e7317948f90373529fde

Observation 528253e7-3801-47df-a54b-7e0035840e8d · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:34.962660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:34.962660Z digest=sha256:49a9be43a904006e359e6163eabc9bac91d9248b73789da52e880ca7a957722e

Observation 5be48bf6-ad10-44dc-9d1a-6adb070acf96 · inbound

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation cites this paper.

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:16.530237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:16.530237Z digest=sha256:4563ead3822994168cfa193d50c819375c4dcb23f0d8ed47a098566f44cecd05

Observation c4bcdf86-1881-40d4-9496-8652eec1d50a · inbound

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs cites this paper.

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:57.803410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:58:57.803410Z digest=sha256:8e7d34a898d006a44eb5b94d269e27c3e6320443e99b325706fd0c28dff38e40

Observation a755dffa-b085-4380-9230-e114f28a74e8 · inbound

Symbiotic Agents: A Novel Paradigm for Trustworthy AGI-driven Networks cites this paper.

Symbiotic Agents: A Novel Paradigm for Trustworthy AGI-driven Networks Evaluating Large Language Models: A Comprehensive Survey

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:23:46.763882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:23:46.763882Z digest=sha256:e15c1f93c9b4ba1f25ffcd94acd59eb43f3ec087ec643f59bf8f951ffd00fffa

Observation 1187c08b-4c1b-4e7e-bd62-b098dec4c48b · inbound

Cognitive Agents Powered by Large Language Models for Agile Software Project Management cites this paper.

Cognitive Agents Powered by Large Language Models for Agile Software Project Management Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:59:24.676630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:59:24.676630Z digest=sha256:27ce29d271ae27845f29b422acb0b64b42f58ad0adffc7d71a19fd87d5a98368

Observation ae6c1ce3-75fb-492e-a33a-e27906e55ab4 · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.813109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.813109Z digest=sha256:1fcf73628b915e4193ee7acbaa71a7711d260bf6929f0ae0eca85a4bc09fd994

Observation 0a6f8ab9-1100-42bc-a5db-c4eab8416ae6 · inbound

Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering cites this paper.

Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:59:56.238159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:59:56.238159Z digest=sha256:5e5a8ac9548dc241cd3b91e3db744e42a2f34871bce88213a5469a77620e5327

Observation 7bf470c6-62b3-40cb-8c0c-45de6b616bc7 · inbound

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment cites this paper.

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:10:18.500779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T21:09:29.669360Z digest=sha256:ebc2741955fdffd9f5a115f41af150800f9f6cc7b1eefda464ceeef764734738

Observation 9877ca53-ae81-4513-ad56-8ba2ef3c504e · inbound

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement cites this paper.

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Evaluating Large Language Models: A Comprehensive Survey

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.758821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:53:10.245859Z digest=sha256:97a2d2857a85127b32a58f3c3ebaebc1230911ebc3194fd85d36c2b5d9c98c42

Observation aa005585-e4db-472f-8abe-51cfaac554ca · inbound

The Generalized Turing Test: A Foundation for Comparing Intelligence cites this paper.

The Generalized Turing Test: A Foundation for Comparing Intelligence Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:26:18.923707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:25:48.145137Z digest=sha256:7f91af5ce1249ab0e5b65b78244e6efc1ffd81c4ec69a8feb108b73652aa45ab

Observation aebc6e71-8859-4a49-bc08-35f047881fe5 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.749420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T06:24:16.157548Z digest=sha256:0e9db69cc1886aca130d6bc32f8799db788c2292f8d8c31baf491a0934a87424

Observation dc48d998-456e-4a0f-8bf0-7edb31d9a7e4 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.695669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T05:44:52.491280Z digest=sha256:22452fbe6c81629c9dc494fd0c95e92981a6abe3b2678d5746d5d666d1a6c16f

Observation ea682319-30d1-4e82-9876-74b6fad82272 · inbound

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation cites this paper.

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation Evaluating Large Language Models: A Comprehensive Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:29:03.254380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T21:26:13.698051Z digest=sha256:fd38269dc35f8cab56b5037ce468d6485cdc794ef7c76f2d6df24fded28c87d0

Observation 111567ad-3fca-4fe1-a757-df8f0a1fddfe · inbound

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts cites this paper.

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts Evaluating Large Language Models: A Comprehensive Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.882307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T12:13:12.018506Z digest=sha256:415a8300f87ba285e01996a69b417827c9546208bf90fb385b035b01ea913171

Observation 99f47601-be53-42f7-8a80-d43eb53c6843 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.500046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:bf4c4b1fa9356a2290b3525bb588b6daaebf4180b927fc24338bab52e1c5393d

Observation edd30526-0248-4d70-8c0d-f389042da49e · inbound

Efficient Sequential Evaluation of Large Language Models cites this paper.

Efficient Sequential Evaluation of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:08:18.676413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:08:18.676413Z digest=sha256:fdd513af5263c0d15a468ecbfc4f57ec097b5c97ac70d985bb1b3fa86a748830

Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.521574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.521574Z digest=sha256:556cab57c142dbc3a02e249d1afdc876e7d9641067ac24bf6f2c8a6b4819e6b0

Observation 30c07076-a48b-46a7-b6dc-857b4affaf53 · inbound

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information cites this paper.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Evaluating Large Language Models: A Comprehensive Survey

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.713700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.713700Z digest=sha256:de1d1dbf9d83347f37de28be3be7369d0219f7a6fd1f97914d8282915666265a