Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models: A Comprehensive Survey

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2310.19736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19736 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:35:36.355786Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.498619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f25dd113-06c5-482c-85d7-af39db2e81b7 · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.777329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:f2f7f035dc124022deab2d2fd15b27d2cbebdfa7a06c72113856d1b6a9b1e0cd

Observation 9e34b787-346b-4e4c-b9b9-2ffb293d6fa7 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Evaluating Large Language Models: A Comprehensive Survey

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.614193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:f224e969b8738bc0931421c0e4c32bc7e35490fbe28cab05b885cd4fc329bf0c

Observation c5cee3cf-8b6c-408a-bb3a-5ad66676cb23 · inbound

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong cites this paper.

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:37:36.032998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T05:37:22.955895Z digest=sha256:807e04dbc933018668b415c57ee7b0edcc46dcb04940b52b238cef1aa78acf4b

Observation a24a0a9f-4ac6-4fea-9d7a-f8c312a689e8 · inbound

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact cites this paper.

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact Evaluating Large Language Models: A Comprehensive Survey

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-08T05:35:36.355786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:35:36.355786Z digest=sha256:bb7408552ab63e58769fbc2eddc25413bfa4bf8af0955bbdb7ecc49c139b0a61

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:96a4617e016965e283bcb7378f63fce9dba7afaff713ee38e842be946b723089

Observation 06f07d09-4a10-4598-9100-985f02458367 · inbound

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data cites this paper.

From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data Evaluating Large Language Models: A Comprehensive Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:03.725406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:03.725406Z digest=sha256:2280c5b403ec108ac7bb5dbd5ab1ab940ed64b0fb39c18e7bc7e132bdb269441

Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · inbound

Human-Centric Evaluation for Foundation Models cites this paper.

Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.124209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.124209Z digest=sha256:95fabcbfa5868c32724f754f57c2ec800eb866993cfd8110f63832a9d601404a

Observation ae2a5157-fa04-47d9-af21-e7a9e5c61a1a · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Evaluating Large Language Models: A Comprehensive Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.263848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.263848Z digest=sha256:9ebe699fe2de94bc7c8355a807217f5115bf50f47acc0a3b3b1c9d3ae18f781f

Observation d8a0c089-0661-4617-a4b5-db7137f29cea · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.624960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.624960Z digest=sha256:374c088f9624fdc44f85b8bd7e7c97b8e91a230181bacd2d01a1d1a4ef706a38

Observation a6a852ba-7372-4622-9b0c-2306905712da · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:16.603597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:16.603597Z digest=sha256:8b20b0cd20b0db2917186434b6ce75ae9e354c5038800b5d9b64a3c2b901b093

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · inbound

Benchmarking the Pedagogical Knowledge of Large Language Models cites this paper.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:082c7d9ba5408d8b3251b9a0917f4aaad32fa3deffeb79196ca11dddadeeadfc

Observation f3488155-84cf-43bb-8959-c2b068c75164 · inbound

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans cites this paper.

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:39.388006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:41:39.388006Z digest=sha256:525bac89a87643ea8d20f55c3b7feef60dec5d996cc93c264c1359dfe60868fc

Observation 360b283c-c81c-4884-a279-da2d7e445a4c · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Evaluating Large Language Models: A Comprehensive Survey

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:31.666043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:31.666043Z digest=sha256:e3d6376805bc2c8c1b32f38be075ca7d1545ed9cc0eadb2a517b4e79670dc1e3

Observation 528253e7-3801-47df-a54b-7e0035840e8d · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:34.962660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:34.962660Z digest=sha256:38a340ab433f3c38ef86ac90cacfc9f9dd4f0b34ca7d72b18a487077943f79fa

Observation 5be48bf6-ad10-44dc-9d1a-6adb070acf96 · inbound

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation cites this paper.

OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:16.530237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:16.530237Z digest=sha256:4a6c6e874876f2d7e4a5c1c5a0da526cd5e5fbd5b32498869c374c8da685f9bd

Observation c4bcdf86-1881-40d4-9496-8652eec1d50a · inbound

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs cites this paper.

SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:57.803410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:58:57.803410Z digest=sha256:987836632e58330f6dcfe8f5fe6536f2bfd281caefa996125966cc74abcd03f6

Observation 1187c08b-4c1b-4e7e-bd62-b098dec4c48b · inbound

Cognitive Agents Powered by Large Language Models for Agile Software Project Management cites this paper.

Cognitive Agents Powered by Large Language Models for Agile Software Project Management Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:59:24.676630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:59:24.676630Z digest=sha256:50b370ca377ecfd87d045f2b77129394d329b21833c183718db2156df250e7b3

Observation ae6c1ce3-75fb-492e-a33a-e27906e55ab4 · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.813109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.813109Z digest=sha256:19cb652640954b92618d677fe65013da16e13bcfef9b3c788182874dbc8f51b9

Observation 7bf470c6-62b3-40cb-8c0c-45de6b616bc7 · inbound

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment cites this paper.

When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment Evaluating Large Language Models: A Comprehensive Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:10:18.500779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:09:29.669360Z digest=sha256:f120a96efcf4f003b49e410d0997a0de119c61f34fef3689bd3ed218705469c2

Observation 9877ca53-ae81-4513-ad56-8ba2ef3c504e · inbound

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement cites this paper.

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Evaluating Large Language Models: A Comprehensive Survey

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.758821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:53:10.245859Z digest=sha256:4b6f75e2191d1ebe32a27db229fdd7fec78cda13556ac8b3ea7a9223a28aba7c

Observation aa005585-e4db-472f-8abe-51cfaac554ca · inbound

The Generalized Turing Test: A Foundation for Comparing Intelligence cites this paper.

The Generalized Turing Test: A Foundation for Comparing Intelligence Evaluating Large Language Models: A Comprehensive Survey

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:26:18.923707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:25:48.145137Z digest=sha256:ae74fead2923f5e8337849688cde0e20b71918b34cb7527e9fa37823e76afc24

Observation aebc6e71-8859-4a49-bc08-35f047881fe5 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.749420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:24:16.157548Z digest=sha256:631cdbff4162d5d828ce31ed17c46842511f3502625141d3026b512ed553d2fc

Observation dc48d998-456e-4a0f-8bf0-7edb31d9a7e4 · inbound

Do Language Models Encode Knowledge of Linguistic Constraint Violations? cites this paper.

Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.695669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:44:52.491280Z digest=sha256:814457073808a475e66640fa1965e0c06cf860195e5195a9adbc96a10f838c1c

Observation ea682319-30d1-4e82-9876-74b6fad82272 · inbound

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation cites this paper.

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation Evaluating Large Language Models: A Comprehensive Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:29:03.254380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T21:26:13.698051Z digest=sha256:ad61f5435fd525462e2de66f07667b468797b93a7ab1358f8bc82a311702e9f0

Observation 111567ad-3fca-4fe1-a757-df8f0a1fddfe · inbound

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts cites this paper.

Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts Evaluating Large Language Models: A Comprehensive Survey

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.882307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T12:13:12.018506Z digest=sha256:c91c9bcdcad1d0469610cb238419633bd9173d755daabf1f46fa403cae5963a6

Observation 99f47601-be53-42f7-8a80-d43eb53c6843 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.500046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:e90d2177b01ea5843b5bf91ce11835846c23758f98d8f792f7d7c816c7cbfd6b

Observation edd30526-0248-4d70-8c0d-f389042da49e · inbound

Efficient Sequential Evaluation of Large Language Models cites this paper.

Efficient Sequential Evaluation of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:08:18.676413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:08:18.676413Z digest=sha256:c993b8028bbc205cb86d1a824d01b51f895ca5bafaa2686da2406b6d5b618b2b

Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:22.521574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:22.521574Z digest=sha256:52a76e58e27d2370863aa1dad0b71190551f04ec38fa647364da758419017200