Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2310.19736.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:35:47.201977Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:49:41.498619Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f25dd113-06c5-482c-85d7-af39db2e81b7 · inbound
A Survey on the Memory Mechanism of Large Language Model based Agents Evaluating Large Language Models: A Comprehensive Survey
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fbb7778b-83db-4141-b825-2edcf7f93bfb · inbound
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab42ba8-c35d-4cc3-8fd7-51d0647ecae7 · inbound
Evaluating Language Models as Synthetic Data Generators Evaluating Large Language Models: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740a16bc-4999-489c-a844-8ecd51b1b38a · inbound
Explingo: Explaining AI Predictions using Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e34b787-346b-4e4c-b9b9-2ffb293d6fa7 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Evaluating Large Language Models: A Comprehensive Survey
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eb47dc57-1fd4-4f9c-b3c6-e33711b03902 · inbound
How to Choose a Threshold for an Evaluation Metric for Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebd7e80-5b9d-468f-adfc-b00a36abe7e8 · inbound
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems Evaluating Large Language Models: A Comprehensive Survey
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee7822b-61b3-4ba6-9fac-f8aef6d6887f · inbound
Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts Evaluating Large Language Models: A Comprehensive Survey
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459a184b-ab83-4925-b692-784313e395a0 · inbound
Validation of GPU Computation in Decentralized, Trustless Networks Evaluating Large Language Models: A Comprehensive Survey
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5cee3cf-8b6c-408a-bb3a-5ad66676cb23 · inbound
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Evaluating Large Language Models: A Comprehensive Survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e78a7fc0-5b4e-4db2-b008-fa580cf2cbbb · inbound
Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework Evaluating Large Language Models: A Comprehensive Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a24a0a9f-4ac6-4fea-9d7a-f8c312a689e8 · inbound
Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact Evaluating Large Language Models: A Comprehensive Survey
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · inbound
The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510c1041-1a1c-4bd7-8645-dfa5d220573b · inbound
Benchmarking LLM-based Relevance Judgment Methods Evaluating Large Language Models: A Comprehensive Survey
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a6b8f2-afa1-4c46-ad3b-5d24bccc4d6a · inbound
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks Evaluating Large Language Models: A Comprehensive Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320b609c-1579-4ee3-afd3-ad55c33d9230 · inbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Evaluating Large Language Models: A Comprehensive Survey
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7765eaf-25a3-4ae2-971b-30f6480a8f1a · inbound
Reasoning Capabilities and Invariability of Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d243a491-2a7e-4761-872b-768ccad37ed5 · inbound
Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets Evaluating Large Language Models: A Comprehensive Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601f26b7-3277-4b41-89f3-f6ef3b87d70c · inbound
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks Evaluating Large Language Models: A Comprehensive Survey
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8be67ff-7b22-4515-8609-98a2f1cfd4aa · inbound
A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges Evaluating Large Language Models: A Comprehensive Survey
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c030bab5-6e7e-41c6-a4c0-6476e242bd4f · inbound
Enigme: Generative Text Puzzles for Evaluating Reasoning in Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0045ef0-0397-4995-8aa2-f362b4d487ed · inbound
Semantic Retention and Extreme Compression in LLMs: Can We Have Both? Evaluating Large Language Models: A Comprehensive Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f07d09-4a10-4598-9100-985f02458367 · inbound
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data Evaluating Large Language Models: A Comprehensive Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · inbound
Research-Oriented Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2a5157-fa04-47d9-af21-e7a9e5c61a1a · inbound
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Evaluating Large Language Models: A Comprehensive Survey
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6665b51-1be9-4974-b389-cd2db10e5faf · inbound
TeleEval-OS: Performance evaluations of large language models for operations scheduling Evaluating Large Language Models: A Comprehensive Survey
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a0c089-0661-4617-a4b5-db7137f29cea · inbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a852ba-7372-4622-9b0c-2306905712da · inbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey Evaluating Large Language Models: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · inbound
Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3488155-84cf-43bb-8959-c2b068c75164 · inbound
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360b283c-c81c-4884-a279-da2d7e445a4c · inbound
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Evaluating Large Language Models: A Comprehensive Survey
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 528253e7-3801-47df-a54b-7e0035840e8d · inbound
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be48bf6-ad10-44dc-9d1a-6adb070acf96 · inbound
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation Evaluating Large Language Models: A Comprehensive Survey
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bcdf86-1881-40d4-9496-8652eec1d50a · inbound
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a755dffa-b085-4380-9230-e114f28a74e8 · inbound
Symbiotic Agents: A Novel Paradigm for Trustworthy AGI-driven Networks Evaluating Large Language Models: A Comprehensive Survey
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1187c08b-4c1b-4e7e-bd62-b098dec4c48b · inbound
Cognitive Agents Powered by Large Language Models for Agile Software Project Management Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6c1ce3-75fb-492e-a33a-e27906e55ab4 · inbound
Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a6f8ab9-1100-42bc-a5db-c4eab8416ae6 · inbound
Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering Evaluating Large Language Models: A Comprehensive Survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bf470c6-62b3-40cb-8c0c-45de6b616bc7 · inbound
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment Evaluating Large Language Models: A Comprehensive Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9877ca53-ae81-4513-ad56-8ba2ef3c504e · inbound
Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Evaluating Large Language Models: A Comprehensive Survey
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aa005585-e4db-472f-8abe-51cfaac554ca · inbound
The Generalized Turing Test: A Foundation for Comparing Intelligence Evaluating Large Language Models: A Comprehensive Survey
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aebc6e71-8859-4a49-bc08-35f047881fe5 · inbound
Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc48d998-456e-4a0f-8bf0-7edb31d9a7e4 · inbound
Do Language Models Encode Knowledge of Linguistic Constraint Violations? Evaluating Large Language Models: A Comprehensive Survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ea682319-30d1-4e82-9876-74b6fad82272 · inbound
A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation Evaluating Large Language Models: A Comprehensive Survey
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 111567ad-3fca-4fe1-a757-df8f0a1fddfe · inbound
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts Evaluating Large Language Models: A Comprehensive Survey
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 99f47601-be53-42f7-8a80-d43eb53c6843 · inbound
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Evaluating Large Language Models: A Comprehensive Survey
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation edd30526-0248-4d70-8c0d-f389042da49e · inbound
Efficient Sequential Evaluation of Large Language Models Evaluating Large Language Models: A Comprehensive Survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c6c9d6-cbec-4b73-bb34-fc4e24f7b3d7 · inbound
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Evaluating Large Language Models: A Comprehensive Survey
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c07076-a48b-46a7-b6dc-857b4affaf53 · inbound
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Evaluating Large Language Models: A Comprehensive Survey
Reference 164
Source-reported events for the cited work
Unavailable: canonical work link unavailable.