Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2306.05087.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:45:45.853992Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
29
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation c7f2840b-0625-461b-b166-3e9f954e8a34 · inbound
Large Language Models are not Fair Evaluators PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8f028035-56a9-427c-97d9-e0cc4538cea5 · inbound
Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ca9376c1-50fb-4a7c-a3a3-1d12180566b2 · inbound
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5a615d9c-6817-456f-8d16-de2b6575f431 · inbound
Explainable LLM-driven Multi-dimensional Distillation for E-Commerce Relevance Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe5d0bf-c40b-462d-a63d-2e5ab29e633f · inbound
Natural Language Reinforcement Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a30c9f-f494-458c-9fd3-d7f26a3bae36 · inbound
A Survey on LLM-as-a-Judge PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0221db6b-c1de-4d69-a5e6-f1612d27294a · inbound
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f1b288-458c-4997-8b0d-d16089b18662 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 245
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bfede6a1-675d-4638-943a-23621287a16e · inbound
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4774d543-5df4-4e24-886f-d26bcb351cba · inbound
SedarEval: Automated Evaluation using Self-Adaptive Rubrics PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21bfd10f-c0a8-4abf-8fd7-24edbe2b6879 · inbound
Verifiable Format Control for Large Language Model Generations PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dd17b74-4b40-49ce-88b7-23af43d7aca5 · inbound
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6196b8d1-9e22-4a3a-9b08-9494653f8cda · inbound
PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0e2d8b-5bfd-4410-a804-db642cd30dc2 · inbound
An Empirical Study of Evaluating Long-form Question Answering PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8856035a-d684-4cf4-bcbf-bc7eb0087672 · inbound
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ca64449-a955-447c-9d9c-03e911f9d531 · inbound
am-ELO: A Stable Framework for Arena-based LLM Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a924e3-b953-46bb-b261-45d2f70bc1f7 · inbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e83056-c601-4f4b-9e4c-0b4f6bc2ea90 · inbound
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5947a4-37a9-463a-8ab1-80b4140c2f43 · inbound
LLM-based Evaluation Policy Extraction for Ecological Modeling PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce379c9d-1747-46e9-9e0e-bbdd090ed7be · inbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897d129d-c17a-4482-b52e-b6f713c33977 · inbound
RewardAnything: Generalizable Principle-Following Reward Models PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95bf17a7-838b-45b7-9e5f-4411dfae52a6 · inbound
Unlocking Recursive Thinking of LLMs: Alignment via Refinement PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5690b6-7ff5-447d-a462-7b390e3e9370 · inbound
Enterprise Large Language Model Evaluation Benchmark PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2824a9ce-3796-4668-88e5-2ce2bce15d94 · inbound
ASSURE: Metamorphic Testing for AI-powered Browser Extensions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbaa157b-8620-4055-8c5a-e9776f690c0a · inbound
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f5b672-60a7-49ca-80f8-3ab96fef1669 · inbound
UQ: Assessing Language Models on Unsolved Questions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc22d46-7426-418e-86fd-9bc4f304ab8f · inbound
Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec08c7f-4854-4589-bb93-55abe983f543 · inbound
A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc7892d8-4d59-463e-839d-6b0ab5c7be15 · inbound
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2539b22b-5c7a-41b0-96bf-fd1cf5f2e7e0 · inbound
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e752bfff-e420-4f2a-8fa5-457c6d21100e · inbound
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fcba09-600c-4791-b8ef-e2c79ef17086 · inbound
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b979a458-252f-42df-bff9-3b5337f3b205 · inbound
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5bdb24e0-15b8-4225-be9e-e3e79a9d3821 · inbound
One Year Later...The Harms Persist, But So Do We! PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 24b91149-6b8b-4aad-9592-02e132d1d712 · inbound
One Year Later...The Harms Persist, But So Do We! PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ed7ab0a-0e09-4f00-a851-e537189ed20d · inbound
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 76225087-f4e6-4a2b-bef9-4681601d9bef · inbound
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.