Pith. sign in

Paper Citation Record · LEDGER

PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2306.05087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.05087 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:45:45.853992Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

29
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7f2840b-0625-461b-b166-3e9f954e8a34 · inbound

Large Language Models are not Fair Evaluators cites this paper.

Large Language Models are not Fair Evaluators PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T12:10:42.437190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-17T12:10:42.248005Z digest=sha256:2427c88c1361fae734d1b400968725850604a5780e6b29950baeba52776b5efa

Observation 8f028035-56a9-427c-97d9-e0cc4538cea5 · inbound

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts cites this paper.

Enhancing Instructional Quality: Leveraging Computer-Assisted Textual Analysis to Generate In-Depth Insights from Educational Artifacts PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:08:48.314734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T03:07:32.103513Z digest=sha256:8182f0e435a17e9b1d6eeecec679de2db58327f2d2d729445354f5ced7cf3c26

Observation ca9376c1-50fb-4a7c-a3a3-1d12180566b2 · inbound

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions cites this paper.

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:53:40.681101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T00:52:52.056076Z digest=sha256:d7d58643fbb3eb9d829cb65102d898f1806031677ceeb2b62bca3246e1b59bcd

Observation 5a615d9c-6817-456f-8d16-de2b6575f431 · inbound

Explainable LLM-driven Multi-dimensional Distillation for E-Commerce Relevance Learning cites this paper.

Explainable LLM-driven Multi-dimensional Distillation for E-Commerce Relevance Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T16:58:26.711052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:58:26.711052Z digest=sha256:ef92dd4ef90bbcab4413ecf8045e57b799816dfa915d066bc7e6feb7f60178af

Observation 9fe5d0bf-c40b-462d-a63d-2e5ab29e633f · inbound

Natural Language Reinforcement Learning cites this paper.

Natural Language Reinforcement Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:26:45.128901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:26:45.128901Z digest=sha256:27d120423464e1b4aeedda3503154307ae306cc735b8c9076bb8fdd215b16a89

Observation 01a30c9f-f494-458c-9fd3-d7f26a3bae36 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.114172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:714c5355e7b3d25c58ea01a6828c3231c7cc8a33aa34dc06e8103e49097957de

Observation 0221db6b-c1de-4d69-a5e6-f1612d27294a · inbound

Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability cites this paper.

Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:18:01.931025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:18:01.931025Z digest=sha256:377332162b9955b98d931adeb1e5260f1d32a8e8941885276e672e94ca7b62ea

Observation 42f1b288-458c-4997-8b0d-d16089b18662 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.657694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:0ea15549a9d8175600a3962c6c69f95a984aabb8d743c5b7a7a566efd08701c0

Observation bfede6a1-675d-4638-943a-23621287a16e · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.855139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.855139Z digest=sha256:b0a53594efb017ca72d37d8a38ada5d8b027b13f12de153ad3f5191b9ea362e9

Observation 4774d543-5df4-4e24-886f-d26bcb351cba · inbound

SedarEval: Automated Evaluation using Self-Adaptive Rubrics cites this paper.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.969582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.969582Z digest=sha256:371cf7928eef0475b98d29a56411924a7026712968793715535e1d44815a163a

Observation 21bfd10f-c0a8-4abf-8fd7-24edbe2b6879 · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.342852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.342852Z digest=sha256:540a3b3c6e393c2ce56d4a8760998cedcfc4be65c39e2f8f42fd0e79b68e42dd

Observation 3dd17b74-4b40-49ce-88b7-23af43d7aca5 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.657674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:05846bb3e6a4fa06598ae312d3645ae372c9e39443720e9ed91818a668a5531b

Observation 6196b8d1-9e22-4a3a-9b08-9494653f8cda · inbound

PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines cites this paper.

PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:45:45.853992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:45:45.853992Z digest=sha256:3f88a27563b02773572b96ed6724c2f086d83b55b2baf96292ddbd65326b91ed

Observation ee0e2d8b-5bfd-4410-a804-db642cd30dc2 · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.088890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.088890Z digest=sha256:182639f07611af075bba5e76298b521cf1c9a8c0a11d1ef94849a3a839b352e0

Observation 8856035a-d684-4cf4-bcbf-bc7eb0087672 · inbound

LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning cites this paper.

LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T01:06:34.317570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:06:34.317570Z digest=sha256:0d678062d3e0291e4a3002c2ba330993f0c2f094c7fb72cd6e2d77868efd7634

Observation 2ca64449-a955-447c-9d9c-03e911f9d531 · inbound

am-ELO: A Stable Framework for Arena-based LLM Evaluation cites this paper.

am-ELO: A Stable Framework for Arena-based LLM Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:42.110600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:58:42.110600Z digest=sha256:91799996d2944864b73641a057a66d0f461da4d28ab8421021757b58fbc88f12

Observation d1a924e3-b953-46bb-b261-45d2f70bc1f7 · inbound

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge cites this paper.

J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:10.690874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:50:10.690874Z digest=sha256:9521715551b95c64ac77517263e6524fc016077d26729c3d3e87386356c7583e

Observation 31e83056-c601-4f4b-9e4c-0b4f6bc2ea90 · inbound

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation cites this paper.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.971613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.971613Z digest=sha256:59d79cb9e15b38c84cd3182d50304cda97bad02d8d76d4aa197a741f4545e426

Observation 4f5947a4-37a9-463a-8ab1-80b4140c2f43 · inbound

LLM-based Evaluation Policy Extraction for Ecological Modeling cites this paper.

LLM-based Evaluation Policy Extraction for Ecological Modeling PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:05.526187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:05.526187Z digest=sha256:2473784a916a451a4c2de7d8c94c9804ff5558bdb0326db790d01a147109672e

Observation ce379c9d-1747-46e9-9e0e-bbdd090ed7be · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:06.346963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:06.346963Z digest=sha256:6750a338366c53c4fd7330ed5b614e9a6a2eea0debde6fc0557ad3db793313f7

Observation 897d129d-c17a-4482-b52e-b6f713c33977 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.750627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.750627Z digest=sha256:07b15ed38fee1e68c68176f645b682190f24f89ce073876bb93cb6351d539811

Observation 95bf17a7-838b-45b7-9e5f-4411dfae52a6 · inbound

Unlocking Recursive Thinking of LLMs: Alignment via Refinement cites this paper.

Unlocking Recursive Thinking of LLMs: Alignment via Refinement PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:02.101272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:02.101272Z digest=sha256:d572db05b0eeb6ae22102bab543c0142223dc7038ec7b8e2566cb6aa4bf530fd

Observation bb5690b6-7ff5-447d-a462-7b390e3e9370 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:35.136253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:35.136253Z digest=sha256:1a25bfa594b598b34e7fd940ff527bfe2357b58c76e5dd10621c31e9c4484cfd

Observation 2824a9ce-3796-4668-88e5-2ce2bce15d94 · inbound

ASSURE: Metamorphic Testing for AI-powered Browser Extensions cites this paper.

ASSURE: Metamorphic Testing for AI-powered Browser Extensions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:30.600858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:30.600858Z digest=sha256:80471a9f652c0bcfb0494aa9a6fdae8b337724fdf1afe0fc06b9a8513491cf26

Observation dbaa157b-8620-4055-8c5a-e9776f690c0a · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:25.456372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:25.456372Z digest=sha256:03390acf65cd2126de0440d9d5163bb1eb3a92f4b0c8e0e5125335659849aea3

Observation 24f5b672-60a7-49ca-80f8-3ab96fef1669 · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.828751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.828751Z digest=sha256:579bab474381b4ed8ad92d5ad7c9909044ebc09c51641db3a6f0b5a942365399

Observation 0dc22d46-7426-418e-86fd-9bc4f304ab8f · inbound

Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval cites this paper.

Beyond Text: Aligning Vision and Language for Multimodal E-Commerce Retrieval PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-15T15:00:25.407602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T15:00:25.407602Z digest=sha256:e6a507fb7b1f8fe27e2e167953fddbcb3dd5bc78b33ed92b27d34ba51ba33329

Observation dec08c7f-4854-4589-bb93-55abe983f543 · inbound

A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges cites this paper.

A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:20:58.903293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:a850a0a664116d9f7fb845fcefb85334b830e6cbc408ae314459aad0671687dc

Observation bc7892d8-4d59-463e-839d-6b0ab5c7be15 · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:39.586160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:604e9537f34fa52ce60964a6f58440ba2d1de08a9a8d18cf897b4c7942b6e220

Observation 2539b22b-5c7a-41b0-96bf-fd1cf5f2e7e0 · inbound

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why cites this paper.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:53:58.856496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T21:43:59.462151Z digest=sha256:2b8ecd51457c8717758da48069f9c558c504d54b712b747b20d7464b1497529a

Observation e752bfff-e420-4f2a-8fa5-457c6d21100e · inbound

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why cites this paper.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T02:20:18.474864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:20:18.474864Z digest=sha256:66f350385ec60492f6bf7b5d5d5285587542a78fd3b6df2db96246f5e9671a21

Observation 75fcba09-600c-4791-b8ef-e2c79ef17086 · inbound

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering cites this paper.

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:55:06.075948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T22:54:03.054871Z digest=sha256:d71b4df7e7b24b435f55cf24c20b4c3ede1c7904a312f1624f1eaa907050c5c8

Observation b979a458-252f-42df-bff9-3b5337f3b205 · inbound

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing cites this paper.

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:40.126393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T15:48:26.303462Z digest=sha256:a50decbe0a2bca4c823aa56d814a266c5e8c8644dc1320a465c754ce3a95c681

Observation 5bdb24e0-15b8-4225-be9e-e3e79a9d3821 · inbound

One Year Later...The Harms Persist, But So Do We! cites this paper.

One Year Later...The Harms Persist, But So Do We! PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:35:29.507789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T06:33:58.246849Z digest=sha256:c7c17e2ad3c308ef80e292eac8d75f9f9e03a09ebc75e8d3fd2aa49fdfe437f0

Observation 24b91149-6b8b-4aad-9592-02e132d1d712 · inbound

One Year Later...The Harms Persist, But So Do We! cites this paper.

One Year Later...The Harms Persist, But So Do We! PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:47:27.609506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T21:38:29.468179Z digest=sha256:6cf966402d02274839f16369a32c4a2d785c44ad03cccf37f37a4e04ba69c54c

Observation 1ed7ab0a-0e09-4f00-a851-e537189ed20d · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.528798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:6379699bc0d97ace9101bf9e5c2c69d8767aa5f297dff49c6f8c2cc0fb23b05d

Observation 76225087-f4e6-4a2b-bef9-4681601d9bef · inbound

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability cites this paper.

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T05:56:50.505931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T05:47:32.670216Z digest=sha256:53613e90cba1089e6dc9f3f5a6646453a0365114a25466976666704fe3db6768