Pith. sign in

Paper Citation Record · LEDGER

LLM Evaluators Recognize and Favor Their Own Generations

As of 14 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 100 inbound Pith citation observations for arXiv:2404.13076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.13076 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T18:44:28.766639Z

measured 134 of 134 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 100 of 116 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:34:16.109049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T10:27:02.394079Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact28
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcc07dca-9d29-43a7-8dd4-6c5cc3e93b96 · outbound

This paper cites Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models.

LLM Evaluators Recognize and Favor Their Own Generations Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.852504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:8cf149b82a36bc898eb8d1c466115770685171101edc3caeea210dc65ec2d2c6

Observation 51465281-ecae-4508-998e-3f7b72b3479a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

LLM Evaluators Recognize and Favor Their Own Generations Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.845174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:b969b783bb38a79c218fb02266b2da09dfd75edaf45770becd44b61930357f6d

Observation e6899d29-4ce0-417e-8f85-ca0c29b0360c · outbound

This paper cites Taken out of context: On measuring situational awareness in LLMs.

LLM Evaluators Recognize and Favor Their Own Generations Taken out of context: On measuring situational awareness in LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.849090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:373f082dd75a304cf0412d80542cc9012250de0fb30edd423627e47870885513

Observation 645e2664-f561-446e-b5f4-cb4fb8d7985d · outbound

This paper cites VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use.

LLM Evaluators Recognize and Favor Their Own Generations VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.800274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:5071f25276a97ab9439a5886befcdf63c75d8e33bd52553f519c846b1b7daca0

Observation fdb6f66e-2cf4-4ca9-9fda-67092216283b · outbound

This paper cites doi: 10.1037/h0057532.

LLM Evaluators Recognize and Favor Their Own Generations doi: 10.1037/h0057532

Reference 5

Resolution
verified exact
doi, observed 2026-05-22T18:44:28.792101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:45fb45ce22b336cd83dcc3d1254ffd4cf0ecb19f7dc50962fcb84635e9812fcb

Observation 009d39b9-0577-47c3-a4e4-3805523c6671 · outbound

This paper cites GPTScore: Evaluate as You Desire.

LLM Evaluators Recognize and Favor Their Own Generations GPTScore: Evaluate as You Desire

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.804592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:fc3e594a5c5cee5d551daa870cb4e41987fce0b87e749b6534a61626c13d4aff

Observation 702d8014-845d-46e6-98f1-9a2cd8c3a526 · outbound

This paper cites doi: 10.3389/feduc.2023.

LLM Evaluators Recognize and Favor Their Own Generations doi: 10.3389/feduc.2023

Reference 7

Resolution
verified exact
doi, observed 2026-05-22T18:44:28.794619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:3bbff20ecca3f72b5107df92ef9debfcaac08f9ff172cc68295f8ff8e7eb082f

Observation 54824b3e-a20f-4d0c-b65f-0cc98bde91f4 · outbound

This paper cites Automatic Detection of Machine Generated Text: A Critical Survey.

LLM Evaluators Recognize and Favor Their Own Generations Automatic Detection of Machine Generated Text: A Critical Survey

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.809325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:055b4c9e3cb7f86fed72a3b125e651ab046bd77f886615fae6f61cd62001f069

Observation 1dcfae41-a0d0-4958-86b7-7ad8b65c6949 · outbound

This paper cites Language Models (Mostly) Know What They Know.

LLM Evaluators Recognize and Favor Their Own Generations Language Models (Mostly) Know What They Know

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.812968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:7aa54884035c6b16a58f7c932e27816c7aaa86362be948b98c919c766abd78af

Observation f899883b-0a27-470b-8d38-f46471bf40bf · outbound

This paper cites Benchmarking Cognitive Biases in Large Language Models as Evaluators.

LLM Evaluators Recognize and Favor Their Own Generations Benchmarking Cognitive Biases in Large Language Models as Evaluators

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.816877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:9abab9ca25be982468338330a0d1e33969cd8e6adc43b23005eb9325360da887

Observation 8c3b4fca-8c02-4bb3-9d92-43b1627d2c44 · outbound

This paper cites A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization.

LLM Evaluators Recognize and Favor Their Own Generations A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.820895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:34816cf681853efd2094c83a94ca88eb34583ebd43ae7739daa28c9e12dc0aaf

Observation fff6c62b-8db7-43c0-a1a7-62905b091570 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

LLM Evaluators Recognize and Favor Their Own Generations RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.825111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:68608d58ad8c1c3a2316c58e92b2bb84daf8d059cf87da0d43464f79b1e851db

Observation cb019c6b-8076-4151-a9e9-5fb79ce59bcc · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

LLM Evaluators Recognize and Favor Their Own Generations Scalable agent alignment via reward modeling: a research direction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.828850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:62fc71f2187883358114a4957349d5cee0b1da5c3de1b03ad8170eb4947e9064

Observation d2851bc6-048b-4ff5-8ed3-3ad223b0d344 · outbound

This paper cites original-date: 2023-05- 25T09:35:28Z.

LLM Evaluators Recognize and Favor Their Own Generations original-date: 2023-05- 25T09:35:28Z

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T18:44:28.908782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:d9be2ee0fc5e68ac73b0dfa49b4084fde8a4f5e8562034e48740d7358e5e53af

Observation d2160a96-009f-4a07-81c8-52bb9ac2f147 · outbound

This paper cites LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores.

LLM Evaluators Recognize and Favor Their Own Generations LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.834486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:7f8e59ef13531429e2c6fec5c556cb0c4d32bf046d20c795664b684b9a3ee0ec

Observation 29194dcf-ade2-4c4c-8384-ae1fbbaa8982 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

LLM Evaluators Recognize and Favor Their Own Generations Self-Refine: Iterative Refinement with Self-Feedback

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:44:28.838271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:69b057986d822c505635c6ba8ef618824c54b9966762abb72252ecb17684d5a9

Observation 80d8964d-6ab0-4501-92ff-6e823b96f15c · outbound

This paper cites doi: 10.18653/v1/K16-1028.

LLM Evaluators Recognize and Favor Their Own Generations doi: 10.18653/v1/K16-1028

Reference 17

Resolution
verified exact
doi, observed 2026-05-22T18:44:28.787534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:2f264d68f5825dd3df3d1c6fd39c02e56b44ee7480f4c80aa157e9e0bfb01527

Observation a2a9abef-832b-469f-8455-021f8dfc1a48 · outbound

This paper cites Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization.

LLM Evaluators Recognize and Favor Their Own Generations Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.855759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:4e437052b058f4ea005f73ef961f3e3668ad3b09dd28e43b35322efb879c1a77

Observation 095a6bf5-3ce4-4e29-ac5e-e8b24df512bb · outbound

This paper cites GPT-4 Technical Report.

LLM Evaluators Recognize and Favor Their Own Generations GPT-4 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.858878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:d61af80baae776d23c3cdf4f3296822ff593782066a30f7d266f163f1c005f84

Observation d582b9b6-d5ed-4597-91f2-37a359ee5adf · outbound

This paper cites Feedback Loops With Language Models Drive In-Context Reward Hacking.

LLM Evaluators Recognize and Favor Their Own Generations Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.862606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:fef36b4af8547ef7613c919bc2775f513a1f7d6d92cef70f329960a555a28bc3

Observation 8b7a79d1-ea2a-4d6e-b455-5176afe3af84 · outbound

This paper cites Towards Evaluating AI Systems for Moral Status Using Self-Reports.

LLM Evaluators Recognize and Favor Their Own Generations Towards Evaluating AI Systems for Moral Status Using Self-Reports

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.866163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:73662db61d3d2f8090b09b204b1c39ff64f30772243790f780a0f885ab96e8f6

Observation f9a0386e-0bc0-4eac-97a0-b2285f92fbad · outbound

This paper cites Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions.

LLM Evaluators Recognize and Favor Their Own Generations Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.869826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:c10c777d0a9e0b61d97bce83e20cf0a41bbc6ded6a52093cdfa32002f04ab981

Observation 5366f98b-251d-442b-82df-442143ec8918 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

LLM Evaluators Recognize and Favor Their Own Generations Self-critiquing models for assisting human evaluators

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.873093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:d7ccfe39324c2a10a63eb60d8ab5c464b855d3909cfc94fde3a39c76dddabf1b

Observation 3cf4097b-9dce-4f2d-aa9a-407bb4794189 · outbound

This paper cites Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models.

LLM Evaluators Recognize and Favor Their Own Generations Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.876395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:18cfaf97b214bbbb697935d5eb3a4a8c367adc4fa56c203b5ff3d84f0aa22923

Observation bc9ed97e-d90f-4a90-99fe-84a9c60fcef2 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

LLM Evaluators Recognize and Favor Their Own Generations Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:44:28.879481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:ab5f72ae7835b83f8fe3ac5c3de6fa094319127b548b97a472f72aaa65ee9822

Observation 5349df24-9bfb-4cd8-a64d-f8a025190f46 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LLM Evaluators Recognize and Favor Their Own Generations Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:44:28.882961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:b54268c07a0d0756c01068edf0b811b1542491f1f28c656f6949c42671e85075

Observation c6838485-6b2c-4ede-bb90-d3c423c6922f · outbound

This paper cites MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception.

LLM Evaluators Recognize and Favor Their Own Generations MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.886552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:56e04a6b8972c9554d747c978c1f6a5cd35abc0a710c2b30ca7a32ddc87d1dd5

Observation 007423c5-0146-4ca1-83c2-8c5cd31ee503 · outbound

This paper cites Recursively Summarizing Books with Human Feedback.

LLM Evaluators Recognize and Favor Their Own Generations Recursively Summarizing Books with Human Feedback

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:44:28.889779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:ced3aa2b1a2d6c0248cc3a95d90d92c29b7839cbe9fbfd715bf72a54d7b16e58

Observation 84c582e3-e9ab-4b39-85c1-ce314e3dcb98 · outbound

This paper cites A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions.

LLM Evaluators Recognize and Favor Their Own Generations A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.893321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:98217ccd8ac9871c54ddbad85b739833ec8d75973ad2a54c15c5b91c779918c1

Observation de6edc4f-4e31-4076-b32c-c778efd9f3e6 · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

LLM Evaluators Recognize and Favor Their Own Generations Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.896390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:f35c30767b1dd0ef380844f0199eb03b9f34a51668402d1bf1c9b244b3bae766

Observation 045ac1d4-3916-4db9-aacf-473ab1fb8b69 · outbound

This paper cites A Survey on Detection of LLMs-Generated Content.

LLM Evaluators Recognize and Favor Their Own Generations A Survey on Detection of LLMs-Generated Content

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.899519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:1701463e3050ef6471ad7f6b7a28fdc106ef066eed773c46a4623e0e844efc3e

Observation 5ea14b52-8aee-431c-83c4-14dab7ddef84 · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

LLM Evaluators Recognize and Favor Their Own Generations Do Large Language Models Know What They Don't Know?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.903039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:809265711bd9507d31133dcc017235293da67f35d9f8a7441df1adf1d6830e41

Observation bde149ff-327d-4cbf-93c6-85e846f42169 · outbound

This paper cites Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation.

LLM Evaluators Recognize and Favor Their Own Generations Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.906549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:09762f2fc96db91160dd29a28fee6c7ba2dc8e80f7a57ed1f7b024f7bcbd1e92

Observation 8ba96294-b41e-4a08-8522-45e72df13aac · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

LLM Evaluators Recognize and Favor Their Own Generations Evaluating Large Language Models at Evaluating Instruction Following

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.841917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:44:28.766639Z digest=sha256:4ff1834864d7bd4833f3e2b7a005f46d602879afffda3952bcc53acb4da6610a

Pith citing papers

Observation b6ad9c75-d36d-41fe-b11b-7d1cdf70caca · inbound

Inertia in Moral and Value Judgments of Large Language Models cites this paper.

Inertia in Moral and Value Judgments of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:05:50.087921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-23T22:05:17.165589Z digest=sha256:9f4117e0f415671ad57af61b032ef3cfb579c5c72decb67b59152b0f6a1911b7

Observation 8ad9e232-bdf9-4907-b2f8-29f0e184af6e · inbound

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct cites this paper.

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct LLM Evaluators Recognize and Favor Their Own Generations

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T19:53:23.262358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-23T19:51:14.630756Z digest=sha256:cddbc8b18c78142114c40bb82a3e237d9050f74056ac1c88cbf9eee7c237ef4a

Observation 3d787a9f-3e54-494b-aa6a-a31a0b37c7ed · inbound

Self-Preference Bias in LLM-as-a-Judge cites this paper.

Self-Preference Bias in LLM-as-a-Judge LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T14:46:32.515449Z digest=sha256:c5b3d5a7318b7e71151edfc0d14fe98ef5f68b5f7c4eec777783e12814b5975e

Observation 64c53fff-a207-48da-97f7-0de4cc5d41b3 · inbound

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas cites this paper.

SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas LLM Evaluators Recognize and Favor Their Own Generations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:16.109049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:16.109049Z digest=sha256:09b558c43950d6e449e232d470d6da30beecb06ecee140cf4bf590143dd41953

Observation c9a68d38-3b3a-4b4d-bed0-ef4950ac7d86 · inbound

The Illusion of Empathy: How AI Chatbots Shape Conversation Perception cites this paper.

The Illusion of Empathy: How AI Chatbots Shape Conversation Perception LLM Evaluators Recognize and Favor Their Own Generations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:08:26.654572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:08:26.654572Z digest=sha256:4ba55852c4a643c040863100b8e076bc3c131628079c77f5b7e3231307ad95af

Observation 00bab622-5de5-4eaf-ad4d-0c121f53aa78 · inbound

Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models cites this paper.

Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models LLM Evaluators Recognize and Favor Their Own Generations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T15:24:17.506298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:24:17.506298Z digest=sha256:33c6627aaac9c4498cde48084d4e9e35ce798d2320e9e1ee1b498a197af48e04

Observation eb71aab2-9895-4280-bedc-5cf82cbbfa18 · inbound

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text cites this paper.

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:37:15.817222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:37:15.817222Z digest=sha256:769ddac4740f2c11720b5e21a87597ac2eaf4185add3712a77f1daf75a2eafaa

Observation aa766792-8143-4ef5-bf0e-bf3cb1f7ab7f · inbound

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation cites this paper.

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation LLM Evaluators Recognize and Favor Their Own Generations

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T22:37:55.664752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:37:55.664752Z digest=sha256:41228ecc7d5c6ab4b19c237615bffd9f3e499802cb96e4cb80a6c38ebe63f8be

Observation 91e4fdbf-9aba-42a9-869e-91d59075299d · inbound

Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs cites this paper.

Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs LLM Evaluators Recognize and Favor Their Own Generations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:28:19.134006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:28:19.134006Z digest=sha256:2ef6c0a59fdb176e8d271851913e144068b0911ace8ccad33cafceab6f39e02e

Observation 3f9f0f66-694b-4dfe-828d-9e95598adb77 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods LLM Evaluators Recognize and Favor Their Own Generations

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:226c309b513920a8ebabc263e2cff3e2427a772491b2566bc1480480d505a5b1

Observation 49fd37c9-6cc2-4171-96c6-5553d8b27e9d · inbound

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs cites this paper.

QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs LLM Evaluators Recognize and Favor Their Own Generations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T14:40:29.503176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:40:29.503176Z digest=sha256:335411f70f19ca6abcec7fde5a1591fd37b862659681b619b8f0c2bbd7a666ea

Observation 24791312-1495-4c26-8b59-d9a64dea32eb · inbound

A Distributed Collaborative Retrieval Framework Excelling in All Queries and Corpora based on Zero-shot Rank-Oriented Automatic Evaluation cites this paper.

A Distributed Collaborative Retrieval Framework Excelling in All Queries and Corpora based on Zero-shot Rank-Oriented Automatic Evaluation LLM Evaluators Recognize and Favor Their Own Generations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:36:16.059013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:36:16.059013Z digest=sha256:eef75dd3ceab53bdc28177f2c04e66dcead0103b73843073e187dbde1021446c

Observation b46ea476-4908-44e3-ad12-f0438f16939e · inbound

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation cites this paper.

Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation LLM Evaluators Recognize and Favor Their Own Generations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:26.937180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:26.937180Z digest=sha256:4382cd30eb8a8bf07e35ed77c4126281bf54f2c371d0ed797420500450d03f69

Observation aa79b3a9-6e4c-4b1b-b5a9-b7fa7962d0a3 · inbound

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong cites this paper.

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong LLM Evaluators Recognize and Favor Their Own Generations

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:37:36.055475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T05:37:22.955895Z digest=sha256:762c365f3055ab6fcd353038ebe514bcee146048c095e0b127fa57c253daa8bb

Observation 0f1a9170-4b93-41c0-9fd0-fa6ecc3be8dd · inbound

Tuning LLM Judge Design Decisions for 1/1000 of the Cost cites this paper.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost LLM Evaluators Recognize and Favor Their Own Generations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.274806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.274806Z digest=sha256:2758406820007f1dc38d6521c4ff5d40fadc63ead5e8ea0134075a0369fc3a1d

Observation 11b10f94-2a9c-451e-aee3-6ca9f4d3ae77 · inbound

How do Humans and Language Models Reason About Creativity? A Comparative Analysis cites this paper.

How do Humans and Language Models Reason About Creativity? A Comparative Analysis LLM Evaluators Recognize and Favor Their Own Generations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T05:23:59.966207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:23:59.966207Z digest=sha256:f9bbfd9cbed876d40b010d398cccc308910ec406afda97af6877345cc02ce6f0

Observation 53b8c123-a100-4905-ad58-b225e8b1981d · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion LLM Evaluators Recognize and Favor Their Own Generations

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.437069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.437069Z digest=sha256:bd06ef5ef3799cabf81f15e4e2d7db8e3aaf38b751c1d8d3eda9663ad46d5ff4

Observation 4270e169-fa7e-4639-aa44-edc42768674f · inbound

Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators cites this paper.

Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators LLM Evaluators Recognize and Favor Their Own Generations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:23.990465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:55:23.990465Z digest=sha256:c466cf5e84b0e6fb14c4306d391b2ab4825515ccfdb92e685169ea6558313f98

Observation 22e590c6-4e69-47fd-9f01-a28c08ad1826 · inbound

Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy cites this paper.

Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy LLM Evaluators Recognize and Favor Their Own Generations

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T18:27:19.450554Z digest=sha256:89af1171365765e98ed045e3e9d6f046353dafd32d37f32306add3b110b8f56e

Observation 1675159f-a359-4928-b17f-45b930d2eaeb · inbound

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation cites this paper.

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation LLM Evaluators Recognize and Favor Their Own Generations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:17.178061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:17.178061Z digest=sha256:653b1168af908a28e85c11bf6fbc5ec9a0cd950c021990602b1f4f4a37538b82

Observation 4631f2bc-8179-4f1e-8198-984328c59ca8 · inbound

Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks cites this paper.

Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks LLM Evaluators Recognize and Favor Their Own Generations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:36.804850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:36.804850Z digest=sha256:4b459af21e7c79bdfccb2d8335b890d94699dd59d5032e958fa491a202b05187

Observation a5c44c37-4683-4d05-b1b3-15e79cd9f1b0 · inbound

Large Language Models for Predictive Analysis: How Far Are They? cites this paper.

Large Language Models for Predictive Analysis: How Far Are They? LLM Evaluators Recognize and Favor Their Own Generations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:24.225917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:05:24.225917Z digest=sha256:5b2f4e3c6cfe2f8b16150e35e720bbf6ebdac8c091f14d8cb786a2d708e671c2

Observation ff79056f-6f10-4e40-8146-140ca201ef0a · inbound

syftr: Pareto-Optimal Generative AI cites this paper.

syftr: Pareto-Optimal Generative AI LLM Evaluators Recognize and Favor Their Own Generations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:50.789445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:50.789445Z digest=sha256:b8fb92a3d2aff0b06e46c9b7ecad627e1373af8ad52c2582f7eac01bd973e784

Observation 47c1e098-0725-4fd0-8b0d-eb49ce71aa30 · inbound

Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement cites this paper.

Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement LLM Evaluators Recognize and Favor Their Own Generations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:59.026744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:59.026744Z digest=sha256:9d3fa0d7540b1f80d4be658b0eab9e5b1246d234379d93933fcd3ae339f25962

Observation aa0623f9-8c14-4145-b51f-12e4a0a176ba · inbound

SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL cites this paper.

SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL LLM Evaluators Recognize and Favor Their Own Generations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:27.078452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:27.078452Z digest=sha256:9a493f83e48091198c3be445ef8f3891eb712db161acf06229fb668947a9b1b3

Observation d381cae9-b08a-46da-97a9-afd1bd5c7c5f · inbound

Does It Make Sense to Speak of Introspection in Large Language Models? cites this paper.

Does It Make Sense to Speak of Introspection in Large Language Models? LLM Evaluators Recognize and Favor Their Own Generations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:37.672436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:30:37.672436Z digest=sha256:5204e21e81a694b347cb81143d3c1ee5fac2ec6baec3e347f5679b205f8a1ddf

Observation e104101e-e157-47ab-9789-edbe58235f57 · inbound

Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning cites this paper.

Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning LLM Evaluators Recognize and Favor Their Own Generations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:49.513946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:49.513946Z digest=sha256:ace42967096556493ff1d0aaffdae21ff805431f052f00dd152ae537284ca85b

Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.937721Z digest=sha256:c4aa580a418a965ee81c641ddf6fe602458c609e7d2542484be9670c1e953537

Observation a14b3870-6674-4ae9-bbd0-3c95a9225e71 · inbound

Evaluating LLM Agent Collusion in Double Auctions cites this paper.

Evaluating LLM Agent Collusion in Double Auctions LLM Evaluators Recognize and Favor Their Own Generations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:57:41.793378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:57:41.793378Z digest=sha256:26de3fb6c1a9355404ac200f002fc322a83f860e0f3170147b7b792683f777c6

Observation c2fac121-3779-4f6e-9a6b-077b66871fe1 · inbound

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations cites this paper.

The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations LLM Evaluators Recognize and Favor Their Own Generations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:02.884796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:02.884796Z digest=sha256:c6fede8fae203145b0cf6b5af8868c8b7b006b4604adbd12a898afdc10334540

Observation ff84316f-d482-4e36-9d8b-12c0ac14e873 · inbound

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? cites this paper.

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? LLM Evaluators Recognize and Favor Their Own Generations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:21.733580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:21.733580Z digest=sha256:c685287fe3fd9c3df58295b8507290dac487d567f12201f79d45607abebd111c

Observation 0112b2ea-cc66-4340-9db0-44ed6570f111 · inbound

Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities cites this paper.

Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities LLM Evaluators Recognize and Favor Their Own Generations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:55.416142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:55.416142Z digest=sha256:0b733dee141a77e634f385f20d9d85c0374145a8be2e42b11748bd0a132b8ce9

Observation e1b4a454-ab3b-4ee8-b80d-7bc864ed07a2 · inbound

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge cites this paper.

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge LLM Evaluators Recognize and Favor Their Own Generations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:40:42.222615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:40:42.222615Z digest=sha256:0c9253a37cdf1ad744d5b9663ccadf1bf07b3d2c2e38ae45a8cc7d7cb7510268

Observation 41f18b2d-7d8e-4ec7-9fe3-c26d127b57bf · inbound

Hermes 4 Technical Report cites this paper.

Hermes 4 Technical Report LLM Evaluators Recognize and Favor Their Own Generations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:54.605015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:32:54.605015Z digest=sha256:87b3b963523cc8bb202d7268f3f83adff387725e040a620fdb9fb19063e7c7ca

Observation d3be7bdb-19ae-4b70-853f-6e0d66ffbc5c · inbound

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators cites this paper.

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:44.497879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:44.497879Z digest=sha256:f10122e9e045636bb0d20397b2efe7af43201aca4881552d9457d15683589826

Observation f9e2d1a7-cba9-4927-9f62-d1632947881c · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering LLM Evaluators Recognize and Favor Their Own Generations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:49.140681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:49.140681Z digest=sha256:1422545c6f4999cf25eb4196e0bfc06eb026821a51db893c49d5b60095f32dce

Observation 964cce02-b2a8-4eff-a5cc-077d67ab2ccb · inbound

Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning cites this paper.

Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning LLM Evaluators Recognize and Favor Their Own Generations

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T01:17:12.349379Z digest=sha256:5a5a6a92d6a9ca05b6e44e25175708add59b42c6bee48ace23251e7c3c181dbc

Observation e05ec9cd-5555-4d59-9cce-b93d2cab1d80 · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications LLM Evaluators Recognize and Favor Their Own Generations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.318507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.318507Z digest=sha256:3bf6fe772c89ce606690216b17de90e86efa95485acf4234474a69fe3da8556c

Observation fe068385-672f-4acc-8838-5f54f5456a33 · inbound

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines cites this paper.

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines LLM Evaluators Recognize and Favor Their Own Generations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T17:52:54.781379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:52:54.781379Z digest=sha256:c510e4e0719874e1d727c12e83fd0f3ee920e80b22bf50990eb0efc307d2d897

Observation 3ee0af1d-7a91-4b7e-93e7-1f082686a1c9 · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules LLM Evaluators Recognize and Favor Their Own Generations

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:76c9ae798612e86cf72c221ae4e694b5ec75ae4e5b60e4afe9783005b5f0b00a

Observation 1155bc7e-469a-40ca-ae71-d1a593c9050f · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:f238b1aaeaea61f6c86d631d1cba87890f152468e8312ef91a4473be9a8eda5c

Observation e3ee2514-3842-445b-af4c-ae31f1626fc7 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.436057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.436057Z digest=sha256:f4b045225c90f213be6bd24b8b0a74e8f67642e9e53822aaeabfbd5e1b479d74

Observation b9785487-4c8a-40aa-bb92-59263e52e98a · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.346642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.346642Z digest=sha256:9bd52f13820cc0642762383b4be51bbd4574efd3eb3bb2d9f546be21d3b667e3

Observation 7dfb4733-05f4-4488-8edd-af6e6f648f84 · inbound

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation cites this paper.

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation LLM Evaluators Recognize and Favor Their Own Generations

Reference 24

Resolution
malformed identifier
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T09:13:55.204404Z digest=sha256:252f2e90ff0804de98b2c023f6ab96efc870a070f67fe5db5b191b887f806f79

Observation 61d69125-6ba6-47c1-87cd-e55399f84958 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines LLM Evaluators Recognize and Favor Their Own Generations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:522b1813d8fe62fc2b2faca3cbbaf6098b4407fb4915f9c98dd3b3945ab440a8

Observation 3b3a299b-074e-41ac-96ba-05882f1d4594 · inbound

SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters cites this paper.

SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T04:44:35.423054Z digest=sha256:83aa3fe71a750e0c1fc43af5a8fc8e6cf317b0ad688bd1fbb736ce86db8d3010

Observation 7c76ed3c-2c4e-4e46-84a1-5b478ca39ff6 · inbound

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator cites this paper.

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator LLM Evaluators Recognize and Favor Their Own Generations

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T03:39:30.528601Z digest=sha256:e47741046b8e27313d9b9ae243c6e5d0368bfa2966d32a95607b12a7e9e4b5d1

Observation 92271494-216a-49d7-a05d-28c2a6d139ad · inbound

Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces cites this paper.

Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces LLM Evaluators Recognize and Favor Their Own Generations

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:25:39.921635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T09:23:20.044291Z digest=sha256:327744184fc32b4ec748aa37c4cdb67079cec5dcfb840d75a9824af2977b8d0d

Observation bd81b812-3647-492b-8910-0e72f0fd08fb · inbound

When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents cites this paper.

When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents LLM Evaluators Recognize and Favor Their Own Generations

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T16:54:17.319933Z digest=sha256:9cdd3904b22c25ee114fc934e57626daf4b05c687ca24cea945de681dbbaa07e

Observation 42bb5091-77d0-40b4-9126-0cd81c93d2a3 · inbound

The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice cites this paper.

The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice LLM Evaluators Recognize and Favor Their Own Generations

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-09T14:30:58.313848Z digest=sha256:32f895e7e16c447ace61ba13032adf738d57a4396ed5c685ffd37b551e1e6f31

Observation 5585b901-868a-4eb7-8b74-35f06feef6c3 · inbound

Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks cites this paper.

Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks LLM Evaluators Recognize and Favor Their Own Generations

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T00:25:09.454515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T00:23:36.245000Z digest=sha256:60ee981cffa8f51f993935e0eabf044b4d19248c58f3df9d7e27de2ce2bbd353

Observation bd80a5bd-5a76-44e7-9ab9-29add114a840 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T16:57:49.396570Z digest=sha256:9b0042acb1f279e46a3d169eccf31fb63e2727250770aaaced80e576b10f0cc4

Observation 84b2d1c8-e56d-405f-bc81-680bc3bec2ab · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:26:54.426050Z digest=sha256:c70b383387ff8f7a620bfbceed5bea519e8d078fa3da6327239d93f7ff93d6ca

Observation 0df374ed-2447-44b8-af40-ff713cfc46ce · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:08:39.328446Z digest=sha256:8ca2f34c4c22a34fa34dfec50bb76cd8c2c49e34d512a717b66b1728281c86c6

Observation 3a9ffa62-4e76-445e-a94f-ef20c61f42d8 · inbound

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization cites this paper.

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T14:53:22.998757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:53:22.998757Z digest=sha256:1ff02fcc4928604d82cf123c65f8503b1b4fd46e12568541c1f89673499d9304

Observation 1a772f4b-7f63-44ca-84da-ab9c91adba96 · inbound

Automated alignment is harder than you think cites this paper.

Automated alignment is harder than you think LLM Evaluators Recognize and Favor Their Own Generations

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T09:47:05.526632Z digest=sha256:fe2632fb56cbdd73632d84bd7bc95e90fab604f8f77d04a62d7cc237e49c6ade

Observation c8c56dd2-ae08-41c1-a880-a984a517a4e2 · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LLM Evaluators Recognize and Favor Their Own Generations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:c1eef16db91776c158ae91e197db848877143eb7978a6a6dacea07e7da4f4c5a

Observation d9389ff7-0297-4d30-932e-5906c5000064 · inbound

Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation cites this paper.

Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation LLM Evaluators Recognize and Favor Their Own Generations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T01:50:01.032543Z digest=sha256:6db738b19824d544bdd225563648640fe5b311ff92156e1eeced6f87f8a8840d

Observation dc9c9793-796b-4c6b-94c6-8d9f4300128d · inbound

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps cites this paper.

Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps LLM Evaluators Recognize and Favor Their Own Generations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T19:05:01.031097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:58:47.229407Z digest=sha256:b1142410d7075b2ab15479a420b9cf454736cacd9cfae3168445d18d2450e1c2

Observation 03c187cb-2cb8-4f6d-b9de-c559218bc649 · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows LLM Evaluators Recognize and Favor Their Own Generations

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:7e936b85881d194237036a1e04535c9147388bf95582d0e9a7522ad8b5cea62e

Observation 1e9f3457-1218-40e9-991f-581a7c04d50c · inbound

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering cites this paper.

Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering LLM Evaluators Recognize and Favor Their Own Generations

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T06:53:44.993529Z digest=sha256:7ed4519e70338c8a3afe93b720a12bf9d55fb158ace659890f1ecb16181a57a4

Observation be36840e-8f59-4008-853d-320446c05b05 · inbound

Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment cites this paper.

Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment LLM Evaluators Recognize and Favor Their Own Generations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T06:13:06.598618Z digest=sha256:87c494b00aed2ad6523209d9457d2fb514ff8aadb1597f9a7c98a8366c822e37

Observation 086b1bbf-d5c1-4c91-9528-cb556e90c2df · inbound

Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks cites this paper.

Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks LLM Evaluators Recognize and Favor Their Own Generations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T02:10:13.030288Z digest=sha256:caf8c9c7317a99b83568c32ebcbcf7f38b5dc916a5057570e6cb39fe5a9ea1e2

Observation d4871818-cf6b-4abf-8172-bddee35d49a2 · inbound

Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost cites this paper.

Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost LLM Evaluators Recognize and Favor Their Own Generations

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-22T06:26:31.841138Z digest=sha256:e510eeee679020328c65c773582d338d6ec07d562ca3c5b92329a399bbd9f5ab

Observation 4602ff13-cb26-43ab-ac7d-c4c35a43d82e · inbound

AMEL: Accumulated Message Effects on LLM Judgments cites this paper.

AMEL: Accumulated Message Effects on LLM Judgments LLM Evaluators Recognize and Favor Their Own Generations

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:44:28.909456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T05:08:30.607268Z digest=sha256:94add58b363abc5efb6b5286e84d4a123726f348dcec58cb8ab4a58317ae1b43

Observation 224160ec-611c-4dd2-8bfc-bce2f95a16d5 · inbound

AMEL: Accumulated Message Effects on LLM Judgments cites this paper.

AMEL: Accumulated Message Effects on LLM Judgments LLM Evaluators Recognize and Favor Their Own Generations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:04:56.460770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T17:04:22.688250Z digest=sha256:ed8c62cf2ab218dabc154b5f421c7929f67b64e9bef09a42feb38e075a5900ef

Observation f0fccc42-0e8b-4263-ad5b-51cebae75a85 · inbound

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays cites this paper.

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays LLM Evaluators Recognize and Favor Their Own Generations

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.045484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c6a7be691a4b5e8e9a885df09a4bd59ddbc57e4086b4a806dfd1b1a88790ddb3

Observation 5d1384b0-29d4-42e9-9394-7b9f78a1df6d · inbound

Plans for Evaluating Structured Generative Search Summaries cites this paper.

Plans for Evaluating Structured Generative Search Summaries LLM Evaluators Recognize and Favor Their Own Generations

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T16:33:39.098123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T16:26:58.658036Z digest=sha256:c9fa5962938b1a1d9b58cf6ca3dd2bb0b67c4075f2741a08b45b1959179674e5

Observation ca11954a-5440-4534-a444-be37301db93d · inbound

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation cites this paper.

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation LLM Evaluators Recognize and Favor Their Own Generations

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:23:50.870842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T18:17:52.284353Z digest=sha256:676db435d45b6637941a96e34575ef0d4ec821b68d7cf62d0ae5f2308af99d3f

Observation aac68ec3-2357-412c-af72-2410bce24d70 · inbound

Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering cites this paper.

Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering LLM Evaluators Recognize and Favor Their Own Generations

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T16:53:40.892169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T16:47:30.514271Z digest=sha256:d9eb004ab38685d911c75c26ecfcf05ed9582bad3b8528b3d59d824034647a63

Observation f9458aa7-00b6-4ba5-aa06-3e9c3e730ffd · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm LLM Evaluators Recognize and Favor Their Own Generations

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:03:26.708876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:2e3b6fd129f62bc9f70858766b1a437533c46e1ea0bf7ae089a55e103d9f6142

Observation 612f9d56-5b7b-46e9-b572-d33c35edc7f3 · inbound

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law cites this paper.

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law LLM Evaluators Recognize and Favor Their Own Generations

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T12:53:26.595076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T12:50:07.673561Z digest=sha256:98a2e6e53cf07de647f18f447a989773fea74a4df0f94a12d94052ef7328b4a3

Observation 76640b33-bdb6-4b2e-bdd6-6760516ccf37 · inbound

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law cites this paper.

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law LLM Evaluators Recognize and Favor Their Own Generations

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T08:05:31.149817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T07:59:42.981532Z digest=sha256:74df45ff8361701bea2999f869e2a84cac3568c2d635fcc59a9d2ca75fa97ba0

Observation 0c048db2-ceb2-4b26-9b35-27f32a6b8f03 · inbound

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning LLM Evaluators Recognize and Favor Their Own Generations

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:16:43.747093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:e7587555dc176e3629a06cab3ff790ed6d43c9c3fe45e3939c659b66d00d42a5

Observation 993b4273-bacb-4f2c-b3da-ac882ca7ee59 · inbound

MIRAI: Prediction and Generation of High-Impact Academic Research cites this paper.

MIRAI: Prediction and Generation of High-Impact Academic Research LLM Evaluators Recognize and Favor Their Own Generations

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:06:55.988288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:29:38.654022Z digest=sha256:c09ef3bf37a085187002b7354b3c5ccd2162c7c6f3944bff8146e37801ea7491

Observation f52edc08-0430-40ac-9ac5-bba9bbaeb894 · inbound

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) cites this paper.

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) LLM Evaluators Recognize and Favor Their Own Generations

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:26:59.396726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T01:15:20.867740Z digest=sha256:01ddb05a00edfd51c0c2eff482c49003bac8fcbb896d8ae65e29d17692e51fc3

Observation 6bc326df-a6e9-42cf-ae0d-0403ce696e1c · inbound

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) cites this paper.

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) LLM Evaluators Recognize and Favor Their Own Generations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T12:21:42.740103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:21:42.740103Z digest=sha256:667ed796f025b8a5e9431563334996fcc1e8d58924708f7775b9e5c73066fc6e

Observation f1b2ea92-b418-4767-9c43-48bde695430c · inbound

Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship cites this paper.

Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship LLM Evaluators Recognize and Favor Their Own Generations

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:59:33.325115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T17:26:22.687010Z digest=sha256:2547b4e6c787591cc49495a697d3bfcbf3a8e231cc13b8ed46de5edbea8930b4

Observation c92eec94-0e14-465b-80b7-70387c349623 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:41.502683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:5c2386813d674d938a57d2c14b81f5f709a0f971d8fd799ef6a240ea2ac06ae7

Observation c2d4a730-8c8d-4242-b191-cd558e79200b · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning LLM Evaluators Recognize and Favor Their Own Generations

Reference 207

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:59:44.859847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:a4ad526fc46c7e5863c0f83d95955abb248b924f09d8f101b4f22efafc75f574

Observation e98114e5-c12e-4aaf-b6c4-49f6d4fb1000 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning LLM Evaluators Recognize and Favor Their Own Generations

Reference 206

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:55:59.685310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:e0fdcfce8bc60b03ed05d6aa56b9840da51a6d1deca9457f66cdd97acbaba1e2

Observation 7d6d7bdd-aa01-4fa1-9990-692e1e358d81 · inbound

Same question, different history: language, national identity, and credit in large language models cites this paper.

Same question, different history: language, national identity, and credit in large language models LLM Evaluators Recognize and Favor Their Own Generations

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:39:45.890306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T08:35:08.947990Z digest=sha256:95af4b4d20b831208cbeb8267daeda61f3ace684ddf8b5d096866221cb93e70c

Observation fe034b2e-076a-4f39-80f7-6982ff3fc2ce · inbound

Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems cites this paper.

Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:59:46.352976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T08:16:51.419186Z digest=sha256:e476cb819793ef482a81460fdc511cf42839a37c6da5a1237f41b1d6a4e51cb4

Observation 52dfce92-da63-4df8-a78d-5f6854a86556 · inbound

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability? cites this paper.

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability? LLM Evaluators Recognize and Favor Their Own Generations

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:19:57.636033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T00:47:00.310205Z digest=sha256:794525017b6af93ff058c57a7f8648f1ee8414f4641d766604f46dd6dda60209

Observation b95f64ae-201f-41c2-833a-c69eae1ee06e · inbound

Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning cites this paper.

Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning LLM Evaluators Recognize and Favor Their Own Generations

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:14:25.793313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T08:06:18.672841Z digest=sha256:2a71e427489e28db5ba099ee70dde4ae423c04f72f7b40dd1f49933f05b075b5

Observation 57dcb337-3dba-4f3d-bc30-e937444b6dab · inbound

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots cites this paper.

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots LLM Evaluators Recognize and Favor Their Own Generations

Reference 14

Resolution
malformed identifier
local_arxiv, observed 2026-06-30T06:14:18.003107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T06:09:32.771308Z digest=sha256:653ace4d69e771f9ac94740ba8ad97c8cc3246f019208591ac50335c9d3c6a0e

Observation 3a8b2b1a-4bb8-476a-beb4-73b75ee05a76 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? LLM Evaluators Recognize and Favor Their Own Generations

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T06:04:21.542266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:5fd684d119bb0c345f0430ee9b0e4206bd5240afa4e8ec3e4ffa99e7daa053af

Observation eac85c47-0c14-49f1-a73a-0f7a6ec6b203 · inbound

AGC-Bench: Measuring Artificial General Creativity cites this paper.

AGC-Bench: Measuring Artificial General Creativity LLM Evaluators Recognize and Favor Their Own Generations

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:36:56.078980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-02T12:33:53.029578Z digest=sha256:f51b3fe8ad0a8065d360e6735a7394e62e8b8ed1c3ffeed3cdf11d2ad41794bf

Observation 20cf4520-2494-4af9-95fb-95fd090a7aec · inbound

AGC-Bench: Measuring Artificial General Creativity cites this paper.

AGC-Bench: Measuring Artificial General Creativity LLM Evaluators Recognize and Favor Their Own Generations

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T21:28:58.100940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T21:25:14.030920Z digest=sha256:2c87d823620e0780e4a3a2f5420b37f0910feed097951c4aef5e035b88bd77eb

Observation fb9b4ffe-3174-4067-8e82-d2a3d73095a1 · inbound

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems cites this paper.

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems LLM Evaluators Recognize and Favor Their Own Generations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T04:31:30.304599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T04:31:30.304599Z digest=sha256:65de66d14758bb5049796cd06ceca5ee55aa03178596cbc877d39db350bb5c3f

Observation d12c5f09-2d28-41b5-9b0a-af4ef3273d9b · inbound

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals cites this paper.

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals LLM Evaluators Recognize and Favor Their Own Generations

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:40.949844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T00:56:31.193905Z digest=sha256:87d11c6d3bb34d913efeed313819178e8cc137ba0c3e1ee32aa3d4888869456d

Observation 6f5c6fd4-a7d4-42c4-bb3d-51443e668ded · inbound

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals cites this paper.

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals LLM Evaluators Recognize and Favor Their Own Generations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:59:16.438403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:59:16.438403Z digest=sha256:a891b4c955153f520d5860f044196d9d9ff5f809a02f3eef6cf7d534a77ad200

Observation dabcdac3-ffb5-45fd-9477-fe485d4f41f5 · inbound

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring cites this paper.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring LLM Evaluators Recognize and Favor Their Own Generations

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:56:41.062491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T00:52:47.537142Z digest=sha256:9532e5882a4c9c6dc9719def898632ce45deeec4117d7941c253520991e1392d

Observation 8d280af0-842a-48a9-b803-0013f8727146 · inbound

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment cites this paper.

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment LLM Evaluators Recognize and Favor Their Own Generations

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T10:27:02.395348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T10:20:46.103409Z digest=sha256:e12b4ee3db52ce0ce752330e67c9c74220337fe61fd5fbb7f1ab8f3e33a2dc23

Observation b6db7c52-4cb7-4940-8b8b-7490af6ffdea · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins LLM Evaluators Recognize and Favor Their Own Generations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:08.574095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:08.574095Z digest=sha256:703decbc1b0c1e82ad0b93d13a1e8f3910d91064525fc642c646c9d01865df27

Observation b9547897-a3d4-4137-8cd2-34840a19cedd · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ LLM Evaluators Recognize and Favor Their Own Generations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:b44b3964dbc968cfec22f456e589f294bfeeb6d96bd595b92c7651fb09eafa74

Observation a867c62c-c34c-4039-b64c-6050286e24eb · inbound

Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques cites this paper.

Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques LLM Evaluators Recognize and Favor Their Own Generations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T01:18:31.914413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:18:31.914413Z digest=sha256:d69e331c07c14dd1e3e13507cdb4f170ff305fa5d6975e03bcddcc36ae51408e

Observation 45b5ff09-409e-4fe8-aadf-9ec1e549ff9c · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation LLM Evaluators Recognize and Favor Their Own Generations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:069fb8c08da4088aec5998e44531b529d0d8d39b09dcc2fb0ad64517b932f348

Observation a788b32a-97bc-4c67-8963-7817758cecf1 · inbound

Does Multi-Agent Debate Improve AI Feedback on Research Papers? cites this paper.

Does Multi-Agent Debate Improve AI Feedback on Research Papers? LLM Evaluators Recognize and Favor Their Own Generations

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T01:21:37.369091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:21:37.369091Z digest=sha256:290075d9d983d2df78637dc9e3394f09c2d19840691a5eb2ed09342c4426a56b

Observation 2c6e025c-a15a-4108-849f-d1a3973c8ea3 · inbound

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models cites this paper.

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models LLM Evaluators Recognize and Favor Their Own Generations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:08.913869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:08.913869Z digest=sha256:7b462eb5717759bc8e5a871aaa7fb549ec260d56e16b57212b4d6d20778ea551