Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T18:44:28.766639Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 100 inbound Pith citation observations for arXiv:2404.13076.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T18:44:28.766639Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:34:16.109049Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T10:27:02.394079Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dcc07dca-9d29-43a7-8dd4-6c5cc3e93b96 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51465281-ecae-4508-998e-3f7b72b3479a · outbound
LLM Evaluators Recognize and Favor Their Own Generations Constitutional AI: Harmlessness from AI Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6899d29-4ce0-417e-8f85-ca0c29b0360c · outbound
LLM Evaluators Recognize and Favor Their Own Generations Taken out of context: On measuring situational awareness in LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 645e2664-f561-446e-b5f4-cb4fb8d7985d · outbound
LLM Evaluators Recognize and Favor Their Own Generations VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fdb6f66e-2cf4-4ca9-9fda-67092216283b · outbound
LLM Evaluators Recognize and Favor Their Own Generations doi: 10.1037/h0057532
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 009d39b9-0577-47c3-a4e4-3805523c6671 · outbound
LLM Evaluators Recognize and Favor Their Own Generations GPTScore: Evaluate as You Desire
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 702d8014-845d-46e6-98f1-9a2cd8c3a526 · outbound
LLM Evaluators Recognize and Favor Their Own Generations doi: 10.3389/feduc.2023
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 54824b3e-a20f-4d0c-b65f-0cc98bde91f4 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Automatic Detection of Machine Generated Text: A Critical Survey
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1dcfae41-a0d0-4958-86b7-7ad8b65c6949 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Language Models (Mostly) Know What They Know
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f899883b-0a27-470b-8d38-f46471bf40bf · outbound
LLM Evaluators Recognize and Favor Their Own Generations Benchmarking Cognitive Biases in Large Language Models as Evaluators
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8c3b4fca-8c02-4bb3-9d92-43b1627d2c44 · outbound
LLM Evaluators Recognize and Favor Their Own Generations A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fff6c62b-8db7-43c0-a1a7-62905b091570 · outbound
LLM Evaluators Recognize and Favor Their Own Generations RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb019c6b-8076-4151-a9e9-5fb79ce59bcc · outbound
LLM Evaluators Recognize and Favor Their Own Generations Scalable agent alignment via reward modeling: a research direction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d2851bc6-048b-4ff5-8ed3-3ad223b0d344 · outbound
LLM Evaluators Recognize and Favor Their Own Generations original-date: 2023-05- 25T09:35:28Z
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d2160a96-009f-4a07-81c8-52bb9ac2f147 · outbound
LLM Evaluators Recognize and Favor Their Own Generations LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 29194dcf-ade2-4c4c-8384-ae1fbbaa8982 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Self-Refine: Iterative Refinement with Self-Feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 80d8964d-6ab0-4501-92ff-6e823b96f15c · outbound
LLM Evaluators Recognize and Favor Their Own Generations doi: 10.18653/v1/K16-1028
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a2a9abef-832b-469f-8455-021f8dfc1a48 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 095a6bf5-3ce4-4e29-ac5e-e8b24df512bb · outbound
LLM Evaluators Recognize and Favor Their Own Generations GPT-4 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d582b9b6-d5ed-4597-91f2-37a359ee5adf · outbound
LLM Evaluators Recognize and Favor Their Own Generations Feedback Loops With Language Models Drive In-Context Reward Hacking
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b7a79d1-ea2a-4d6e-b455-5176afe3af84 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Towards Evaluating AI Systems for Moral Status Using Self-Reports
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9a0386e-0bc0-4eac-97a0-b2285f92fbad · outbound
LLM Evaluators Recognize and Favor Their Own Generations Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5366f98b-251d-442b-82df-442143ec8918 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Self-critiquing models for assisting human evaluators
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3cf4097b-9dce-4f2d-aa9a-407bb4794189 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bc9ed97e-d90f-4a90-99fe-84a9c60fcef2 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5349df24-9bfb-4cd8-a64d-f8a025190f46 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c6838485-6b2c-4ede-bb90-d3c423c6922f · outbound
LLM Evaluators Recognize and Favor Their Own Generations MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 007423c5-0146-4ca1-83c2-8c5cd31ee503 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Recursively Summarizing Books with Human Feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 84c582e3-e9ab-4b39-85c1-ce314e3dcb98 · outbound
LLM Evaluators Recognize and Favor Their Own Generations A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation de6edc4f-4e31-4076-b32c-c778efd9f3e6 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 045ac1d4-3916-4db9-aacf-473ab1fb8b69 · outbound
LLM Evaluators Recognize and Favor Their Own Generations A Survey on Detection of LLMs-Generated Content
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5ea14b52-8aee-431c-83c4-14dab7ddef84 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Do Large Language Models Know What They Don't Know?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bde149ff-327d-4cbf-93c6-85e846f42169 · outbound
LLM Evaluators Recognize and Favor Their Own Generations Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ba96294-b41e-4a08-8522-45e72df13aac · outbound
LLM Evaluators Recognize and Favor Their Own Generations Evaluating Large Language Models at Evaluating Instruction Following
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6ad9c75-d36d-41fe-b11b-7d1cdf70caca · inbound
Inertia in Moral and Value Judgments of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ad9e232-bdf9-4907-b2f8-29f0e184af6e · inbound
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct LLM Evaluators Recognize and Favor Their Own Generations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3d787a9f-3e54-494b-aa6a-a31a0b37c7ed · inbound
Self-Preference Bias in LLM-as-a-Judge LLM Evaluators Recognize and Favor Their Own Generations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 64c53fff-a207-48da-97f7-0de4cc5d41b3 · inbound
SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas LLM Evaluators Recognize and Favor Their Own Generations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a68d38-3b3a-4b4d-bed0-ef4950ac7d86 · inbound
The Illusion of Empathy: How AI Chatbots Shape Conversation Perception LLM Evaluators Recognize and Favor Their Own Generations
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bab622-5de5-4eaf-ad4d-0c121f53aa78 · inbound
Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models LLM Evaluators Recognize and Favor Their Own Generations
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb71aab2-9895-4280-bedc-5cf82cbbfa18 · inbound
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text LLM Evaluators Recognize and Favor Their Own Generations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa766792-8143-4ef5-bf0e-bf3cb1f7ab7f · inbound
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation LLM Evaluators Recognize and Favor Their Own Generations
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e4fdbf-9aba-42a9-869e-91d59075299d · inbound
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs LLM Evaluators Recognize and Favor Their Own Generations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f9f0f66-694b-4dfe-828d-9e95598adb77 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods LLM Evaluators Recognize and Favor Their Own Generations
Reference 177
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 49fd37c9-6cc2-4171-96c6-5553d8b27e9d · inbound
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs LLM Evaluators Recognize and Favor Their Own Generations
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24791312-1495-4c26-8b59-d9a64dea32eb · inbound
A Distributed Collaborative Retrieval Framework Excelling in All Queries and Corpora based on Zero-shot Rank-Oriented Automatic Evaluation LLM Evaluators Recognize and Favor Their Own Generations
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46ea476-4908-44e3-ad12-f0438f16939e · inbound
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation LLM Evaluators Recognize and Favor Their Own Generations
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa79b3a9-6e4c-4b1b-b5a9-b7fa7962d0a3 · inbound
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong LLM Evaluators Recognize and Favor Their Own Generations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0f1a9170-4b93-41c0-9fd0-fa6ecc3be8dd · inbound
Tuning LLM Judge Design Decisions for 1/1000 of the Cost LLM Evaluators Recognize and Favor Their Own Generations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b10f94-2a9c-451e-aee3-6ca9f4d3ae77 · inbound
How do Humans and Language Models Reason About Creativity? A Comparative Analysis LLM Evaluators Recognize and Favor Their Own Generations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b8c123-a100-4905-ad58-b225e8b1981d · inbound
AI Alignment at Your Discretion LLM Evaluators Recognize and Favor Their Own Generations
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4270e169-fa7e-4639-aa44-edc42768674f · inbound
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators LLM Evaluators Recognize and Favor Their Own Generations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e590c6-4e69-47fd-9f01-a28c08ad1826 · inbound
Multi-Stage Retrieval for Operational Technology Cybersecurity Compliance Using Large Language Models: A Railway Casestudy LLM Evaluators Recognize and Favor Their Own Generations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1675159f-a359-4928-b17f-45b930d2eaeb · inbound
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation LLM Evaluators Recognize and Favor Their Own Generations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4631f2bc-8179-4f1e-8198-984328c59ca8 · inbound
Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks LLM Evaluators Recognize and Favor Their Own Generations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c44c37-4683-4d05-b1b3-15e79cd9f1b0 · inbound
Large Language Models for Predictive Analysis: How Far Are They? LLM Evaluators Recognize and Favor Their Own Generations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff79056f-6f10-4e40-8146-140ca201ef0a · inbound
syftr: Pareto-Optimal Generative AI LLM Evaluators Recognize and Favor Their Own Generations
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c1e098-0725-4fd0-8b0d-eb49ce71aa30 · inbound
Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement LLM Evaluators Recognize and Favor Their Own Generations
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0623f9-8c14-4145-b51f-12e4a0a176ba · inbound
SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL LLM Evaluators Recognize and Favor Their Own Generations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d381cae9-b08a-46da-97a9-afd1bd5c7c5f · inbound
Does It Make Sense to Speak of Introspection in Large Language Models? LLM Evaluators Recognize and Favor Their Own Generations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e104101e-e157-47ab-9789-edbe58235f57 · inbound
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning LLM Evaluators Recognize and Favor Their Own Generations
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · inbound
How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14b3870-6674-4ae9-bbd0-3c95a9225e71 · inbound
Evaluating LLM Agent Collusion in Double Auctions LLM Evaluators Recognize and Favor Their Own Generations
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2fac121-3779-4f6e-9a6b-077b66871fe1 · inbound
The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations LLM Evaluators Recognize and Favor Their Own Generations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff84316f-d482-4e36-9d8b-12c0ac14e873 · inbound
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? LLM Evaluators Recognize and Favor Their Own Generations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0112b2ea-cc66-4340-9db0-44ed6570f111 · inbound
Cascaded Information Disclosure for Generalized Evaluation of Problem Solving Capabilities LLM Evaluators Recognize and Favor Their Own Generations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b4a454-ab3b-4ee8-b80d-7bc864ed07a2 · inbound
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge LLM Evaluators Recognize and Favor Their Own Generations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f18b2d-7d8e-4ec7-9fe3-c26d127b57bf · inbound
Hermes 4 Technical Report LLM Evaluators Recognize and Favor Their Own Generations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3be7bdb-19ae-4b70-853f-6e0d66ffbc5c · inbound
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e2d1a7-cba9-4927-9f62-d1632947881c · inbound
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering LLM Evaluators Recognize and Favor Their Own Generations
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964cce02-b2a8-4eff-a5cc-077d67ab2ccb · inbound
Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning LLM Evaluators Recognize and Favor Their Own Generations
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e05ec9cd-5555-4d59-9cce-b93d2cab1d80 · inbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications LLM Evaluators Recognize and Favor Their Own Generations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe068385-672f-4acc-8838-5f54f5456a33 · inbound
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines LLM Evaluators Recognize and Favor Their Own Generations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee0af1d-7a91-4b7e-93e7-1f082686a1c9 · inbound
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules LLM Evaluators Recognize and Favor Their Own Generations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1155bc7e-469a-40ca-ae71-d1a593c9050f · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e3ee2514-3842-445b-af4c-ae31f1626fc7 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9785487-4c8a-40aa-bb92-59263e52e98a · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfb4733-05f4-4488-8edd-af6e6f648f84 · inbound
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation LLM Evaluators Recognize and Favor Their Own Generations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 61d69125-6ba6-47c1-87cd-e55399f84958 · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines LLM Evaluators Recognize and Favor Their Own Generations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b3a299b-074e-41ac-96ba-05882f1d4594 · inbound
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters LLM Evaluators Recognize and Favor Their Own Generations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7c76ed3c-2c4e-4e46-84a1-5b478ca39ff6 · inbound
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator LLM Evaluators Recognize and Favor Their Own Generations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 92271494-216a-49d7-a05d-28c2a6d139ad · inbound
Noncrossing Duality and the Geometry of Positive Tropical Linear Spaces LLM Evaluators Recognize and Favor Their Own Generations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bd81b812-3647-492b-8910-0e72f0fd08fb · inbound
When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents LLM Evaluators Recognize and Favor Their Own Generations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 42bb5091-77d0-40b4-9126-0cd81c93d2a3 · inbound
The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice LLM Evaluators Recognize and Favor Their Own Generations
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5585b901-868a-4eb7-8b74-35f06feef6c3 · inbound
Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks LLM Evaluators Recognize and Favor Their Own Generations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bd80a5bd-5a76-44e7-9ab9-29add114a840 · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 84b2d1c8-e56d-405f-bc81-680bc3bec2ab · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0df374ed-2447-44b8-af40-ff713cfc46ce · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a9ffa62-4e76-445e-a94f-ef20c61f42d8 · inbound
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization LLM Evaluators Recognize and Favor Their Own Generations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a772f4b-7f63-44ca-84da-ab9c91adba96 · inbound
Automated alignment is harder than you think LLM Evaluators Recognize and Favor Their Own Generations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c8c56dd2-ae08-41c1-a880-a984a517a4e2 · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities LLM Evaluators Recognize and Favor Their Own Generations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9389ff7-0297-4d30-932e-5906c5000064 · inbound
Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation LLM Evaluators Recognize and Favor Their Own Generations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dc9c9793-796b-4c6b-94c6-8d9f4300128d · inbound
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps LLM Evaluators Recognize and Favor Their Own Generations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 03c187cb-2cb8-4f6d-b9de-c559218bc649 · inbound
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows LLM Evaluators Recognize and Favor Their Own Generations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1e9f3457-1218-40e9-991f-581a7c04d50c · inbound
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering LLM Evaluators Recognize and Favor Their Own Generations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be36840e-8f59-4008-853d-320446c05b05 · inbound
Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment LLM Evaluators Recognize and Favor Their Own Generations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 086b1bbf-d5c1-4c91-9528-cb556e90c2df · inbound
Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks LLM Evaluators Recognize and Favor Their Own Generations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d4871818-cf6b-4abf-8172-bddee35d49a2 · inbound
Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost LLM Evaluators Recognize and Favor Their Own Generations
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4602ff13-cb26-43ab-ac7d-c4c35a43d82e · inbound
AMEL: Accumulated Message Effects on LLM Judgments LLM Evaluators Recognize and Favor Their Own Generations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 224160ec-611c-4dd2-8bfc-bce2f95a16d5 · inbound
AMEL: Accumulated Message Effects on LLM Judgments LLM Evaluators Recognize and Favor Their Own Generations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0fccc42-0e8b-4263-ad5b-51cebae75a85 · inbound
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays LLM Evaluators Recognize and Favor Their Own Generations
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d1384b0-29d4-42e9-9394-7b9f78a1df6d · inbound
Plans for Evaluating Structured Generative Search Summaries LLM Evaluators Recognize and Favor Their Own Generations
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ca11954a-5440-4534-a444-be37301db93d · inbound
Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation LLM Evaluators Recognize and Favor Their Own Generations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aac68ec3-2357-412c-af72-2410bce24d70 · inbound
Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering LLM Evaluators Recognize and Favor Their Own Generations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9458aa7-00b6-4ba5-aa06-3e9c3e730ffd · inbound
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm LLM Evaluators Recognize and Favor Their Own Generations
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 612f9d56-5b7b-46e9-b572-d33c35edc7f3 · inbound
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law LLM Evaluators Recognize and Favor Their Own Generations
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 76640b33-bdb6-4b2e-bdd6-6760516ccf37 · inbound
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law LLM Evaluators Recognize and Favor Their Own Generations
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0c048db2-ceb2-4b26-9b35-27f32a6b8f03 · inbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning LLM Evaluators Recognize and Favor Their Own Generations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 993b4273-bacb-4f2c-b3da-ac882ca7ee59 · inbound
MIRAI: Prediction and Generation of High-Impact Academic Research LLM Evaluators Recognize and Favor Their Own Generations
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f52edc08-0430-40ac-9ac5-bba9bbaeb894 · inbound
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) LLM Evaluators Recognize and Favor Their Own Generations
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6bc326df-a6e9-42cf-ae0d-0403ce696e1c · inbound
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) LLM Evaluators Recognize and Favor Their Own Generations
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b2ea92-b418-4767-9c43-48bde695430c · inbound
Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship LLM Evaluators Recognize and Favor Their Own Generations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c92eec94-0e14-465b-80b7-70387c349623 · inbound
BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c2d4a730-8c8d-4242-b191-cd558e79200b · inbound
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning LLM Evaluators Recognize and Favor Their Own Generations
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e98114e5-c12e-4aaf-b6c4-49f6d4fb1000 · inbound
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning LLM Evaluators Recognize and Favor Their Own Generations
Reference 206
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7d6d7bdd-aa01-4fa1-9990-692e1e358d81 · inbound
Same question, different history: language, national identity, and credit in large language models LLM Evaluators Recognize and Favor Their Own Generations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe034b2e-076a-4f39-80f7-6982ff3fc2ce · inbound
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems LLM Evaluators Recognize and Favor Their Own Generations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 52dfce92-da63-4df8-a78d-5f6854a86556 · inbound
Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability? LLM Evaluators Recognize and Favor Their Own Generations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b95f64ae-201f-41c2-833a-c69eae1ee06e · inbound
Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning LLM Evaluators Recognize and Favor Their Own Generations
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 57dcb337-3dba-4f3d-bc30-e937444b6dab · inbound
EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots LLM Evaluators Recognize and Favor Their Own Generations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a8b2b1a-4bb8-476a-beb4-73b75ee05a76 · inbound
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? LLM Evaluators Recognize and Favor Their Own Generations
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eac85c47-0c14-49f1-a73a-0f7a6ec6b203 · inbound
AGC-Bench: Measuring Artificial General Creativity LLM Evaluators Recognize and Favor Their Own Generations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 20cf4520-2494-4af9-95fb-95fd090a7aec · inbound
AGC-Bench: Measuring Artificial General Creativity LLM Evaluators Recognize and Favor Their Own Generations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fb9b4ffe-3174-4067-8e82-d2a3d73095a1 · inbound
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems LLM Evaluators Recognize and Favor Their Own Generations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12c5f09-2d28-41b5-9b0a-af4ef3273d9b · inbound
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals LLM Evaluators Recognize and Favor Their Own Generations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6f5c6fd4-a7d4-42c4-bb3d-51443e668ded · inbound
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals LLM Evaluators Recognize and Favor Their Own Generations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dabcdac3-ffb5-45fd-9477-fe485d4f41f5 · inbound
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring LLM Evaluators Recognize and Favor Their Own Generations
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8d280af0-842a-48a9-b803-0013f8727146 · inbound
Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment LLM Evaluators Recognize and Favor Their Own Generations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6db7c52-4cb7-4940-8b8b-7490af6ffdea · inbound
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins LLM Evaluators Recognize and Favor Their Own Generations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9547897-a3d4-4137-8cd2-34840a19cedd · inbound
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ LLM Evaluators Recognize and Favor Their Own Generations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a867c62c-c34c-4039-b64c-6050286e24eb · inbound
Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques LLM Evaluators Recognize and Favor Their Own Generations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b5ff09-409e-4fe8-aadf-9ec1e549ff9c · inbound
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation LLM Evaluators Recognize and Favor Their Own Generations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a788b32a-97bc-4c67-8963-7817758cecf1 · inbound
Does Multi-Agent Debate Improve AI Feedback on Research Papers? LLM Evaluators Recognize and Favor Their Own Generations
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6e025c-a15a-4108-849f-d1a3973c8ea3 · inbound
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models LLM Evaluators Recognize and Favor Their Own Generations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.