Pith. sign in

Paper Citation Record · LEDGER

Engineering AI Judge Systems

As of 13 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2411.17793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17793 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:59:18.434546Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:32:47.756343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:33:18.382028Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e07a633d-a329-4f9a-822d-47e8a45c6cad · outbound

This paper cites Claude 3 sonnet has become very lazy,.

Engineering AI Judge Systems Claude 3 sonnet has become very lazy,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.294465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.139088Z digest=sha256:786c64d140a3c6e42a25a46f94b8f4cceba5870bd25e59afa9dd19881e6d6033

Observation 0641d58c-b4d9-45ae-a544-f469463b3b76 · outbound

This paper cites Do you guy think the cost of gpt-4 is high,.

Engineering AI Judge Systems Do you guy think the cost of gpt-4 is high,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.271435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.148769Z digest=sha256:fab6bbd8a2e2fc7e5074f17b5960e5f128417e0043e40a685bbe71156bdd873d

Observation 7e88b40b-4104-444c-a7f7-955aa0354452 · outbound

This paper cites Gpt-4 is crazy expensive,.

Engineering AI Judge Systems Gpt-4 is crazy expensive,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.259355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.152711Z digest=sha256:8cf72f0114a9d60b6b136c5235625f537a387bbe7ed6ad9579ffefb6b909d806

Observation 6e7e3ff0-5b01-49c5-8426-0c64801f4e06 · outbound

This paper cites Open llm leaderboard,.

Engineering AI Judge Systems Open llm leaderboard,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.247901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.157331Z digest=sha256:5c4f3ab33f50e4898cabc9e1132a5dff13a7b0c8d7d16f8e643a39585cf9f1a9

Observation a62f18e3-1362-46cb-9a8a-d5e82ba916dd · outbound

This paper cites Use agent metrics & llm judges to evaluate app performance,.

Engineering AI Judge Systems Use agent metrics & llm judges to evaluate app performance,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.236198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.161619Z digest=sha256:b6f70d09dd433bfa51211758459baa78544cf32c43b42dc6611aacb5c0513c5a

Observation 2d24a91c-5b70-4273-bff3-90967b30d88b · outbound

This paper cites 2030 software engineering,.

Engineering AI Judge Systems 2030 software engineering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.224133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.166669Z digest=sha256:8716030ed670e266f7f86ece3535f2e83605869ffc41a91173acd44f9aa32f4f

Observation 947c5aeb-8423-4bc7-9dda-6df89c092418 · outbound

This paper cites The acm international conference on the foundations of software en- gineering (fse) 2024,.

Engineering AI Judge Systems The acm international conference on the foundations of software en- gineering (fse) 2024,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.210844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.170796Z digest=sha256:b9d9f23fa15e8cd9cc561a37508f0de3ccf3d37b46b05bc877257892d43e71d2

Observation 05ca5043-7480-4d7e-8b16-34110261baea · outbound

This paper cites Fm+se summit 2024,.

Engineering AI Judge Systems Fm+se summit 2024,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.198220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.175887Z digest=sha256:fdd068589a3a41daee295e37cc07bd54f97c3d96f4c409e2319b6e2eb99df88d

Observation 40e1b4f1-a253-4d29-a173-96e35aede6e5 · outbound

This paper cites Opea initiative (open platform for enterprise ai (opea),.

Engineering AI Judge Systems Opea initiative (open platform for enterprise ai (opea),

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.186229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.180239Z digest=sha256:d1391654c3633f5a4cfcffaaf02845752093adfe3ef5cf465ab4bad9a2a46404

Observation c3f25090-1dee-424a-9f4d-ef4075c739ff · outbound

This paper cites Re-thinking data strategy and integration for artificial intelligence: concepts, opportuni- ties, and challenges,.

Engineering AI Judge Systems Re-thinking data strategy and integration for artificial intelligence: concepts, opportuni- ties, and challenges,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.172791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.185399Z digest=sha256:d2aaa483d0b5de1891ef26135c1bcc571be1e51dd1414633c856438cb716d87f

Observation 8a957eb4-3f66-4202-b0f1-8923229c292d · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Engineering AI Judge Systems Constitutional AI: Harmlessness from AI Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.189110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.189110Z digest=sha256:cdb26605931ca42f497fc91bb3d0d2f13f55e559b76b1bd9aec95856d4c3b286

Observation c8023431-371b-4128-96f9-6d4dd42ca5d2 · outbound

This paper cites Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs.

Engineering AI Judge Systems Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.194285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.194285Z digest=sha256:ccef2085ff3df617f8ff692325c95a3f85783b8ce50e40c7cc3496b2496ff02d

Observation 92cf6acd-09fd-4426-b0ef-6bf32fe48823 · outbound

This paper cites Meteor: an automatic metric for MT evalua- tion with improved correlation with human judgments,.

Engineering AI Judge Systems Meteor: an automatic metric for MT evalua- tion with improved correlation with human judgments,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.159428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.198230Z digest=sha256:b0a820ec4b68a22ed3e9652809464f036d56f0164f01f441c4cef06c3b15319d

Observation cfb058a0-0668-4acf-87dd-305edac4a074 · outbound

This paper cites Test driven development: By example addison-wesley,.

Engineering AI Judge Systems Test driven development: By example addison-wesley,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.147267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.201903Z digest=sha256:229a9f7099729b482782f8f9b38adf1f30ff0eef878d9e75d4e9d745831840d5

Observation 71493e67-f94e-479d-a0cd-602bac9facd1 · outbound

This paper cites On the dangers of stochastic parrots: can language models be too big?.

Engineering AI Judge Systems On the dangers of stochastic parrots: can language models be too big?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.135774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.205397Z digest=sha256:8c1ce2f5417b887e7c9b439e4c71c923fc9a07fed9aab28d1f3ce019d6c947b0

Observation c2ac4a3a-4cff-48fd-97bf-e8fd9f8a5dd5 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Engineering AI Judge Systems Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.208969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.208969Z digest=sha256:f0275dd0439d6cdae1be4de78e2b97b5d01e6e3a0ddf401b851f0e5234c54af8

Observation 62e3339c-7943-452c-847f-8cbde46d5111 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Engineering AI Judge Systems ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.212705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.212705Z digest=sha256:0fe86916fe2418e3fb0f5a4ae44b697ce58ce22078a3056c5a68ad77095e8eb8

Observation 0935c594-8e41-416a-95db-9f86a061e380 · outbound

This paper cites A survey on evaluation of large language models,.

Engineering AI Judge Systems A survey on evaluation of large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.122927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.217544Z digest=sha256:87e3e1c708bd5010cc0779b4a2c0d7ed061c26f86159c7c2d015e0fb8afb9ab8

Observation e65b98a8-a1b8-4100-a638-188fd934f478 · outbound

This paper cites Unleashing the potential of prompt engineering for large language models.

Engineering AI Judge Systems Unleashing the potential of prompt engineering for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.221024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.221024Z digest=sha256:4c58783f9875bc753352bb4e41b10514fac0fdc4f21ad5616761f37b4205222c

Observation a970bc9c-49d3-465b-9583-2f2c3cd9b86e · outbound

This paper cites Towards training reproducible deep learning models,.

Engineering AI Judge Systems Towards training reproducible deep learning models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.110698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.224849Z digest=sha256:e2c7dc9a45c89618a2a8ab7991f58fbf0365254463a18543a54ea287c74d5a4a

Observation d2fa95c7-8c31-498e-8bdf-dec0558784c8 · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

Engineering AI Judge Systems Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.228507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.228507Z digest=sha256:928bfd24e7cf46f67d7a9ca73834e47c960a4b1c8f2bd1eaa4b1a4f4db42c3a9

Observation c5ad56e2-3a5c-4c91-a84c-08c4ee58a9c2 · outbound

This paper cites How is ChatGPT's behavior changing over time?.

Engineering AI Judge Systems How is ChatGPT's behavior changing over time?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.233491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.233491Z digest=sha256:363a5586d47deba9dab3c2256696005484ae9e6a295a251ff7accdc3cb12ed0d

Observation 64e25685-2f9e-4c1b-9113-cc8e95572a23 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Engineering AI Judge Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.237211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.237211Z digest=sha256:b8cead4de8bf2d093a7d9933f5f421a78841198b534dd2dd2f2177a7531ba28a

Observation 2ddc37fd-3306-4aac-bcb2-a6cc389bc29f · outbound

This paper cites Available: https://www.reddit.com/r/ClaudeAI/comments/ 1bv8ww5/claude 3 sonnet has become very lazy/.

Engineering AI Judge Systems Available: https://www.reddit.com/r/ClaudeAI/comments/ 1bv8ww5/claude 3 sonnet has become very lazy/

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.282971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.143956Z digest=sha256:ff8f73e751aa97869e9ae7852894fb50175a333421ad40a5f2fa1300af71974e

Observation 10d1d72c-38ce-4f89-8809-379d44341fc4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Engineering AI Judge Systems Training Verifiers to Solve Math Word Problems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.241194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.241194Z digest=sha256:113104859f6cb3980259f6105079093384b27e7309807873614bb158a46401bd

Observation 54a73f2e-70a0-43f2-89e7-684da527c5e6 · outbound

This paper cites Evalullm: llm assisted evaluation of generative outputs,.

Engineering AI Judge Systems Evalullm: llm assisted evaluation of generative outputs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.098115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.245859Z digest=sha256:1693048b15cbed62782dbe06962f1b97a7a078bd982bc2c0987c0574b1097c6b

Observation 2ab6465c-c725-4a04-b939-394da8855761 · outbound

This paper cites Qlora: ef- ficient finetuning of quantized llms,.

Engineering AI Judge Systems Qlora: ef- ficient finetuning of quantized llms,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.086396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.249801Z digest=sha256:c3e4e33fe989b402de0a1f8378cc5893a40c9d9f85c1d8102c08451702524b24

Observation aec2f2f2-46ed-4a8e-9983-d24f1b299583 · outbound

This paper cites Fira: fine-grained graph-based code change representation for automated com- mit message generation,.

Engineering AI Judge Systems Fira: fine-grained graph-based code change representation for automated com- mit message generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.075156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.253234Z digest=sha256:2e4784a7ca556f8c45cc795e6f1fd922d213193a04b076220b586938671f9564

Observation 2c4f6a65-7ce3-4ec6-a895-0cbd0da38f37 · outbound

This paper cites Alpacafarm: a simulation framework for methods that learn from human feedback,.

Engineering AI Judge Systems Alpacafarm: a simulation framework for methods that learn from human feedback,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.063807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.256879Z digest=sha256:eea7c44912df6c7b2c1d62a78ca19474169da0e7a98c8e3713c2233d6dd8183a

Observation 9831402a-94fc-4b4a-bb73-8d52cc27b3c7 · outbound

This paper cites Bias and fairness in large language models: a survey,.

Engineering AI Judge Systems Bias and fairness in large language models: a survey,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.052740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.260862Z digest=sha256:73f2b304fa36e0d3dc4b31a43a1d30d84ec039c514407e778aa6dace55df2da5

Observation 535c660a-2eb5-4e4a-8c8c-2d803839f901 · outbound

This paper cites Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning.

Engineering AI Judge Systems Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.266978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.266978Z digest=sha256:1ec7b5c6e9692a71f535e19507c66cb1c3ca8af6c9d18d4a0048ce911f921adf

Observation 4944e0f9-35d0-4078-8829-e8f5a610a8e9 · outbound

This paper cites LLM-based NLG Evaluation: Current Status and Challenges.

Engineering AI Judge Systems LLM-based NLG Evaluation: Current Status and Challenges

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.271869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.271869Z digest=sha256:a2814ca7e720520a1f411464d8125b39be612afd096efeb3c7ff1626036bb882

Observation 964d0df2-5403-4e1b-8136-a64903194b90 · outbound

This paper cites Fm+se vision 2030,.

Engineering AI Judge Systems Fm+se vision 2030,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.041370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.276954Z digest=sha256:4aae2fe650330e02f5e66ce4541d660d7be4c8f1ad4f6eefd106e8d80e0220fd

Observation 80f2e391-f7a0-4f0b-9247-39c261d7d047 · outbound

This paper cites Rethinking software engineering in the foundation model era: a curated catalogue of challenges in the development of trustworthy fmware,.

Engineering AI Judge Systems Rethinking software engineering in the foundation model era: a curated catalogue of challenges in the development of trustworthy fmware,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.029239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.280771Z digest=sha256:8856607a8652563bc65c4f0017ed84310c700d97baf44c169154a9cebe55fe6e

Observation 38c394b4-fb5d-4586-b09d-45b121f5a5ef · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Engineering AI Judge Systems Measuring Massive Multitask Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.285054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.285054Z digest=sha256:8670c74a8c0a850f86ea172731ed3978ea28eebda48bd0c25e372ffce5a7de65

Observation abc6d780-04f7-4f85-bcb2-8b0d8f8b217b · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Engineering AI Judge Systems A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.289124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.289124Z digest=sha256:2a48adeb10926c5f772a94a2f3f43c2af399535b99dcf848de64a3f0fc68219b

Observation 8f6c5c9f-b355-4868-9765-adf08a204dd1 · outbound

This paper cites AI safety via debate.

Engineering AI Judge Systems AI safety via debate

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.293117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.293117Z digest=sha256:018a37630c393db108971f2a7ac44ad4cae1690b7ab401cc5caadc6cfdf0e861

Observation 601eb689-23ad-4590-b3ee-368c505574dd · outbound

This paper cites Kejriwal, Domain-specific knowledge graph construction.

Engineering AI Judge Systems Kejriwal, Domain-specific knowledge graph construction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.018229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.297827Z digest=sha256:c8b2c24b4e04cba851162af41f0880b07a5f404e754379355ade701109230d38

Observation dce9d1bd-1dd9-479a-9f89-db458e0f2469 · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

Engineering AI Judge Systems On scalable oversight with weak LLMs judging strong LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.301527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.301527Z digest=sha256:189cc4fa9d8658295b9cffa840907588b5728e25ffbb04fd61d59ad0cf3886e6

Observation 28284a8d-410a-4e89-b44b-1d1c75aff2f3 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Engineering AI Judge Systems Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.305344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.305344Z digest=sha256:8e1664c2d0c733fe6c455425b42ec0c06bf92e47132e5f663f4592c39c47ae0f

Observation 84a337a4-3562-4833-9efb-ddcd9b7d00f5 · outbound

This paper cites Dspy: compiling declarative language model calls into state-of-the-art pipelines,.

Engineering AI Judge Systems Dspy: compiling declarative language model calls into state-of-the-art pipelines,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:19.007287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.309227Z digest=sha256:8616fadfc651c0f9b912d415ab57cf727a70cc40380d02527d148d9f76c912e8

Observation f9386766-e1fd-461f-9013-e120f2ed1260 · outbound

This paper cites Software engineering for machine learning applications (semla) 2023,.

Engineering AI Judge Systems Software engineering for machine learning applications (semla) 2023,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.993839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.313972Z digest=sha256:e9093dbe65f7cc0d71f264c64756eedd3b90fd788e53070db8bb8b02ef342aad

Observation dac9e283-2527-47fb-8842-c2df3acbb6cc · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Engineering AI Judge Systems Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.317823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.317823Z digest=sha256:87fe099e754f027951fdf4fdbab73ba0013e168c61b36bd1912c82ad214b5d28

Observation 3aec817e-5fca-4029-af5c-2709ab851d84 · outbound

This paper cites On the role of knowledge graphs in explainable ai,.

Engineering AI Judge Systems On the role of knowledge graphs in explainable ai,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.981263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.322382Z digest=sha256:41a6208f184b87d9495e0771ff20743d0e27ac066a5d82d23e4466c17a654fb7

Observation 2104f4a8-8d00-479d-aa47-f19387188b8c · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Engineering AI Judge Systems Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.326017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.326017Z digest=sha256:3911a2dafda0f097329cc9e7489ea88b365707448b26653f5aca58b12839e9ac

Observation 9724063b-752b-41cc-92a5-2e8b9e3fa414 · outbound

This paper cites Rouge: a package for automatic evaluation of summaries,.

Engineering AI Judge Systems Rouge: a package for automatic evaluation of summaries,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.968535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.330973Z digest=sha256:00dcde917af9a3f31b5a823aac145bfa668036f4dbc1e951597b97fd20d5bb66

Observation 56a5924d-8edc-4c5f-8af5-dc72367b80be · outbound

This paper cites Best Practices and Lessons Learned on Synthetic Data.

Engineering AI Judge Systems Best Practices and Lessons Learned on Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.334631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.334631Z digest=sha256:ad2f4e0488dc2755a66cc2b0663d6edbe65e3ceb1cf0e1d6885c187949017ec8

Observation a57f28fd-7676-4364-9c5f-b975c931315e · outbound

This paper cites Calibrating LLM-Based Evaluator.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:877098cd0dc28eec03f60f75dd4a557ad0dd24b041ec85caec3e3b3de549d8d2

Observation ae9b3172-cb8b-4264-b00e-b1fc5795ef68 · outbound

This paper cites Generating training data with language models: towards zero-shot language understanding,.

Engineering AI Judge Systems Generating training data with language models: towards zero-shot language understanding,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.956636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.343429Z digest=sha256:d798eaebca686cc7e78379f0e57df6083b8556da77aea2e9b0bce0b16cf4998d

Observation eb717707-4edb-4c03-aeb5-23bf499ef0a4 · outbound

This paper cites Cider: robust consensus-based image description evaluation,.

Engineering AI Judge Systems Cider: robust consensus-based image description evaluation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.944650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.348047Z digest=sha256:e1166b51f6d5316c3daf6ef15813ab44e53757be10c400829a1312568ac4fe8d

Observation 95a55474-bc6c-4ec1-81f6-73192c62e639 · outbound

This paper cites A framework for evaluating and improving requirements specifications based on the developers and testers perspective,.

Engineering AI Judge Systems A framework for evaluating and improving requirements specifications based on the developers and testers perspective,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.932435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.352122Z digest=sha256:e6d5f9e7dd2f727d5929667332b6292556e4f312ffe9269fb653b24381f2e928

Observation 698c9c48-3d7c-4663-9190-028b2013d533 · outbound

This paper cites Human-Centered Design Recommendations for LLM-as-a-Judge.

Engineering AI Judge Systems Human-Centered Design Recommendations for LLM-as-a-Judge

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.356623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.356623Z digest=sha256:7023a273f454c0a382320d77f7557a84b169e2afafbd9682542f3445d0560dd7

Observation d380be56-886c-43af-a80c-d04386398aaf · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Engineering AI Judge Systems Bleu: a method for automatic evaluation of machine translation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.920316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.360737Z digest=sha256:8a51204b4392d753d67b17030d2249962be36ed579b11b8e684564759ac5216b

Observation a5a90d46-9e76-4ecb-94ce-4bf898a63156 · outbound

This paper cites The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities.

Engineering AI Judge Systems The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.365781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.365781Z digest=sha256:ba8e2847488dcc00a9bef6e2527b7629b0c5fde557c46c3b32ad2dd08407f329

Observation 57920572-3461-49cf-b1b9-5a97faa05cee · outbound

This paper cites Verbosity Bias in Preference Labeling by Large Language Models.

Engineering AI Judge Systems Verbosity Bias in Preference Labeling by Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.369902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.369902Z digest=sha256:8706cae2f6528b5139d53f4916ae316dd0a512e3785f540577b14c3fc05f7548

Observation 0f1b10a0-d706-48ae-a6fd-e0e483db6f0e · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

Engineering AI Judge Systems BLEURT: Learning Robust Metrics for Text Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.374841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.374841Z digest=sha256:2d5a3ce7ae895e08704c105840124a4a6d14c6d2f1ebe18b271be6c64e879e21

Observation 1cd8717b-f6b1-4089-aec4-d5b796dbcae0 · outbound

This paper cites On automatic summarization of what and why information in source code changes,.

Engineering AI Judge Systems On automatic summarization of what and why information in source code changes,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.906996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.379788Z digest=sha256:33cc0718e51a84cbe5df0fa7483086dee1037fd25d8b48df038e19e88453e631

Observation ff1d24d0-ab0f-4ea3-92d4-57d3d1841f88 · outbound

This paper cites On the evaluation of neural code summarization,.

Engineering AI Judge Systems On the evaluation of neural code summarization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.894287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.383721Z digest=sha256:1b695d5513f2ba9985a9787248fa8dd4ac8dfcc9dba7a7767c2b050d3415306d

Observation f1f2be3f-b8ef-4e6c-9a72-9c183b1bb0f9 · outbound

This paper cites RACE: Retrieval-augmented commit message generation,.

Engineering AI Judge Systems RACE: Retrieval-augmented commit message generation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.882388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.388673Z digest=sha256:092360ea5098ed09c2ab4e263244d6d4aa1f2045c0e0af9149742412304a233a

Observation 9de57533-72ab-447b-9c2e-9a16d1936911 · outbound

This paper cites On the evaluation of commit message generation models: an experimental study,.

Engineering AI Judge Systems On the evaluation of commit message generation models: an experimental study,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.870426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.392463Z digest=sha256:3bcbe0166a1985057fb28510b67d23caf3628e31c3317f66a1e0e5b84a1ecaf0

Observation 68bbccb3-869f-49bb-aa65-0f68b69ef797 · outbound

This paper cites Alpaca: a strong, replicable instruction-following model,.

Engineering AI Judge Systems Alpaca: a strong, replicable instruction-following model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.858724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.397190Z digest=sha256:80d2001e9cc4f858a107a2f516c92ceb02cf4cd1f44e4a800ad9e44db8f2b54b

Observation 196eee81-2147-4992-9d96-4f60defcddda · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Engineering AI Judge Systems Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.400788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.400788Z digest=sha256:77c784edd9a02778366ed9204252fd7d1fff071c418635a8c6c1dc56fb994bf8

Observation d64aa7a2-c591-4f8b-8884-5cb6100a9cc3 · outbound

This paper cites Synthetic data, real errors: how (not) to publish and use synthetic data,.

Engineering AI Judge Systems Synthetic data, real errors: how (not) to publish and use synthetic data,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.844266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.405510Z digest=sha256:596c265f0795b8ab4449d17708684f16fc49f07fb8e03eba951b14e840925281

Observation a7d2c3be-6852-4eea-b3be-9e62162e1d76 · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Engineering AI Judge Systems Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.409126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.409126Z digest=sha256:2733f70684fac470bb9b6f77b8d0a4fc03fcea33d30370c33a2d84f94bcf3b2e

Observation 02d0b65c-f5f1-4087-b51f-22bad0f786a2 · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Engineering AI Judge Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.413886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.413886Z digest=sha256:daa905ab4a2b21b8e0dcadea2b411f331c8fe082216fb8fb4b896d3139a772ce

Observation 5746f32f-31dc-40ec-b425-a3aa15be497f · outbound

This paper cites Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems.

Engineering AI Judge Systems Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.417892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.417892Z digest=sha256:e31abc4debc9a210ba31d1d680b6721182f13303a7e1bebcecbf5f3ae22934df

Observation 99eedc02-be11-4d1c-bb22-864eaef6e7e4 · outbound

This paper cites Fake it till you make it: face analysis in the wild using synthetic data alone,.

Engineering AI Judge Systems Fake it till you make it: face analysis in the wild using synthetic data alone,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.831987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.422862Z digest=sha256:d125ad169c892ce9c0e8cc7596c981d51c64476215205155097801ca338fb24c

Observation 219aebda-04d4-455d-b19a-6145c4cbea0a · outbound

This paper cites Commit message generation via chatgpt: how far are we?.

Engineering AI Judge Systems Commit message generation via chatgpt: how far are we?

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.820449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.427237Z digest=sha256:36e055e6c8eb0a5873cf6f08d69ad638fea349ee5206c89e127233134f17cae0

Observation f890a40e-3532-4f3d-b0f0-0e99dbe9cb2c · outbound

This paper cites Automatic commit message generation: a critical review and directions for future work,.

Engineering AI Judge Systems Automatic commit message generation: a critical review and directions for future work,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.808693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.431034Z digest=sha256:b2d77f78f4f6750d7c72d3effe2ec8c421433d6ca68361d682a6527993754b01

Observation b3e532f0-e1c9-45c0-8e81-6851a8f47262 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Engineering AI Judge Systems Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:59:18.795757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:59:18.434546Z digest=sha256:6bc9bb0e694aa5a386355b48922b3658f455b43a24094210ecca47a6ab0eff1a

Pith citing papers

Observation a1f08563-e1d9-450b-b1fb-942ab696aa28 · inbound

Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers cites this paper.

Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers Engineering AI Judge Systems

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:32:15.515399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T09:32:08.596259Z digest=sha256:7f621bb5a320c8dbcdb4ce3a300f4d20e6e71d3b174453aeccf9313cf271f0ad

Observation c3ec2d89-6bea-4e6b-8f3b-09f152b10195 · inbound

An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges cites this paper.

An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges Engineering AI Judge Systems

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:18.383835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T10:32:47.756343Z digest=sha256:343a8a023ced552f080248c475e20507f5f8eb82e25a3dfd7e6346635beae86f