Pith. sign in

Paper Citation Record · LEDGER

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2601.22025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22025 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:46:26.478418Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:28.587806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d4274b2-69b1-4e55-912c-3f70d875307e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.181481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.181481Z digest=sha256:29291d9150125a9b54aac9d5a4721497cde8668d2cb57a4f04c903c094599d3c

Observation 6fec83e7-30de-4b38-b3b0-dcd8c80cd51f · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.193439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.193439Z digest=sha256:137965a48da5aaa0381c732269800e7b084fb1e08d804aa8ced22cfcabccceeb

Observation 1d7edf6f-ede9-48e4-a462-222bdd539e28 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.200818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.200818Z digest=sha256:d824dab905848029b13b81b241db3f8843d2a515cfe35314eebc344e428ce148

Observation 7d21660b-26a8-4905-baf1-ca082d8d8a86 · outbound

This paper cites Language model evaluation harness.https://github.com/EleutherAI/lm-evaluation-harness, 2023.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Language model evaluation harness.https://github.com/EleutherAI/lm-evaluation-harness, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.207610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.207610Z digest=sha256:ca1d440bb1eaa68913702c4677bea0ee4ba8f2c1d1d0e7a78b66fb9a54199105

Observation fa6ab860-6b49-4130-b8ec-0d2b137789e0 · outbound

This paper cites On calibration of modern neural networks.International Conference on Machine Learning, pages 1321–1330, 2017.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications On calibration of modern neural networks.International Conference on Machine Learning, pages 1321–1330, 2017

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.212907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.212907Z digest=sha256:55738adab26bae09a96889de1b558378b9dbd9d66a2cdab0a2408143a0f307bc

Observation d7fb54a3-5fe5-485c-b28a-39c729601e6e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.218264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.218264Z digest=sha256:c1db0edbf78bd55af23939fd4ad0a24c5ec8ac0d74794dc1b4e3085e600f0a55

Observation 1b7beccb-dbfb-4fc2-913d-02deb0567a92 · outbound

This paper cites Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.230145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.230145Z digest=sha256:8a21676ab6267bb03ef15cb200339225942df20f839e2a0b215e733ce0c2bbf6

Observation b5619446-54b0-4b95-b434-b9a6fe6be5bd · outbound

This paper cites Survey of hallucination in natural language generation.ACM Computing Surveys, 55(12):1–38, 2023.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Survey of hallucination in natural language generation.ACM Computing Surveys, 55(12):1–38, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.236101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.236101Z digest=sha256:4a9d5435482cfdbe077fb1e120ed7896e536f647db119de1383cf0751f66897e

Observation e2f44361-f2f2-46ae-a306-e09af85d1fd0 · outbound

This paper cites Language Models (Mostly) Know What They Know.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Language Models (Mostly) Know What They Know

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.243906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.243906Z digest=sha256:f3b717a6642dc1df556132987692c394074aa8e0be55c52bb0885a019f55c3b1

Observation ec87f45a-0667-481b-ab1a-6f8785d22127 · outbound

This paper cites Computing krippendorff’s alpha-reliability.Departmental Papers (ASC), 2011.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Computing krippendorff’s alpha-reliability.Departmental Papers (ASC), 2011

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.252662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.252662Z digest=sha256:7aba0b992b6b1a2f812960b7efeb2f48d3773fbc9083485fc6a592eeb6f4c544

Observation dede9957-29c3-4500-81c8-dbb3fa3cf791 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.Advances in Neural Information Processing Systems, 33:9459–9474, 2020.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Retrieval- augmented generation for knowledge-intensive nlp tasks.Advances in Neural Information Processing Systems, 33:9459–9474, 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.259826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.259826Z digest=sha256:f8687cc7e52068e2a29a88210b6245255aa75438caecc4f97234254fcb063237

Observation 48a12c09-5ca5-4925-8609-f68c7a6a6754 · outbound

This paper cites Holistic evaluation of language models.Transactions on Machine Learning Research, 2023.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Holistic evaluation of language models.Transactions on Machine Learning Research, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.268546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.268546Z digest=sha256:2d0a033d7e0a5cd9a62aca6fa14113a5155456ee5f648206a83ab4079c5f37f3

Observation 937a9e9b-e375-45a1-a24e-186803836d2c · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.Text Summa- rization Branches Out, pages 74–81, 2004.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Rouge: A package for automatic evaluation of summaries.Text Summa- rization Branches Out, pages 74–81, 2004

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.278646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.278646Z digest=sha256:6d2de1b09fa4f59eae517392483601b48476cb95137d8845038f31dd0d8285de

Observation 9c635d45-dfee-4172-8dd1-b51b48722f11 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.284150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.284150Z digest=sha256:2efe31128dfc3ea83f21dae574411479c4d0f5afe363858b28199d403bbe6978

Observation b580fb96-f051-4051-9703-84073ec4ce2d · outbound

This paper cites On Faithfulness and Factuality in Abstractive Summarization.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications On Faithfulness and Factuality in Abstractive Summarization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.291385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.291385Z digest=sha256:a2379425a5a3047b4ed9a5ea0b1450e19099e70806c0a9bf6c7c2a6a9d5b85b5

Observation bf6067e1-4326-4f96-939f-5a8c9c8cc09c · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.296715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.296715Z digest=sha256:d654b686e4733d4af91e61323bdb8484e35dd9a12dd9ef5c4a20b7c99e1e63bc

Observation 0fa099e0-a7b9-496c-b7ab-4e2fad7d3189 · outbound

This paper cites Openai evals.https://github.com/openai/evals, 2023.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Openai evals.https://github.com/openai/evals, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.305008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.305008Z digest=sha256:00695dfafb91f467253ac52069c39ba130c2f4c67c28464a35afe8713062b0ba

Observation cff1e5be-7173-4e4b-bd59-7f86a4ad03dd · outbound

This paper cites Training language mod- els to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Training language mod- els to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.313308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.313308Z digest=sha256:6f2b775b3fd17ed27a741736e871ef8b67e5f3ebf85593df43fce5c9dadfa99a

Observation e05ec9cd-5555-4d59-9cce-b93d2cab1d80 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications LLM Evaluators Recognize and Favor Their Own Generations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.318507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.318507Z digest=sha256:c60ae6e3ad4b53bbfc7e18803e4cd9a9b51da9777118a4313598112bfe687a93

Observation 4d0f27fb-637b-4b23-b91b-06d707a3b1a9 · outbound

This paper cites Red Teaming Language Models with Language Models.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Red Teaming Language Models with Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.324859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.324859Z digest=sha256:14c70de097b6294fcb346b3124d94fc1912dc704c581fc171a80f99dc8dd6382

Observation 0377750d-6f53-440b-9e31-f8b4592ec00d · outbound

This paper cites ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.331261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.331261Z digest=sha256:64764f3fa1cdad72ab01e6dababf0e34f143fcada767d9ca4ee2df8c27cf398d

Observation d6048266-81af-4ed5-9c79-d42da57ac594 · outbound

This paper cites NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.337035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.337035Z digest=sha256:ee896227a6f1872da511b56bbf327daaa2562f83be66de9abd893147db497650

Observation c8b1f316-d9e5-45ad-af54-0cc5541e8488 · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications BLEURT: Learning Robust Metrics for Text Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.342276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.342276Z digest=sha256:c01de76804d2173612d9687b87475904d25d242615b6c06ed4d65b25b28a0879

Observation 52b1f8f9-cd37-4b3d-9806-07d0c9eaf17f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.348309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.348309Z digest=sha256:6182fc7ef12a6df5f341a763f3b8f94b77963d2bb62a65a0326f6519639e0c72

Observation 554301a6-be59-4a82-9304-0c40a860e416 · outbound

This paper cites Large Language Models are not Fair Evaluators.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Large Language Models are not Fair Evaluators

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.353366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.353366Z digest=sha256:882c21f030befc6d3431ca9d4427b84ddd3a979c3a382e3065274c1e37b0ca9d

Observation 3967a47a-f691-4b57-b57c-860e8a86fd68 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.361211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.361211Z digest=sha256:843276a5e5adbd7125021feb00c9a4159985f92da0fda3bc3772c6f7e0731c31

Observation e028793b-e67c-49a0-aeb8-c55dc563c994 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.366420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.366420Z digest=sha256:6e377e12a6f594eaab8f69a3bd8c95399af91a868084ac732a3f04bf41b744f1

Observation 0495afe7-d37d-4a4b-8fe3-b7e7c29526c7 · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.373904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.373904Z digest=sha256:c39e882135ce5e07e13b845f8b9386366329cb226963d98c125462a453158932

Observation f69b204c-5028-4705-99bf-df87c86cbc5f · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.379891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.379891Z digest=sha256:116e6cfaa7ee8efd236fbe24b71f6a8306cc6baccbf126bba700eb3dbd65eb20

Observation d1ecf084-a41d-4aef-9793-e71613aa2689 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.391383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.391383Z digest=sha256:dc372733c83f3dafd85694038bb84fb90e9a7c49136ba4f82f29665f09d35ef1

Observation 81b93a26-7363-47dd-b7a3-179827c81889 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.402681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.402681Z digest=sha256:afb86cef2a6c705035a8466d7c4a4c1f63d957701cbe59818fc9927748e8d616

Observation dedd61d1-4cff-465a-a81b-48a554abcda8 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.410237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.410237Z digest=sha256:fe2fbc6fc7b7262f6a027ba8b6458c8803b67fc03096ce61beca2d5e67c51d56

Observation 333ea756-0285-427a-86d2-16c488817dd4 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.418141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.418141Z digest=sha256:989aeb95018fc48539c8b7ac85daf2179a1d8fcaf4e8696bdca2151c1a2f5298

Observation ad4d3492-84ed-4b7e-84c5-1e7563b2f7f0 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.425217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.425217Z digest=sha256:d23cb675ab5e89c01e20c3a362a7a91c786b8423a659f05c42cf9c6c6123ecd5

Observation 221e6b5e-ca51-41ce-8a36-f075d138d4e1 · outbound

This paper cites 38 A.5 LLM-as-Judge Checklist.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications 38 A.5 LLM-as-Judge Checklist

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.435130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.435130Z digest=sha256:9b246fb526d2fc65909ce9879eb690db46464aa239459284c3d57ff1ad630478

Observation 23232521-b3c1-489a-9270-047535cf5155 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.444650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.444650Z digest=sha256:845dc3574e4e49b4a3b3fec57bf7bf5ff6879ecab3167d93d7b9e0eea297370b

Observation 52f2671f-e169-4b06-9c3a-df1a8b0b211a · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.451123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.451123Z digest=sha256:14249fdd39a6a26daf19813cee4bb28238ad0c23fc74c45675c0576d92d59a7d

Observation 27e39043-40f1-4236-acc4-0665905b6fe1 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.456928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.456928Z digest=sha256:a566f6c417b45cfde985a929f35c7361a8da1cde5d0b5ad7932d8cea1352b32e

Observation 585f1f71-b53e-4979-a2b2-ebd1bc7f65a5 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.463034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.463034Z digest=sha256:4906bd3b520e75d0733bea350df01ffcce75e4430986cdac20620c65809bb634

Observation 38f0d459-79b2-49dc-badf-5834603dbe0a · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.467732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.467732Z digest=sha256:283760eecb014af442459e5888cfdf61b92074c7e481809d7556dac320ef4236

Observation c9b374fc-ed31-415c-83aa-0682f972ba28 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.472653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.472653Z digest=sha256:567e13f64a74eff424406ebfc7515786c09bd074918f3ee6baaadbfdebadd57f

Observation 25d8ffb2-91e7-4a5b-81c7-56d5f0af33a3 · outbound

This paper cites an unresolved cited work.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.478418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.478418Z digest=sha256:a66cab92aaaa024c112d5aa6ec8f838880df3e53e11daa140e01f54f9d63df32

Pith citing papers

Observation 56e0dc3f-31bb-4b50-9c3b-b8108b784197 · inbound

Mi-Memory: A Lifecycle Memory Framework for Personal AI cites this paper.

Mi-Memory: A Lifecycle Memory Framework for Personal AI When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:28.587806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:28.587806Z digest=sha256:57393912176cace9e4fb20d0aab88ef9885e5f9ca4d4c92b5aef34f1dc8a567b