Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:46:26.478418Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2601.22025.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:46:26.478418Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T13:52:28.587806Z
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0d4274b2-69b1-4e55-912c-3f70d875307e · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Constitutional AI: Harmlessness from AI Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fec83e7-30de-4b38-b3b0-dcd8c80cd51f · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Ragas: Automated Evaluation of Retrieval Augmented Generation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7edf6f-ede9-48e4-a462-222bdd539e28 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d21660b-26a8-4905-baf1-ca082d8d8a86 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Language model evaluation harness.https://github.com/EleutherAI/lm-evaluation-harness, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6ab860-6b49-4130-b8ec-0d2b137789e0 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications On calibration of modern neural networks.International Conference on Machine Learning, pages 1321–1330, 2017
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7fb54a3-5fe5-485c-b28a-39c729601e6e · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Measuring Massive Multitask Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b7beccb-dbfb-4fc2-913d-02deb0567a92 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5619446-54b0-4b95-b434-b9a6fe6be5bd · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Survey of hallucination in natural language generation.ACM Computing Surveys, 55(12):1–38, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2f44361-f2f2-46ae-a306-e09af85d1fd0 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Language Models (Mostly) Know What They Know
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec87f45a-0667-481b-ab1a-6f8785d22127 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Computing krippendorff’s alpha-reliability.Departmental Papers (ASC), 2011
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dede9957-29c3-4500-81c8-dbb3fa3cf791 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Retrieval- augmented generation for knowledge-intensive nlp tasks.Advances in Neural Information Processing Systems, 33:9459–9474, 2020
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a12c09-5ca5-4925-8609-f68c7a6a6754 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Holistic evaluation of language models.Transactions on Machine Learning Research, 2023
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937a9e9b-e375-45a1-a24e-186803836d2c · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Rouge: A package for automatic evaluation of summaries.Text Summa- rization Branches Out, pages 74–81, 2004
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c635d45-dfee-4172-8dd1-b51b48722f11 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b580fb96-f051-4051-9703-84073ec4ce2d · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications On Faithfulness and Factuality in Abstractive Summarization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6067e1-4326-4f96-939f-5a8c9c8cc09c · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa099e0-a7b9-496c-b7ab-4e2fad7d3189 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Openai evals.https://github.com/openai/evals, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff1e5be-7173-4e4b-bd59-7f86a4ad03dd · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Training language mod- els to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05ec9cd-5555-4d59-9cce-b93d2cab1d80 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications LLM Evaluators Recognize and Favor Their Own Generations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d0f27fb-637b-4b23-b91b-06d707a3b1a9 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Red Teaming Language Models with Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0377750d-6f53-440b-9e31-f8b4592ec00d · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6048266-81af-4ed5-9c79-d42da57ac594 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b1f316-d9e5-45ad-af54-0cc5541e8488 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications BLEURT: Learning Robust Metrics for Text Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b1f8f9-cd37-4b3d-9806-07d0c9eaf17f · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554301a6-be59-4a82-9304-0c40a860e416 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Large Language Models are not Fair Evaluators
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3967a47a-f691-4b57-b57c-860e8a86fd68 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications BERTScore: Evaluating Text Generation with BERT
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e028793b-e67c-49a0-aeb8-c55dc563c994 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0495afe7-d37d-4a4b-8fe3-b7e7c29526c7 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69b204c-5028-4705-99bf-df87c86cbc5f · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ecf084-a41d-4aef-9793-e71613aa2689 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b93a26-7363-47dd-b7a3-179827c81889 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedd61d1-4cff-465a-a81b-48a554abcda8 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 333ea756-0285-427a-86d2-16c488817dd4 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4d3492-84ed-4b7e-84c5-1e7563b2f7f0 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221e6b5e-ca51-41ce-8a36-f075d138d4e1 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications 38 A.5 LLM-as-Judge Checklist
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23232521-b3c1-489a-9270-047535cf5155 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f2671f-e169-4b06-9c3a-df1a8b0b211a · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27e39043-40f1-4236-acc4-0665905b6fe1 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585f1f71-b53e-4979-a2b2-ebd1bc7f65a5 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f0d459-79b2-49dc-badf-5834603dbe0a · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9b374fc-ed31-415c-83aa-0682f972ba28 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d8ffb2-91e7-4a5b-81c7-56d5f0af33a3 · outbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e0dc3f-31bb-4b50-9c3b-b8108b784197 · inbound
Mi-Memory: A Lifecycle Memory Framework for Personal AI When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.