Pith. sign in

Paper Citation Record · LEDGER

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild

As of 6 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 0 inbound Pith citation observations for arXiv:2604.16304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16304 v1

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T11:26:38.634540Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

76 of 76 outbound references displayed

  • verified exact31
  • verified fuzzy6
  • unresolved23
  • parse uncertain0
  • malformed identifier8
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80e341ab-260a-428b-9938-e79d6df64d11 · outbound

This paper cites 2001.Introduction to measurement theory.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild 2001.Introduction to measurement theory

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:27:48.496686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:2ffdef86443cd9507d45483725997bf5dfd101039ced5c9560695b90f9b5e774

Observation 8933a2d6-0df9-4728-9308-db5098a0cf22 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.453775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:3d257305d0478277a03fdc63ede4f2e211477f027b19e60cba0dd2d92f7cd373

Observation 851e59cc-2593-4d09-8730-60a0d96bab1e · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.509630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:44680472d2b7c7d52d23c3d65c0d3e6971edce1af120842570aab7eb9cca52b2

Observation cd535416-a21e-41e0-9f10-63735a5f3287 · outbound

This paper cites LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks

Reference 5

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.904536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0d89153696942eba671bccff3c6d292ced6d72e9791d3959af8d02ffd1211e38

Observation b33a07bc-e042-402b-bafd-39b36c57c02b · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.507918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:a37e27284bbc86612b2ded237ca2112b99559964a2e8f41f68365dc74ad6b38b

Observation 077db636-8303-496c-aea2-f7cc50494ea7 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:44:50.211576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:5952104fa76d028342dab05e197aed4d1236ac84bedc1d1c5649ecfe13862c93

Observation 9372f1ff-ce68-4137-a29b-9dc82731c75a · outbound

This paper cites 2022.Thematic Analysis: A Practical Guide.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild 2022.Thematic Analysis: A Practical Guide

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:27:48.483982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0e276cabff60c8bf074347f7d0c1713a24c2823f6227639558994e44c07d3769

Observation 8b667205-3b4a-43b0-958a-96527d2cab0c · outbound

This paper cites Yu, Qiang Yang, and Xing Xie.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Yu, Qiang Yang, and Xing Xie

Reference 9

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.878496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:df5fb5a4918b68b347c09847870cf4191147557fa9b2bef29c8e022c96b6f60e

Observation f065e130-ba71-43de-8fa9-6c8a0cfac72e · outbound

This paper cites All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.159299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:f349050452ce95dbbd80c6231112026a845ef62cd409ea4a37811788cef2e97a

Observation b84cc13c-78a3-4706-a23b-b4c6dfb83173 · outbound

This paper cites Collins, Albert Q.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Collins, Albert Q

Reference 11

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.834809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:1bfabf7303c7921f34c8b59f86a0245f4a3310e6c80d197796e630d8b5bb24ba

Observation ba750c34-295a-404a-831f-a388cd25c403 · outbound

This paper cites Bennett, Gary Hsieh, and Sean A.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Bennett, Gary Hsieh, and Sean A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:27:48.488384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:be9dd7b9144993963cc9bb0d644ef783adb404b37de353542554cfdcdc082284

Observation 8b1fa7d3-34e7-4c4f-a915-3acf674e0c35 · outbound

This paper cites The WyCash portfolio management system.Proceedings of the Conference on Object-Oriented Programming Systems, Languages, and Applications, OOPSLA.1992;Part F1296(October):29–30.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild The WyCash portfolio management system.Proceedings of the Conference on Object-Oriented Programming Systems, Languages, and Applications, OOPSLA.1992;Part F1296(October):29–30

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.864303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:899cc3d266cf499d9e78f2f2ac0494443e78730e35f020792b4b8b78840a7487

Observation f5df0b8b-740b-4758-b416-7d3930930720 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.490381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:660075009b1ddf4dbac89b5224e25c7668726e2b4e58cf0577c25aea6b746d7d

Observation c1c6f0f1-d791-4621-a60a-03844a4c333a · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.492227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:e021b673398b914692390ed0d86beecb7a8a5a0663312e26fcdf1b7eeb1b2c92

Observation 1800c2d2-f378-4ad2-83b4-64e245562f74 · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Towards A Rigorous Science of Interpretable Machine Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:27:48.162805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:d290a8ca7cc096bfe6c57b7e9682952998882775f01944fbce3b29988c670349

Observation 867b3613-c244-4fde-8aea-314055d7f3ef · outbound

This paper cites Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y

Reference 17

Resolution
malformed identifier
doi_truncated, observed 2026-05-16T11:27:47.845342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:e0b98506f63597837c644811ec9643aa2e432e67651c91dc2b8e7424b87d4dc2

Observation 642437ac-0201-4472-a60f-bdd951fd3cf2 · outbound

This paper cites An extension of HybridSynchAADL and its application to collaborating au- tonomous UA Vs.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild An extension of HybridSynchAADL and its application to collaborating au- tonomous UA Vs

Reference 18

Resolution
malformed identifier
doi_truncated, observed 2026-05-16T11:27:47.874813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:f7d06f099e7d6f5dede5a78a2aeaaa57b4f69022ee0b4f3110217d2d795f35bd

Observation 68611810-516a-47da-a2b2-e4c4faf8dfdb · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.504268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:7916d8f08454f1ff459455267f69387f7dcf4085d25c2c3244c6c8f281278654

Observation 8e8082fe-c8f1-41f5-a457-ba29a69b883d · outbound

This paper cites Gallagher, Jasmine Ratchford, Tyler Brooks, Bryan P.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Gallagher, Jasmine Ratchford, Tyler Brooks, Bryan P

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.893352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:b66f5edaaf5e525e483695abf873e3bfe15830c8ed0b02eaeb55f3ada5217897

Observation 5ea49024-9aa2-41c8-929e-ea11e617409d · outbound

This paper cites Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text

Reference 21

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.867364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:69e306f0c9947dc541e86e9a2eb0e6d2ee3566c4a6db6348d468a1c3f7e967e9

Observation f31a23bd-c6bd-43b0-889f-6b740966aec8 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.482170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:47fc3732fe6d62306845ec45a8f3cd66a4f04aa72aafc6537e77bb1d9e714a96

Observation dc70fb52-17e8-4e3c-a259-cdd0d716fc8c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Measuring Massive Multitask Language Understanding

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:27:47.852132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:923d0b454becb09e49bd78275a4782583b3f7cecc5f7189044efa9d3eb6afaf6

Observation 2d582065-7ce6-4686-9177-cc608f00d943 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 25

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T11:27:48.169940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:7c88b812484ffda6048a5115508fdcd66aa1c34629e79a29ccfe1f07d0adb763

Observation 634a7496-4a61-4b54-9ae7-01197c2b8530 · outbound

This paper cites Towards interactive evaluations for interaction harms in human-AI systems.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Towards interactive evaluations for interaction harms in human-AI systems

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.886634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0f8c55de4f2ab34794ab75caf78213d9436521023d90689ac34f65eb16ca0e39

Observation e9667621-0abd-45e1-929f-6d6bbfea9c65 · outbound

This paper cites Jacobs and Hanna Wallach.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Jacobs and Hanna Wallach

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.814181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:354450baacd43c9f7570f6b1105f0c6bcf1d36751c958572ea2a40c01239afdc

Observation 7677e608-811f-490f-8044-63846081f68f · outbound

This paper cites Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

Reference 28

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.896651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:d57fe221a9562f4efd967119054ff8f23611a6c8e0792884648c785368d0c5af

Observation 2c83638c-126d-43d7-bff1-6d1e563fee52 · outbound

This paper cites Evaluating Human-Language Model Interaction.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Evaluating Human-Language Model Interaction

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:27:48.177470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:4e4566323865b41b3d0d20fc93a25618f95ed1a0ad6229ded4806bc53853fd11

Observation b3859dfb-7c22-4824-8991-0fcfd778ef64 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.478434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0d7bdbff78862b350d8e752568e7ba8d54474ca9358fcf41f6cc687e6c3ce363

Observation bfd5ece7-7118-4f4a-8ab9-41bfc0c926ad · outbound

This paper cites Holistic Evaluation of Language Models.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Holistic Evaluation of Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T11:27:47.839549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:c91766347e89e8c6392b42d8a6603c63c9687682af889c19462df8727b0cf8bd

Observation fe8eb688-3b1f-421d-afc1-47a3f38d2548 · outbound

This paper cites Rethinking Model Evaluation as Narrowing the Socio-Technical Gap.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:27:48.140243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:8a519c47d6006317a328c25d1b4bd5f7c9bece0048ec6a0f6f51ba31fb6e4281

Observation 853f6ce4-547b-4940-ae04-1c5c627994a4 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.494327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:6d89b9dbf0cd9988bb3e6ae568f824191776e8f8b7a63e0d79b2640f60118d01

Observation 001f13f0-6994-424d-9679-3127905ec226 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.899203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:342fcafe97d865eba1c19710b63bb851bcfe64055db7abb86267eab973b3e90d

Observation e6bcc486-f2a1-49ba-9502-266cba625410 · outbound

This paper cites In: Zong, C., Xia, F., Li, W., Navigli, R.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild In: Zong, C., Xia, F., Li, W., Navigli, R

Reference 35

Resolution
malformed identifier
doi_truncated, observed 2026-05-16T11:27:47.848306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:ce332f22192da2461aec4a8e2662025465e337b6e2c961e036e23214456dd8bd

Observation 333773b1-c3c1-434b-95d2-a90fbc04dafe · outbound

This paper cites Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.133595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:275f11eb8e45b94de91ca732fbdbc55b7e876ba90b50aeb61202dadc89326c4b

Observation 57d23f4c-4081-458d-8ae2-068ccde76639 · outbound

This paper cites Vera Liao, Alexandra Olteanu, and Ziang Xiao.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Vera Liao, Alexandra Olteanu, and Ziang Xiao

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:27:48.474495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0ee2415412996a03ee5f0042ba92932a6ae29394a708fc904e4f8b4e4e7371af

Observation 27cc4289-3688-47fe-a0f3-1557c1baa35e · outbound

This paper cites Welcome your new AI teammate: On safety analysis by leashing large language mod- els.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Welcome your new AI teammate: On safety analysis by leashing large language mod- els

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.871460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:504272853724fbc7252a50202612ed6c482f2e8e028ac5c05454ff0494484568

Observation ac8613c3-61c6-4564-9e12-62af4a5afc04 · outbound

This paper cites Insightai: Root cause analysis in large log files with private data using large language model.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Insightai: Root cause analysis in large log files with private data using large language model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.787309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:508cc185ae1bb08c91405d76da7527f103dcd26fdd525d9daf2569d601d91dc9

Observation 93cfc402-2e07-4269-8275-7d3c666342eb · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.476622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:875804fbd13b2058b31c79662c54ae339fc618d189577f7b5dcc9e11a45c897e

Observation 7fb0825c-fd3e-4972-8f08-05234f5e7fa3 · outbound

This paper cites McIntosh and Teo Susnjak and Nalin A.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild McIntosh and Teo Susnjak and Nalin A

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.826375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:c823a35e4805762517e45431bf5c40d83d09217a21a3c3428a9170b7258d3770

Observation 27063029-dd91-45fb-8585-cf3d24a224c8 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.470277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:8d29648535d8a5576ad8995cfda9c586fd1ebb2b61bcaea8d75ec757bd5c8a97

Observation fbddf141-cf64-494d-948c-6e7632a1ba2f · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 43

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T11:27:48.472442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0c9b4c681a5f6abe83aa7cdf5f609b9d38d258cf98c2b336e14ea63da092673a

Observation f0848d9d-3205-4797-bc79-af2aa97ef347 · outbound

This paper cites In2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP).

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild In2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP)

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:27:47.856810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:4c8b04a1bba97ebe831386727a41b1366308f3b33fcb0f9418eb7dcf4ca4524c

Observation 566e65d9-87a1-48b9-901d-a80e896c55e6 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.511279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:ead537d4c2d93c8f231aaaecfba114434ba2d27f10a65154dce188520cdc05db

Observation 218778fb-06fb-4b8a-941c-aa30ea9dfed1 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.451590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:8c92613ef85ed6af743ad892f0988d7097e3c9eb8cea831ed52ea32ec3adcef0

Observation 75230a94-7698-4b96-a6b1-4f25dd1cc06a · outbound

This paper cites Human-Centered Design Recommendations for LLM-as-a-Judge.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Human-Centered Design Recommendations for LLM-as-a-Judge

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.184935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:0806916b1c6c3b1072559619d53371ef778f5997e62e40af6b56defa98d3e0fb

Observation 297542bf-2184-4d8a-8a04-f51344b29ff3 · outbound

This paper cites Jianming Chang, Songqiang Chen, Chao Peng, Hao Yu, Zhiming Li, Pengfei Gao, and Tao Xie.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Jianming Chang, Songqiang Chen, Chao Peng, Hao Yu, Zhiming Li, Pengfei Gao, and Tao Xie

Reference 48

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T11:27:48.155468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:a254980d5e449a98a75a5d253d7f988019ff3e69ebe4f2d48fd77c3db9643d78

Observation 13c8994e-228b-4a65-826a-a2d2d8fdaf64 · outbound

This paper cites White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:27:48.137236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:c66a5b3fbe3d820018f6da57de21936cf79eac6bac1df5bbfc6359d841bad785

Observation d346d19b-b601-4c32-abc0-d40d1b223dcc · outbound

This paper cites Roedl and Erik Stolterman.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Roedl and Erik Stolterman

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.809588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:461a483ba1c0ea1eae7c49f935e1441c1aa28684a1d81e7b2e4c88a89051ab70

Observation 6aecd7c5-8c64-474c-aa7e-3604f74b1899 · outbound

This paper cites The user experience of ChatGPT: Findings from a questionnaire study of early users, in: Proc.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild The user experience of ChatGPT: Findings from a questionnaire study of early users, in: Proc

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:47.821769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:b84bb7ddc55652d7c525b1083f4a8bd4c7b268c736ae534bb6e0d582951e0d6e

Observation f2e49c1e-2fa3-4470-baae-0463ec1a19ff · outbound

This paper cites Emerging roles and relationships among humans and interactive ai systems.International Journal of Human–Computer Interaction, 41(17):10595–10617.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Emerging roles and relationships among humans and interactive ai systems.International Journal of Human–Computer Interaction, 41(17):10595–10617

Reference 52

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T11:27:48.166766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:a56b3e5c6a829012a5fd375ffbf41cbdfd19f0f7162194cbe7c3b894a6e74a53

Observation f4010a13-58e0-44f0-8430-e0057f263945 · outbound

This paper cites Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.147823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:73227e0440696be91f4ebcddbb12d0019291a154941b480c5dabf00b802784a6

Observation 0871d58b-932b-4134-ae2e-06c0b6dbb904 · outbound

This paper cites Shergadwala, Himabindu Lakkaraju, and Krishnaram Kenthapadi.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Shergadwala, Himabindu Lakkaraju, and Krishnaram Kenthapadi

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:27:48.465991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:f29b589c730df83e824b5a39f09496093cbc5370d02497537495912cdd7e5347

Observation dd7e5557-0e1f-4db9-87ef-50de955665f5 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.467993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:3c26bc07be91f7174b12f4bee75290d57077762054f04e2a32d039e9f8c1cb85

Observation 810bb39f-09bc-441f-9959-2707bf7a9c9a · outbound

This paper cites What factors might be causing the significant deviations in my circadian rhythm patterns over the past 30 days?.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild What factors might be causing the significant deviations in my circadian rhythm patterns over the past 30 days?

Reference 56

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.816958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:24543ece5e46a80989f3b45af8cd411852a8f214735e80d3a5adb1bac5db5ebc

Observation 09f64f09-602a-467d-b4fc-b9500cd8316f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 57

Resolution
malformed identifier
local_arxiv, observed 2026-05-16T11:27:47.797554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:9ceba5efff168ec43f2050836b72592bb860a0e699fc35c7a16441a51f97ddfa

Observation 4d4c71f7-8d3b-4242-9ea4-84eec71e99a0 · outbound

This paper cites Stolyar, Katelyn Polanska, Karleigh R.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Stolyar, Katelyn Polanska, Karleigh R

Reference 58

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.881004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:9c59d18db935cddcc8c88496e7f78cc167ab1001afa556da6cd6af1efbddf0bd

Observation deb2eb5c-4aa6-4b73-a29f-4af409cb3453 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.455510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:a21c2661af750d1349bfcc61abb50dde567e7c3d79b894d1889821cddbec036b

Observation 992133ee-f6b2-402f-8aec-2bc377c8635b · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.461524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:a47d18d69eeadd9e8ccadaf214fae8b525798345a99566e49b578fb425ce2f22

Observation ca2c015b-5ab0-4fed-9b85-141c840b6543 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 61

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.883237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:ab2417da5f032f4f7abd49065379a62e222f3b91e92b182a4c5993c792556a4f

Observation 25e632c4-0b85-4311-a715-4bc877246762 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.459656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:383f37845011dd1bde792542058450b1e8e8410d0224487b62569dce27724537

Observation 09bc4aa2-72d8-479d-853b-731026a23d6b · outbound

This paper cites Proceedings of the 2018.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Proceedings of the 2018

Reference 63

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.802686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:650b260b54badc3750679b0fb07d6f0ffabacb95e5c23d97a60d7893fd4eef16

Observation 74507c95-25d9-4b6a-a01b-f6e7661780b6 · outbound

This paper cites Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T11:27:48.181459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:6475a128f13483bc15c1888f7f0f92245ddcc12eb2e58507ed5345efe10b563c

Observation 167c89ad-4491-4e22-b6a5-37350ef5bc41 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.506150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:8fd2512a68c9a8f0146a40c292e19f418b94a5e5338daf622a43e4ceabda1bc9

Observation f5951944-213e-44ba-bd14-2857391ddc4b · outbound

This paper cites MacLellan.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild MacLellan

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:27:47.832098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:d668edcc898d6b4d92b506285884854b8df1751b56cf096d3e3acc6ff5b7a930

Observation ca0a90b5-20ac-4025-ac26-bc4addbb0079 · outbound

This paper cites A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.173597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:1d62b4f5991d30693152be65ceeb1e43731ae197d4f7d01388dcfc2caf30317e

Observation 74fcc166-da4b-4d85-933d-994eb24b6569 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.457454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:a9ff49f638723a9e64bf5cb5580c6d7b75de17708063775719f0c3bcae6bfcfd

Observation d78b3bbd-e1c3-46ce-aaf5-4b5b5e18a3dd · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.480383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:95aa1274d7f3d13aa636a932133f32bd2de61c520492325faced6a9278a3a18e

Observation f0167c0e-0463-44fa-a1b8-6bf56655e503 · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Sociotechnical Safety Evaluation of Generative AI Systems

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.144345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:261f534395cca9e4af54d4b3df7e3be2fbeeb7fd4376fb956f0d2f853273a886

Observation e7312317-cefb-4177-bf4e-91e5db0b9a18 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.498973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:1fce0866b5655ed8d9d45ae114e7405a4843fcc08d3bdc48f4229f4f00f4edd8

Observation fa00ad77-93ce-4e73-b336-d056a3b35f51 · outbound

This paper cites Vera Liao.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Vera Liao

Reference 72

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.901787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:15991cae9d6c1d68e822a1e8a5a57000a40f57d157414dea47f6a38b24d0a5e6

Observation 0510db8b-0d87-476b-8657-23546a52059e · outbound

This paper cites Hashimoto.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Hashimoto

Reference 73

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.889515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:e1e11ea3796e19e58d7bd2e274ae9debdc79335c0127981d352e4e3997e779ff

Observation 5eac3c4a-458e-4c6f-a953-2af68497b5d0 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.463651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:4b1340ac9e2d2d38ce1f4367f224e4fd88bf2fceae34a5a15ca7f078a80e67d5

Observation 4bac9148-1a9f-4b4c-83a0-f9f3c95bc09c · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T11:27:47.860553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:ad249bbb1db5e76ded8fd03e60759bf3b5e18863c18d72decd88e78c02d812ea

Observation 4a5f1225-deb4-4932-81c5-da2e02e5226a · outbound

This paper cites Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

Reference 76

Resolution
verified exact
doi, observed 2026-05-16T11:27:47.842619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:4fd7376f03e96facfb3c4d2047671d6f5d4c7f8345978e2f2e06ee4e56f18d3f

Observation 274674f2-c35f-45b2-b5b7-fc26280b0466 · outbound

This paper cites an unresolved cited work.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-16T11:27:48.486248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:384ba5c084f70bb7e0b615dac7580d00ef1b545158210a755b9d4e6d0de28401

Observation b2324964-4818-47e1-ad9b-514c271ced54 · outbound

This paper cites InProceedings of the Thirteenth International Conference on Learning Representations.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild InProceedings of the Thirteenth International Conference on Learning Representations

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:27:48.501760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:bca9a9e86667c2b1976bfa9958b1373df9226350466699e1ecd64bd78fee4dfa

Pith citing papers

No inbound Pith citation observations are available.