Pith. sign in

Paper Citation Record · LEDGER

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

As of 6 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.21010.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21010 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:46:17.594637Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ba862811-772a-4f4e-b97d-8401b36248bd · outbound

This paper cites A comprehensive survey on legal summarization: Challenges and future directions.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers A comprehensive survey on legal summarization: Challenges and future directions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:12.710693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:12.710693Z digest=sha256:5c16a6cfa7f7d9e459f05f0157d37be1641ec1543241f5d4dabaa7032942319c

Observation 4517b5ee-256d-4809-98e2-046cbb1cefde · outbound

This paper cites Large language models robustness against perturbation: S.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Large language models robustness against perturbation: S

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:12.775600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:12.775600Z digest=sha256:a484d70e8e3ecada0ecccdc2eb313ab70e75e1b09067b92e95c737a78228a3e2

Observation 79f977dd-d613-41a7-929e-e0a1eac5cea5 · outbound

This paper cites Using llm (large language model) to improve efficiency in literature review for undergraduate research.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Using llm (large language model) to improve efficiency in literature review for undergraduate research

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:12.857316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:12.857316Z digest=sha256:14e8cf03216a7fad56ab8236fb3b8e7cd9d854455ac77858bbbe5778cc1d24d6

Observation 4c16b665-e324-4179-a1d2-9e266f785d28 · outbound

This paper cites Robust tests for the equality of variances.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Robust tests for the equality of variances

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:12.918614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:12.918614Z digest=sha256:d9458ca7d660f9a5480a4fcf4bf47d99a1d0ec59dbb5100372d4e87dd9ea1f30

Observation 7d7d2ada-1917-4f51-a68e-cdf0e061cb7b · outbound

This paper cites Efficient inference for noisy llm-as-a-judge evaluation.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Efficient inference for noisy llm-as-a-judge evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.001406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.001406Z digest=sha256:04abd9f4a57ab8bf7e25a78cea792bbf3d1f5476e68a631b22ee67ab22e0ae87

Observation 5d7adfc0-191e-41c0-84f9-cc58b96d029e · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.056939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.056939Z digest=sha256:d729833488fbf987cd8345c610efbeea72e6d2b07c4317aa67d1640d271082d1

Observation 419922fc-7cf8-43b4-83f8-35130c2586e0 · outbound

This paper cites Statistical comparisons of classifiers over multiple data sets.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Statistical comparisons of classifiers over multiple data sets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.137024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.137024Z digest=sha256:494647c560d6bc4f67ab5a2bdb8a462c0e9d5161aa35be0ad3eb44694f008cd2

Observation 35c44920-4697-49da-99a8-8f8c807b6887 · outbound

This paper cites Applicability of large language models and generative models for legal case judgement summarization.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Applicability of large language models and generative models for legal case judgement summarization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.242928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.242928Z digest=sha256:5cf50d2ca11426fc1c412c664565871615efdb9fc23aca1320c6f8499a305ccc

Observation 34c92790-8a86-4ac3-94c8-19b2a16735ec · outbound

This paper cites Explainability meets text summarization: A survey, in: Proceedings of the 17th International Natural Language Generation Conference, pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Explainability meets text summarization: A survey, in: Proceedings of the 17th International Natural Language Generation Conference, pp

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.328803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.328803Z digest=sha256:e6879c513f87991c3d5817452bd5704996aa32598df9157af6fea072d1d6db8d

Observation cabed90e-9dca-40e7-9d8c-2ea4f23e8236 · outbound

This paper cites Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Consistency Evaluation of News Article Summaries Generated by Large (and Small) Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.384716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.384716Z digest=sha256:8f460d1f830f75a3ea8f78703bc0782c66a5b340455f1a67f333b936c8ed33d8

Observation 809600eb-c6dd-4c0b-8520-469eb2391571 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.464978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.464978Z digest=sha256:e15aa764f3f5b54a5042d1e980878f69c1760aac84d43b6309f29042206cafc9

Observation 5fb3496f-d1c1-427d-8a82-5f6890f80ebc · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.526027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.526027Z digest=sha256:e7380b711037a50d63746ddd8336f7ecf973a2987cdd0deb4aadbb9e61a2ad86

Observation aba51b5b-1f25-4649-91a3-294785f454d0 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers TrustLLM: Trustworthiness in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.605521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.605521Z digest=sha256:601ef49775086def478a5f1b02a238e016dd5748915771f1479a8e9947725f76

Observation d060040c-c398-43a2-bf95-9d03be857d02 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.655817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.655817Z digest=sha256:12a43edbb4ceac6fc295ea08632314c20c5093362a7c49026d94abfe97d33714

Observation 95f74450-44a6-4caf-a246-cc60b71ae1ab · outbound

This paper cites Consistency analysis of chatgpt, in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Consistency analysis of chatgpt, in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.743863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.743863Z digest=sha256:b33606be2b6f8a6db4b349223aac26f94ff77a24df942916616cecfe8f5231b0

Observation c3386394-0a16-4ba4-85ab-2f6af86a7628 · outbound

This paper cites Context-Aware Sports Highlight Generation Leveraging Large Language Models.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Context-Aware Sports Highlight Generation Leveraging Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.826062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.826062Z digest=sha256:847656864cb6a757358a7884444888ab2223fe3a0565c5259ee711e1153d7651

Observation 7bd3b05b-7ab4-400f-8a84-06ac5f5d09a7 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.908595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.908595Z digest=sha256:e1d93997eeb59eacc850a27b293b03a17c394e516a0397831e6691f50db7f72a

Observation 126b27e8-6410-4740-b003-af20fc2c6b6e · outbound

This paper cites Llms cannot reliably judge (yet?): A comprehensive assessment on the robustness of llm-as-a-judge.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Llms cannot reliably judge (yet?): A comprehensive assessment on the robustness of llm-as-a-judge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:13.964911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:13.964911Z digest=sha256:a8d0795c74e519f92217416f73d35da8c2b8c9cc8bd74c4d37dce749ff1ee618

Observation e9df35f9-1f64-44ac-add0-315d9a909051 · outbound

This paper cites Do not abstain! identify and solve the uncertainty, in: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Do not abstain! identify and solve the uncertainty, in: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), pp

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.045044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.045044Z digest=sha256:7c4b67360428cf97a1c1de4f7a44b1f6fb5a5faaba0cdeab51aa26c1019dffdf

Observation 4b6ac1fc-ceec-4945-83a1-0f0c3c5671c6 · outbound

This paper cites Sumsurvey: An abstractive dataset of scientific survey papers for long document summarization, in: Findings of the ACL 2024, pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Sumsurvey: An abstractive dataset of scientific survey papers for long document summarization, in: Findings of the ACL 2024, pp

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.129595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.129595Z digest=sha256:3b75b24e9486bb761b188512569cbe7fd4979bc238fae72ab6f0f965166812a7

Observation 343ff456-85ca-420d-abde-9d4959a2f7d4 · outbound

This paper cites Low-resource court judgment summarization for common law systems.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Low-resource court judgment summarization for common law systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.214472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.214472Z digest=sha256:33e0038cc36dfb40634ffd60ff12f84c58057857047ec19c72267b62a0526b80

Observation e3b93c37-3db2-4232-a37d-e3c3a644c50d · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.297541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.297541Z digest=sha256:8a32b84e79dc3f87a43e63e05f51b7fa921739ceb4c4eb8ad6e092f022df5d07

Observation 10f62dba-8f8e-4273-a3e8-f9a326ee90f8 · outbound

This paper cites Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.378320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.378320Z digest=sha256:663a5643170aec751bcc7d3cc2dfb896208691478e6077844af22b9242d9803d

Observation f824fb54-7766-4edd-aa89-b918c519da33 · outbound

This paper cites Effectiveness in retrieving legal precedents: exploring text summarization and cutting-edge language models toward a cost-efficient approach.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Effectiveness in retrieving legal precedents: exploring text summarization and cutting-edge language models toward a cost-efficient approach

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.438296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.438296Z digest=sha256:ee0a3d26bf90d114322e90deaf8ed1a6b10aaed0d53338dd617ee98326d97f0b

Observation 06d89a72-c36d-4ab3-972a-98e481ec4533 · outbound

This paper cites State of what art? a call for multi-prompt llm evaluation.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers State of what art? a call for multi-prompt llm evaluation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.496388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.496388Z digest=sha256:d0241b3d8baec0e505708007921b908b7d66c737cfdf37abac42aaf600fc1400

Observation b0193b8d-9fdb-48fd-922e-c76541e7a415 · outbound

This paper cites Abstractive text summarization using sequence-to-sequence rnns and beyond, in: Proceedings of the 20th SIGNLL conference on computational natural language learning, pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Abstractive text summarization using sequence-to-sequence rnns and beyond, in: Proceedings of the 20th SIGNLL conference on computational natural language learning, pp

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.547936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.547936Z digest=sha256:46110b3c250e779bce93cdd7697a1780add64cd3e389340d5620b7a7b7de5ae5

Observation 202d3509-ebfc-412a-9779-072afd236e84 · outbound

This paper cites Evaluating Variance in Visual Question Answering Benchmarks.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Evaluating Variance in Visual Question Answering Benchmarks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.599132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.599132Z digest=sha256:a05b8dee5cc3f5706ecbb035200d52e96b8d0e96edb13c2b95b26e6d3a618710

Observation 974951f5-4266-4556-ab26-7c495e28da2b · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.681971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.681971Z digest=sha256:150bfe97e3b6b9125d4ea5f3fb0a5f454da8105c71ac2b1d5213a22fb0593dab

Observation 2f8529f4-8b82-46e4-a21f-bd02679036bd · outbound

This paper cites Efficient multi-prompt evaluation of llms, in: Proceedings of the 38th International Conference on Neural Information Processing Systems, pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Efficient multi-prompt evaluation of llms, in: Proceedings of the 38th International Conference on Neural Information Processing Systems, pp

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.757339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.757339Z digest=sha256:f14d2b9b2401d8812c437f2331d1785b68991a0376408869253fbd7f281ac1dd

Observation 7cc65411-dfe0-4a9b-ba71-48ab1b18503e · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.842281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.842281Z digest=sha256:6ccb555a94b1b63980c89f76e84c27237ad6fe459c0df14993e894c11dca681e

Observation e60734a2-dc61-4c68-952b-2471bbe58cd7 · outbound

This paper cites Leveraging large language models on the traditional scientific writing workflow, in: 2024 Conference on AI, Science, Engineering, and Technology (AIxSET), IEEE.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Leveraging large language models on the traditional scientific writing workflow, in: 2024 Conference on AI, Science, Engineering, and Technology (AIxSET), IEEE

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.903488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.903488Z digest=sha256:05308fdd329654f3c3574a89183a3d0a2edcbb5aaa6e44e10727f5440ebca78d

Observation 7ae50f9e-85e8-4705-b520-83a86618b120 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.953010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.953010Z digest=sha256:b3b455053039e2bfe88b590ade618169faf24a35279db3e2faf6a6517ee546a7

Observation 16d23be8-631b-4a6b-8395-42dc8c138b97 · outbound

This paper cites How resilient are language models to text perturbations?, in: International Conference on Intelligent Data Engineering and Automated Learning, Springer.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers How resilient are language models to text perturbations?, in: International Conference on Intelligent Data Engineering and Automated Learning, Springer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.030965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.030965Z digest=sha256:8a5c10b8a33e07e6e1210d71a51203630764fcd4fd34c96677b3ef7da0351c6f

Observation 7de55c08-2825-4751-9ff4-6e4fcd5c413a · outbound

This paper cites Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.119758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.119758Z digest=sha256:9bf555316795d7f6563506ba6c3fbda69625b8cf0f959073816271b76b027540

Observation 16003fb3-6a20-4e28-a272-1381f5281c10 · outbound

This paper cites A coin flip for safety: Llm judges fail to reliably measure adversarial robustness.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers A coin flip for safety: Llm judges fail to reliably measure adversarial robustness

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.195403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.195403Z digest=sha256:97b73a8c781a7e6538fd9cb9cb4cbeb4390f3e6ac905b6bef20ec5bccb4fdc03

Observation 2d827d14-b1c1-4bd3-928a-5498c3224a70 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.357112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.357112Z digest=sha256:7c0a94bc5b9ea2137a3391fdfee843587c7a214e63248e8f43708c064438fa72

Observation b067281a-a2b1-40b2-8680-38783516dc41 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.442442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.442442Z digest=sha256:5654c5fe035076966ea2a4efc11f8c45706df0bdd9e29213ed04add1c27def6a

Observation 03ef4978-01b3-4b3d-bb78-18fa5c883a29 · outbound

This paper cites Robustness of large language models to perturbations in text.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Robustness of large language models to perturbations in text

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.611011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.611011Z digest=sha256:6bc4a02dbc9c19d65015f3e384d585c7bc45c5659d72fb7bb5ae5ab04cb82cd5

Observation edb6d993-40e4-4486-88d7-22a441d7efb7 · outbound

This paper cites Legal text summarization via judicial syllogism with large language models.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Legal text summarization via judicial syllogism with large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.768217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.768217Z digest=sha256:b89eddd8c8aa88de249eb0841ce915361a9cfc76f3b9c3acf245b16731eb384c

Observation c3afb610-bf4a-44da-995a-d0367277be74 · outbound

This paper cites Artificial intelligence risk management framework (ai rmf 1.0).(2023).

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Artificial intelligence risk management framework (ai rmf 1.0).(2023)

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:15.873375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:15.873375Z digest=sha256:bf02f5a379995084fd6bebee2d6152187774d971048d3a21f280fa86e4c32a7b

Observation 3fe07b83-615a-413d-a7cb-e1c110969d97 · outbound

This paper cites Evaluating the factual consistency of large language models through news summarization, in: Findings of the Association for Computational Linguistics: ACL 2023, pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Evaluating the factual consistency of large language models through news summarization, in: Findings of the Association for Computational Linguistics: ACL 2023, pp

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.037489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.037489Z digest=sha256:7de23f76bd63ae9481ae098f88785602ae12117ecd3e0060ec517735d93c75b5

Observation aaf3ebbc-c2ec-49e6-a647-898ea236ded5 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.098103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.098103Z digest=sha256:41dbcbe7f63772c3dd43dba73544562d481056243cd5df5fa76d2367c44383e4

Observation 43e5511e-d97d-4d8b-89df-5a7c1429c8dd · outbound

This paper cites Prompt engineering in consistency and reliability with the evidence-based guideline for llms.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Prompt engineering in consistency and reliability with the evidence-based guideline for llms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.188424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.188424Z digest=sha256:5202b4b10010991432a46aae2f35c126edc9ca0fb02976171b2bfe40dfbf8db9

Observation a1ea5236-7a7f-46ce-9374-8b81b6d1da4c · outbound

This paper cites Using llm-supported lecture summarization system to improve knowledge recall and student satisfaction.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Using llm-supported lecture summarization system to improve knowledge recall and student satisfaction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.271309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.271309Z digest=sha256:6525c1e8ba1c1dd50c222be5b8a1df3b71b398135fc7da3c6db343d688812af1

Observation ed5ebf97-8d16-4988-b3a8-d53c42bc02a2 · outbound

This paper cites Comparisons of various types of normality tests.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Comparisons of various types of normality tests

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.352441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.352441Z digest=sha256:139deddcdf129ce286b657dafff69da6c4a1910514ccefd02b53e2b20f71303c

Observation 4e76288c-eb55-4ac2-a35f-b22e8972999a · outbound

This paper cites Event-based evaluation of abstractive news summarization, in: Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2), pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Event-based evaluation of abstractive news summarization, in: Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2), pp

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.413745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.413745Z digest=sha256:cf3b3a7bb49f21ee0a4eecf124377a165068bd6d336f24330c4d317f7e572a42

Observation 9e5fd6cd-414d-4bc3-bcf3-e01497129cb9 · outbound

This paper cites Alignscore: Evaluating factual consistency with a unified alignment function, in: Proceedings of the 61st Annual Meeting of the ACL (V olume 1: Long Papers), pp.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Alignscore: Evaluating factual consistency with a unified alignment function, in: Proceedings of the 61st Annual Meeting of the ACL (V olume 1: Long Papers), pp

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.565036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.565036Z digest=sha256:40cc041448dc33e9fbc498955cf66fb396feb3ead9e44e90a206d3b6f75e61b0

Observation d92b3102-aed1-4c81-9747-fbfc74587ea9 · outbound

This paper cites A systematic survey of text summarization: From statistical methods to large language models.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers A systematic survey of text summarization: From statistical methods to large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.726842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.726842Z digest=sha256:a99123f4d46992fa72fa0f03b17ba7249395ca70e9198f204ac211e0350d2462

Observation 84908382-9703-49db-9ac0-85d915a637cf · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers BERTScore: Evaluating Text Generation with BERT

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:16.840476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:16.840476Z digest=sha256:81b8890c156bd1e6268a0bdf594f0194e4b7bd9ee8c766e846e1d2264a126f00

Observation dc0fea92-c0bd-4775-8685-7f08ff666f63 · outbound

This paper cites Benchmarking large language models for news summarization.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Benchmarking large language models for news summarization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:17.000413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:17.000413Z digest=sha256:4b35c75a72c3942bf3aaeb2772a314a920d3f6f84ec8ac484058cf5ca0bd556c

Observation 8dbd9767-522a-4ddd-89e7-84bfdc68063f · outbound

This paper cites Trustworthy evaluation of large language models.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Trustworthy evaluation of large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:17.115073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:17.115073Z digest=sha256:dcf21a7394e3de5d42397128c14ff6a532108b9b2831d1c918dc43e55f65bb44

Observation 5146d519-3232-4e64-a9ca-5721081d5061 · outbound

This paper cites A comprehensive survey on process-oriented automatic text summarization with exploration of llm-based methods.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers A comprehensive survey on process-oriented automatic text summarization with exploration of llm-based methods

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:17.278674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:17.278674Z digest=sha256:70aaa9c1552afceeda644717f1dc93f5c1fe367c47e9fe0e78487e0e080f307b

Observation 451aec2d-0551-4685-a02b-cbde61db5322 · outbound

This paper cites Assessing the accuracy of artificial intelligence-generated clinical summaries from ambulatory glaucoma subspecialty clinical encounters.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Assessing the accuracy of artificial intelligence-generated clinical summaries from ambulatory glaucoma subspecialty clinical encounters

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:17.438254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:17.438254Z digest=sha256:af78e1e9ddc00f22e192d338b5c264e401d469a9eb3e9d4fb6c01bc8d7d8b1bb

Observation b4f7f642-9430-4782-96af-eb8d3d58dd67 · outbound

This paper cites an unresolved cited work.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:17.594637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:17.594637Z digest=sha256:73f3606683e661ddc87d8421c4347f2697aafebfa3bcf95c7b08deaa4f9c5b52

Pith citing papers

No inbound Pith citation observations are available.