Pith. sign in

Paper Citation Record · LEDGER

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations

As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2412.09269.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09269 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:09:54.183251Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T12:59:04.484292Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8060070-e346-4d86-8112-2b0807bdc0d8 · outbound

This paper cites GPT-4 Technical Report.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.009467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.009467Z digest=sha256:221f6e5b56bee850e4ed4d70d590c8b081b2287cae36c653c4d7fb2d286c9974

Observation bf234a74-bd4b-4c8f-be6c-e91854d47810 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.017438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.017438Z digest=sha256:7f90a1dd78d20c5f5630e398cf1cf90d50bb4f52a5f9e4698ccd7e77c3eec012

Observation 394daea4-e4f2-4157-891f-5069d9181aa9 · outbound

This paper cites Re-evaluating Evaluation in Text Summarization.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Re-evaluating Evaluation in Text Summarization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.025882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.025882Z digest=sha256:0f056ace0761154a26cf7d50159efa1d6688fc4921521da3ade4243ea1ff80c1

Observation 5e13a94e-cc17-42bc-a087-23fa2c759481 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.035630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.035630Z digest=sha256:5d507d81a66e5b9dbab7f1bf74d13567d5d84f948076af59d0de38c11e9b49e7

Observation 66dcf312-e0a5-4df1-a7d5-c43f3c795ee5 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.042824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.042824Z digest=sha256:a1e637537e06b7884ab2d172f43c594f27ec2d1f182efb372acbbb539097778e

Observation 803a6656-b02b-41ea-a204-03070e5f380c · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Gemma: Open Models Based on Gemini Research and Technology

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.048030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.048030Z digest=sha256:6ed141407df92091bcab966c3d9798596ad8bdf841a2eaae34c430e46222f35d

Observation 61379a8e-e48e-4be6-8009-d2fa18d4b866 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations LLaMA: Open and Efficient Foundation Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.061600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.061600Z digest=sha256:1831d1c8ad5f3b11c8a64fbe099c9a13300116707a1007fa58959727d3747f5a

Observation a79839ef-05db-4754-a65b-28c436494747 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.067255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.067255Z digest=sha256:9884ab58e6bdb489b19a11cd41621acae83df0044f2fc06120f8327005241f78

Observation ac2e9577-33ff-482f-9938-941b3e43c6f8 · outbound

This paper cites Language Models are Few-Shot Learners.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Language Models are Few-Shot Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.072999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.072999Z digest=sha256:4206f57752d96aee755a1132930825868981efaea31376127a60931afa03a184

Observation ada3f868-7c45-4478-a6e5-ce83389e5456 · outbound

This paper cites SummEval: Re-evaluating Summarization Evaluation.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations SummEval: Re-evaluating Summarization Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.079938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.079938Z digest=sha256:40fc8abfc04f019df3a5a5bd184693a67374838278635e836059c500f5362271

Observation 29c6a6f2-e6ed-495d-8918-9b2a8085e504 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.086772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.086772Z digest=sha256:020c07189eba30b07166aada2605e5e310f3cb85751f576ba02efdc7406f7f43

Observation d72b22a4-244e-4db6-a8f6-a1edb4a22e3a · outbound

This paper cites Scaling Trends in Language Model Robustness.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Scaling Trends in Language Model Robustness

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.093857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.093857Z digest=sha256:0dee8748c11f4dd103b7e5961841572c07c250b2c06be9224601c64abca5712f

Observation ee55629d-f332-4c7d-b797-ee287915b088 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.100225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.100225Z digest=sha256:fbf1a96313752655220ca6a95cbac48065b25213b7fa4f0760e12abc2fb41a45

Observation 650dcf8f-40fe-4dd4-b6f9-b5ff79cee0e9 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.105783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.105783Z digest=sha256:81295cef529720476feb92c326c52dd3633229116f63256c42e8eb066b3369a5

Observation 1fcbb0c0-c9aa-4b5f-914f-8fbe4284272d · outbound

This paper cites Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.111112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.111112Z digest=sha256:10c309766ed2e33622bf01608637dda574dbacef838286233b3d9ca1b36988c1

Observation d450cdc2-c782-482e-ab39-51d89260b056 · outbound

This paper cites Generated Knowledge Prompting for Commonsense Reasoning.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Generated Knowledge Prompting for Commonsense Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.116282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.116282Z digest=sha256:6108e87dc1e4a85c891606eefd99fee7132aa535cb0bc178db310cb01a2944f5

Observation 4b5fe6ae-e82b-410a-93e6-8f52e449a1d1 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:09:54.833397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:09:54.121522Z digest=sha256:ef02a703922324d52b69e301c498e0ae782a1b4d2cbd109b2995446ca2f4cc18

Observation 6b4f08a4-a186-49ee-94b4-56353fec2413 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.126636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.126636Z digest=sha256:2722f7163cf0d1af8a8f99382f029cf87a0282be8c951163f224e0c39896aa6a

Observation 68825cb3-a41d-4a33-b674-2b97329fd8f6 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.132080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.132080Z digest=sha256:860cd2fafa3cc7c397b5e52666e4f2f15f24e3a41d2ab5aa8f49ac2de4d16842

Observation e39f4052-00bc-4d4b-9d89-f8c258d4d13e · outbound

This paper cites Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.139224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.139224Z digest=sha256:5223fa6f22f1a3a026e7e30afc5ca9617df141f3209316176af2523652f52960

Observation a71fa5f8-df0d-49f5-90db-27efe6fb39d4 · outbound

This paper cites an unresolved cited work.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:09:54.806572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T17:09:54.145378Z digest=sha256:7238e47137c0397a9a1a6a8701a99ca1c0862450cfc4ae966948a5669107140c

Observation 2b7379c9-b4c2-4d5a-8704-896c4058af97 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Gemini: A Family of Highly Capable Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.150020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.150020Z digest=sha256:aa3451de22fb2175a131a22f9ef2f08cd40839b7eabc40038380d7e5b133ab09

Observation 85f4836b-a0a9-4aa6-900b-bbeeba2e1ecb · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Finetuned Language Models Are Zero-Shot Learners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.154915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.154915Z digest=sha256:6f2a90d9c45c6e8acb195d6b169cc812674fb52944680bafc85a08694d482430

Observation 4f935100-b3c5-4f29-b2a7-af7b4c810a40 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.159275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.159275Z digest=sha256:2eaa02202073a6cd6f7229266a81c5bb465a182bf8b29329e732bd5d6ffb70da

Observation fc316826-c338-488b-88c9-a5f96bb579d8 · outbound

This paper cites A Comprehensive Assessment of Dialog Evaluation Metrics.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations A Comprehensive Assessment of Dialog Evaluation Metrics

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.164611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.164611Z digest=sha256:2f3ef54d66974aa34643a843e4123a06c3b1a2c8b4fa8f6c26cb7f277d1349a0

Observation 15b32259-2cc1-4c63-a9a4-abd910efec36 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.169733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.169733Z digest=sha256:5723ace3bf01d531191dd7d61df0a586ab9b68145c26a34530ce285f17643392

Observation 879d24d4-9cf0-402b-a4ca-aa003605fb2c · outbound

This paper cites online" 'onlinestring :=.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations online" 'onlinestring :=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.175635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.175635Z digest=sha256:013aec144d2fffc9695337bfa0d031e7b961714b5c1a60bb8b7a312b6bac7662

Observation dc42d19b-bd13-4fab-a9a1-106d04b42aaf · outbound

This paper cites write newline.

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:09:54.183251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:09:54.183251Z digest=sha256:42ba522a8479eb2feacdcb5c0288bef7bf63dde34ff82638493dc3130b281cd2

Pith citing papers

Observation 42e08fb8-2846-423b-9d71-894d5f1b9948 · inbound

Learning When to Trust in Contextual Social Bandits cites this paper.

Learning When to Trust in Contextual Social Bandits Towards Understanding the Robustness of LLM-based Evaluations under Perturbations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T12:59:04.484292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:59:04.484292Z digest=sha256:93873701c15867693eb12a67a60f84d543cfd089ed73d1d2cc3c454eb9c2f432