Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:09:54.183251Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2412.09269.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:09:54.183251Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-15T12:59:04.484292Z
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f8060070-e346-4d86-8112-2b0807bdc0d8 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf234a74-bd4b-4c8f-be6c-e91854d47810 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 394daea4-e4f2-4157-891f-5069d9181aa9 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Re-evaluating Evaluation in Text Summarization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e13a94e-cc17-42bc-a087-23fa2c759481 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66dcf312-e0a5-4df1-a7d5-c43f3c795ee5 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803a6656-b02b-41ea-a204-03070e5f380c · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Gemma: Open Models Based on Gemini Research and Technology
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61379a8e-e48e-4be6-8009-d2fa18d4b866 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations LLaMA: Open and Efficient Foundation Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79839ef-05db-4754-a65b-28c436494747 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac2e9577-33ff-482f-9938-941b3e43c6f8 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Language Models are Few-Shot Learners
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada3f868-7c45-4478-a6e5-ce83389e5456 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations SummEval: Re-evaluating Summarization Evaluation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c6a6f2-e6ed-495d-8918-9b2a8085e504 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72b22a4-244e-4db6-a8f6-a1edb4a22e3a · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Scaling Trends in Language Model Robustness
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee55629d-f332-4c7d-b797-ee287915b088 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650dcf8f-40fe-4dd4-b6f9-b5ff79cee0e9 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fcbb0c0-c9aa-4b5f-914f-8fbe4284272d · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d450cdc2-c782-482e-ab39-51d89260b056 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Generated Knowledge Prompting for Commonsense Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5fe6ae-e82b-410a-93e6-8f52e449a1d1 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b4f08a4-a186-49ee-94b4-56353fec2413 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68825cb3-a41d-4a33-b674-2b97329fd8f6 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39f4052-00bc-4d4b-9d89-f8c258d4d13e · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71fa5f8-df0d-49f5-90db-27efe6fb39d4 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b7379c9-b4c2-4d5a-8704-896c4058af97 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Gemini: A Family of Highly Capable Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f4836b-a0a9-4aa6-900b-bbeeba2e1ecb · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Finetuned Language Models Are Zero-Shot Learners
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f935100-b3c5-4f29-b2a7-af7b4c810a40 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc316826-c338-488b-88c9-a5f96bb579d8 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations A Comprehensive Assessment of Dialog Evaluation Metrics
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b32259-2cc1-4c63-a9a4-abd910efec36 · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations BERTScore: Evaluating Text Generation with BERT
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879d24d4-9cf0-402b-a4ca-aa003605fb2c · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations online" 'onlinestring :=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc42d19b-bd13-4fab-a9a1-106d04b42aaf · outbound
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations write newline
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e08fb8-2846-423b-9d71-894d5f1b9948 · inbound
Learning When to Trust in Contextual Social Bandits Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.