Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:54:48.503377Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.21735.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:54:48.503377Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d88d1232-5dec-4a04-984b-c4b5a8e0ff33 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Hudson, Ehsan Adeli, et al
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373d399f-b1e6-43d5-8093-c4405b0438c2 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Model cards for model reporting
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc17684-f994-410b-aa15-0a01343edefb · outbound
What AI Red-Team Evaluations Can and Cannot Prove Holistic evaluation of language models, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c09251-4538-47c7-85cf-fdae96f89c24 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Gritsenko, et al
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb6170d-e3d4-476f-b6a1-7d3af53b54bb · outbound
What AI Red-Team Evaluations Can and Cannot Prove Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfbf40ef-b0e4-49f6-8f4b-3b6555b6659d · outbound
What AI Red-Team Evaluations Can and Cannot Prove The structural safety gener- alization problem, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c47ca0-2d57-47cc-a63b-e132dcfd7056 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Adding error bars to evals: A statistical approach to language model evaluations, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d11f44d-a56b-4035-84e3-33a05813ffa2 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Kim and Anthony R
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57a5a564-44a7-4132-ae28-c6dccd5a2e43 · outbound
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d7eebd-aa13-47ce-b215-3e3b5db064f5 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Safety cases: How to justify the safety of advanced AI systems, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c28ae5-00a6-4d28-bbc5-9371f24a0edf · outbound
What AI Red-Team Evaluations Can and Cannot Prove A sketch of an AI control safety case, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd78ca10-2ea9-4e4b-8876-96ee4677661b · outbound
What AI Red-Team Evaluations Can and Cannot Prove Shadish, Thomas D
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd1b70a-b96a-4904-8310-6ff63f2a110d · outbound
What AI Red-Team Evaluations Can and Cannot Prove Dulberg, and George A
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b985ad45-8b24-404f-bd54-efc7efda56bd · outbound
What AI Red-Team Evaluations Can and Cannot Prove Hanley and Abby Lippman-Hand
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f097a5-e73e-4c79-8be4-89db9febbe71 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Brown, and Francis R
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eef9d01-3c9f-4227-ba1f-356c9a457061 · outbound
What AI Red-Team Evaluations Can and Cannot Prove XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35546b27-3658-41db-ad14-ecb340c7aaee · outbound
What AI Red-Team Evaluations Can and Cannot Prove SafetyBench: Evaluating the Safety of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9876aa84-411b-4b94-828d-1e58b0411666 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Zico Kolter, and Matt Fredrikson
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 672f8c41-867c-4c87-b89a-54f65793a14c · outbound
What AI Red-Team Evaluations Can and Cannot Prove HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2fab8a-2f47-4b04-9ebd-f839dc927756 · outbound
What AI Red-Team Evaluations Can and Cannot Prove A StrongREJECT for empty jailbreaks, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6737ef-b31b-4d9d-8da0-0885059f99a3 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f37e2ad5-f7d1-4f60-bcc3-33ae49081e2d · outbound
What AI Red-Team Evaluations Can and Cannot Prove John Wiley and Sons, New York, 1965
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6467b702-3a36-4728-acd7-6519488f324e · outbound
What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec26b57d-c8a9-40b2-a2eb-5fbf0aad3434 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Model card and evaluations for claude models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2743c202-13bb-48ad-9ec3-9d6c206cc841 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Sentence-BERT: Sentence embeddings using siamese BERT-networks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9382bd33-680a-4971-a2c4-76cec48fe29c · outbound
What AI Red-Team Evaluations Can and Cannot Prove LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211ea424-422f-4e2d-b5a6-6eaec1d95e12 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Borgwardt, Malte J
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26b55e6-5b01-46cc-9a5f-dc8f24d4fc7a · outbound
What AI Red-Team Evaluations Can and Cannot Prove UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942aafff-26e4-42eb-bb8c-7f19f5c9ef46 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Choquette-Choo, et al
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a80966c-fb58-438f-9e1b-4b4c63f0615c · outbound
What AI Red-Team Evaluations Can and Cannot Prove AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7130a205-44c7-4b3b-a148-c24277ff8fe8 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Does refusal training in LLMs generalize to the past tense?, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a1a1e4-0715-4b4e-9a2d-87dba58c35be · outbound
What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb41bd9-2072-4694-a6f9-e520ee816c51 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Tree of attacks: Jail- breaking black-box LLMs automatically, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 652b69c2-d087-4678-bab5-18eb0e197c1d · outbound
What AI Red-Team Evaluations Can and Cannot Prove Ignore previous prompt: Attack techniques for lan- guage models, 2022
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ea6870-7599-4e90-8976-8d8e4930c06f · outbound
What AI Red-Team Evaluations Can and Cannot Prove GPT-4 technical report, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d73ba15e-e881-40ee-a011-21c5730c57e0 · outbound
What AI Red-Team Evaluations Can and Cannot Prove GPT-4o system card, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aeced3a-b254-4c43-9259-387270fb29bd · outbound
What AI Red-Team Evaluations Can and Cannot Prove The claude 3 model family: Opus, sonnet, haiku
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d0024a-c2f1-4657-910c-c039b8fdf02f · outbound
What AI Red-Team Evaluations Can and Cannot Prove System card: Claude opus 4 and claude sonnet 4
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a976dac-022b-4d95-a2b6-79f8b8d1df37 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Gemini: A family of highly capable multimodal models, 2023
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d610c4-3c39-4a56-b54c-3403433df20b · outbound
What AI Red-Team Evaluations Can and Cannot Prove Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4954885d-4486-4200-a812-e363906d50eb · outbound
What AI Red-Team Evaluations Can and Cannot Prove Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edfa1227-57cf-4042-8c31-f4869976d8f4 · outbound
What AI Red-Team Evaluations Can and Cannot Prove The llama 3 herd of models, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56e9109f-1aee-4983-9afd-fa36d97d3de9 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Responsible scaling policy
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bacd9a9-bb05-4f97-a16d-74edf5a71c39 · outbound
What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8811b220-1865-4f65-acbe-2ba4172ec09b · outbound
What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models with language models, 2022
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.