Pith. sign in

Paper Citation Record · LEDGER

What AI Red-Team Evaluations Can and Cannot Prove

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.21735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21735 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:54:48.503377Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d88d1232-5dec-4a04-984b-c4b5a8e0ff33 · outbound

This paper cites Hudson, Ehsan Adeli, et al.

What AI Red-Team Evaluations Can and Cannot Prove Hudson, Ehsan Adeli, et al

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.292789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.292789Z digest=sha256:522dc54df0d7229d4e67b5ac0fa86d057581b99748d8baf575bd5d258007d9f8

Observation 373d399f-b1e6-43d5-8093-c4405b0438c2 · outbound

This paper cites Model cards for model reporting.

What AI Red-Team Evaluations Can and Cannot Prove Model cards for model reporting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.298076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.298076Z digest=sha256:03e5bc8e7478abf24d863d641842ebc6b85d969467b533ffc0d8c9e01a298388

Observation 1dc17684-f994-410b-aa15-0a01343edefb · outbound

This paper cites Holistic evaluation of language models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Holistic evaluation of language models, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.302790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.302790Z digest=sha256:06b12e442a53799d8490a6148bfb6aed561f6d2107d43760eb3b25dab51efba6

Observation 77c09251-4538-47c7-85cf-fdae96f89c24 · outbound

This paper cites Gritsenko, et al.

What AI Red-Team Evaluations Can and Cannot Prove Gritsenko, et al

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.307997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.307997Z digest=sha256:71e25870de1e0a477e149fc5c5865deccd7ce1fc1064a8157b5770992ace72c3

Observation dbb6170d-e3d4-476f-b6a1-7d3af53b54bb · outbound

This paper cites Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach.

What AI Red-Team Evaluations Can and Cannot Prove Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.312735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.312735Z digest=sha256:2d1fcfab9dfe5bba40fc8b4689099af9a6ab3aac42618b07b36825e6e4c2496d

Observation bfbf40ef-b0e4-49f6-8f4b-3b6555b6659d · outbound

This paper cites The structural safety gener- alization problem, 2025.

What AI Red-Team Evaluations Can and Cannot Prove The structural safety gener- alization problem, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.317460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.317460Z digest=sha256:5c04da7b54ec1684d362d550ff8ebb1e5e6c8f61b44f3d8402634a0363db7930

Observation a2c47ca0-2d57-47cc-a63b-e132dcfd7056 · outbound

This paper cites Adding error bars to evals: A statistical approach to language model evaluations, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Adding error bars to evals: A statistical approach to language model evaluations, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.322834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.322834Z digest=sha256:c9fd54ce2a4bf99967da838cedbc9de7564548b7c18f594ddf22f919d652edcb

Observation 3d11f44d-a56b-4035-84e3-33a05813ffa2 · outbound

This paper cites Kim and Anthony R.

What AI Red-Team Evaluations Can and Cannot Prove Kim and Anthony R

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.327008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.327008Z digest=sha256:3fe561107c12433d4143cc328803f2bebbd3787e45df122e25f67dc18f6e841f

Observation 57a5a564-44a7-4132-ae28-c6dccd5a2e43 · outbound

This paper cites Prentice.

What AI Red-Team Evaluations Can and Cannot Prove Prentice

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.331485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.331485Z digest=sha256:5a7d13a531f1618cf4beca29bf32da661816bc93be1b5e8733be69829eb2a07b

Observation 15d7eebd-aa13-47ce-b215-3e3b5db064f5 · outbound

This paper cites Safety cases: How to justify the safety of advanced AI systems, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Safety cases: How to justify the safety of advanced AI systems, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.336192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.336192Z digest=sha256:703dcf28ee8458836dcdeed86f840489aa1d87f2cd58fc6db91e09e78487dbc8

Observation d1c28ae5-00a6-4d28-bbc5-9371f24a0edf · outbound

This paper cites A sketch of an AI control safety case, 2025.

What AI Red-Team Evaluations Can and Cannot Prove A sketch of an AI control safety case, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.340824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.340824Z digest=sha256:0dfe55508982f651d4d6c4a57e198f0161356dc5b0105f4fdddb997de53bac5d

Observation bd78ca10-2ea9-4e4b-8876-96ee4677661b · outbound

This paper cites Shadish, Thomas D.

What AI Red-Team Evaluations Can and Cannot Prove Shadish, Thomas D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.346205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.346205Z digest=sha256:d972275dae1de9a8e719289c3c4a3ea085afcfaa941e18fdca35a2a5c29920a0

Observation 8fd1b70a-b96a-4904-8310-6ff63f2a110d · outbound

This paper cites Dulberg, and George A.

What AI Red-Team Evaluations Can and Cannot Prove Dulberg, and George A

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.351912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.351912Z digest=sha256:7383826d028592f7efd9788897bcf57df8c7fe5fb2f451ba796f6f5925cf0cc8

Observation b985ad45-8b24-404f-bd54-efc7efda56bd · outbound

This paper cites Hanley and Abby Lippman-Hand.

What AI Red-Team Evaluations Can and Cannot Prove Hanley and Abby Lippman-Hand

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.356405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.356405Z digest=sha256:33c44d6848e93e4244d7375c7db56a0f355f2dd7f73f229d5ce34c9612109aa1

Observation 80f097a5-e73e-4c79-8be4-89db9febbe71 · outbound

This paper cites Brown, and Francis R.

What AI Red-Team Evaluations Can and Cannot Prove Brown, and Francis R

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.360716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.360716Z digest=sha256:2c6eb04cd36c074bbe301211cc917c5be87f6756f5087733301f9d291f5c7011

Observation 3eef9d01-3c9f-4227-ba1f-356c9a457061 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove XSTest: A test suite for identifying exaggerated safety behaviours in large language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.365061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.365061Z digest=sha256:3555a9911e224fa07c251914f82335b244a8e0e03be2f24ab2e2b729b00f677a

Observation 35546b27-3658-41db-ad14-ecb340c7aaee · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

What AI Red-Team Evaluations Can and Cannot Prove SafetyBench: Evaluating the Safety of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.369854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.369854Z digest=sha256:b4aaf16622d2d572e820e82b3d3dd78453b006a1a2b462ae9761be9e32f249e0

Observation 9876aa84-411b-4b94-828d-1e58b0411666 · outbound

This paper cites Zico Kolter, and Matt Fredrikson.

What AI Red-Team Evaluations Can and Cannot Prove Zico Kolter, and Matt Fredrikson

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.374594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.374594Z digest=sha256:177d5b3f9441480e392d3f6f0c0c1aa9f493b222b643f8369b604734cdf3438a

Observation 672f8c41-867c-4c87-b89a-54f65793a14c · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024.

What AI Red-Team Evaluations Can and Cannot Prove HarmBench: A standardized evaluation framework for automated red teaming and robust refusal, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.378992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.378992Z digest=sha256:33401bc32aed439ea29c0f2ed674ee888e18f769337f52b69033b4ad72364f3c

Observation 5b2fab8a-2f47-4b04-9ebd-f839dc927756 · outbound

This paper cites A StrongREJECT for empty jailbreaks, 2024.

What AI Red-Team Evaluations Can and Cannot Prove A StrongREJECT for empty jailbreaks, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.383131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.383131Z digest=sha256:4a803e27f27136323235a37b2d7bba18a0d2ec716a0d0594a79307d7ec6e0e46

Observation 2d6737ef-b31b-4d9d-8da0-0885059f99a3 · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.387331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.387331Z digest=sha256:d78fc303630c017140170930abe7009eacfb66f13e18612338cca2c07a5e7aa1

Observation f37e2ad5-f7d1-4f60-bcc3-33ae49081e2d · outbound

This paper cites John Wiley and Sons, New York, 1965.

What AI Red-Team Evaluations Can and Cannot Prove John Wiley and Sons, New York, 1965

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.391557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.391557Z digest=sha256:52d9ec905a6d55f501f7f8303eae155fba65afce9d74f7fd18a82a8f5d96e843

Observation 6467b702-3a36-4728-acd7-6519488f324e · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.395950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.395950Z digest=sha256:dec51caef1eee9155d223ad2eda3789fb79926df901cf2513ceeac1efcff3c73

Observation ec26b57d-c8a9-40b2-a2eb-5fbf0aad3434 · outbound

This paper cites Model card and evaluations for claude models.

What AI Red-Team Evaluations Can and Cannot Prove Model card and evaluations for claude models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.400904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.400904Z digest=sha256:d7541f362784781d17114db429cd58f2153f1ad0c2241fbb57473007d6fbc0d2

Observation 2743c202-13bb-48ad-9ec3-9d6c206cc841 · outbound

This paper cites Sentence-BERT: Sentence embeddings using siamese BERT-networks.

What AI Red-Team Evaluations Can and Cannot Prove Sentence-BERT: Sentence embeddings using siamese BERT-networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.406066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.406066Z digest=sha256:cc4facc174700ab02b023786cf316fac929bd5a6ce88c4a8705a9adf36b14939

Observation 9382bd33-680a-4971-a2c4-76cec48fe29c · outbound

This paper cites LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024.

What AI Red-Team Evaluations Can and Cannot Prove LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.410572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.410572Z digest=sha256:19b2b9249086ebcefab7a585120625b6d06848f72d7ea0af28d1622319457b1c

Observation 211ea424-422f-4e2d-b5a6-6eaec1d95e12 · outbound

This paper cites Borgwardt, Malte J.

What AI Red-Team Evaluations Can and Cannot Prove Borgwardt, Malte J

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.415116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.415116Z digest=sha256:e6e7cd843269275d73fafb2d8fd38add468375d3d6a4df834908dbbe3d418d1e

Observation d26b55e6-5b01-46cc-9a5f-dc8f24d4fc7a · outbound

This paper cites UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018.

What AI Red-Team Evaluations Can and Cannot Prove UMAP: Uniform manifold ap- proximation and projection for dimension reduction, 2018

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.420308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.420308Z digest=sha256:fb0b3ed061c87e6ec82fc04c50343fef7a9224ff1e2a3318844f2ea7172a19b1

Observation 942aafff-26e4-42eb-bb8c-7f19f5c9ef46 · outbound

This paper cites Choquette-Choo, et al.

What AI Red-Team Evaluations Can and Cannot Prove Choquette-Choo, et al

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.426517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.426517Z digest=sha256:9e303270d7233d66e88cf977765e3822c7300fa90e9adbc49046e32112c47878

Observation 1a80966c-fb58-438f-9e1b-4b4c63f0615c · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove AutoDAN: Generating stealthy jailbreak prompts on aligned large language models, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.434251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.434251Z digest=sha256:118731b361320c0bcf6673f63d098d02df12b3f46f311d2e9bf5862d39d23025

Observation 7130a205-44c7-4b3b-a148-c24277ff8fe8 · outbound

This paper cites Does refusal training in LLMs generalize to the past tense?, 2025.

What AI Red-Team Evaluations Can and Cannot Prove Does refusal training in LLMs generalize to the past tense?, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.440217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.440217Z digest=sha256:94fa4dd67b2fe780c18ed5c28c84394063114e5faa48fcccf379cf5f99bc08d5

Observation 40a1a1e4-0715-4b4e-9a2d-87dba58c35be · outbound

This paper cites an unresolved cited work.

What AI Red-Team Evaluations Can and Cannot Prove Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.446386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.446386Z digest=sha256:f2abe3994b1a05ac8f36ac3d7b7c349315efce500651e42006ea80cc8de10809

Observation 0eb41bd9-2072-4694-a6f9-e520ee816c51 · outbound

This paper cites Tree of attacks: Jail- breaking black-box LLMs automatically, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Tree of attacks: Jail- breaking black-box LLMs automatically, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.451336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.451336Z digest=sha256:3cdbb1fc2ace03a429db260449b89401b85ff054f6f51f6fb71dbbaf431ce314

Observation 652b69c2-d087-4678-bab5-18eb0e197c1d · outbound

This paper cites Ignore previous prompt: Attack techniques for lan- guage models, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Ignore previous prompt: Attack techniques for lan- guage models, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.455476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.455476Z digest=sha256:a1b1b1d46868fa887b4e1604d6eab5d3e9e3d5c370cba2788135a73953e3fbc6

Observation 66ea6870-7599-4e90-8976-8d8e4930c06f · outbound

This paper cites GPT-4 technical report, 2023.

What AI Red-Team Evaluations Can and Cannot Prove GPT-4 technical report, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.460464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.460464Z digest=sha256:15179ffd40341b0588627c21daae83637f6f81ead20a2879456126ed525703dd

Observation d73ba15e-e881-40ee-a011-21c5730c57e0 · outbound

This paper cites GPT-4o system card, 2024.

What AI Red-Team Evaluations Can and Cannot Prove GPT-4o system card, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.464794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.464794Z digest=sha256:3c4830fccb4518b7d30d29bbd11016f092b5cfead451d88c6778121e544cc7ee

Observation 6aeced3a-b254-4c43-9259-387270fb29bd · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

What AI Red-Team Evaluations Can and Cannot Prove The claude 3 model family: Opus, sonnet, haiku

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.469031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.469031Z digest=sha256:fbe0c07c4d5eb946661afd16063034388602041f63a38aec49a4a31ff313c6c0

Observation b8d0024a-c2f1-4657-910c-c039b8fdf02f · outbound

This paper cites System card: Claude opus 4 and claude sonnet 4.

What AI Red-Team Evaluations Can and Cannot Prove System card: Claude opus 4 and claude sonnet 4

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.473574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.473574Z digest=sha256:e373c2faa9a6732d7066c20bbad1ed9ec4061d5eda1a8c9c85eb55b41796b3cc

Observation 9a976dac-022b-4d95-a2b6-79f8b8d1df37 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Gemini: A family of highly capable multimodal models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.478226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.478226Z digest=sha256:856b7a1b8d01247c5ff79f5b5194952bbc4d6beb2e6434c96d5c2fb01d73b9ae

Observation 43d610c4-3c39-4a56-b54c-3403433df20b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024.

What AI Red-Team Evaluations Can and Cannot Prove Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.482631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.482631Z digest=sha256:216faad3de5eae1b129fdaf8f7896cef3f16f423aa0c18bbc1e5f1af12b6f7c8

Observation 4954885d-4486-4200-a812-e363906d50eb · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

What AI Red-Team Evaluations Can and Cannot Prove Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.486674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.486674Z digest=sha256:0596b9faca6f1031be8194615b3436ce4649c1233486179b398545e3ba1f8eb8

Observation edfa1227-57cf-4042-8c31-f4869976d8f4 · outbound

This paper cites The llama 3 herd of models, 2024.

What AI Red-Team Evaluations Can and Cannot Prove The llama 3 herd of models, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.490984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.490984Z digest=sha256:4775c3d10ef1c5e39672cdf39ebc65d6ed2833175fdafb14d49d302eabd24f74

Observation 56e9109f-1aee-4983-9afd-fa36d97d3de9 · outbound

This paper cites Responsible scaling policy.

What AI Red-Team Evaluations Can and Cannot Prove Responsible scaling policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.495168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.495168Z digest=sha256:7907c823439dd55d0a0d8b0ce1e3c9c019a4d6efed2e92371446440d79fede81

Observation 2bacd9a9-bb05-4f97-a16d-74edf5a71c39 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.499159Z digest=sha256:895716e33a3dcba896eea0e79e88be2e5caf6ecdff9c5daf33f9c77226b1e25f

Observation 8811b220-1865-4f65-acbe-2ba4172ec09b · outbound

This paper cites Red teaming language models with language models, 2022.

What AI Red-Team Evaluations Can and Cannot Prove Red teaming language models with language models, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T06:54:48.503377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:54:48.503377Z digest=sha256:47e79e0ba831e0960460905e2e780eda5f22cf54de8e3300b6940d71aa0d2656

Pith citing papers

No inbound Pith citation observations are available.