Pith. sign in

Paper Citation Record · LEDGER

AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2407.17436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.17436 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:08.125748Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32bb5059-d424-47b3-a16e-80c5ac345f4a · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 242

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.125748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.125748Z digest=sha256:54d768dabb82da99a901f9f7df27a30c3c07f21409300c98e32b49049ef13148

Observation 25b02937-0332-4627-a663-341e2264d9ac · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 230

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.221798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.221798Z digest=sha256:3e6838f81ed72024af38aff90385a7488debcc16bad4d3bffb221bd1ddf62d7f

Observation fd6579a4-a8d8-4ddf-9357-46a11e763ff8 · inbound

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset cites this paper.

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:35:31.636611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:32:52.986523Z digest=sha256:1d068d7b4fd661bd66ef875d7f92a48abf9035215c25db8fb75f97100f5cd248

Observation 93a34e3b-59e6-4fa6-92b9-1d52fe91019f · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:45:14.326403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:815da3353bde857cf33d8db937aa12fcdb6f757dca11e89566abb06006ed5f58

Observation f3818a60-a2e4-466c-80fa-cf4cc7520b4a · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:40:43.456003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:1a11ac3e1b23470fd82d6433a8134952783e9e55e2478ecadb2220a54f6179b5

Observation ad4d6a5d-b8e0-490c-9650-c701483c320c · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.738353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:5efe95c00c0954259d6acab2c82a7fa1c34f473e3e552f1752f66b1027580fcf

Observation 7ebd75bf-b23f-4b84-97de-3bd1b6993d3f · inbound

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety cites this paper.

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:45.799831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T04:55:36.069767Z digest=sha256:f45e861d1e0a16b513d3c7560b4087d8a81858c88af30c2dc334164b4426b7b0

Observation a768578b-b939-42c5-a201-531f761575b0 · inbound

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety cites this paper.

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T16:49:32.395383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:49:32.395383Z digest=sha256:05c773fd6ab6d8b8a9d8428937b770263795bf289eb80067894951524a02740f

Observation 361f769d-7269-4c3e-8ecf-29bfd8626531 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.882817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:67e1596f887a56e3bba8c72ede71b0daef7c9cc3c5f3cc9d130f89ef814993a3

Observation 5671be64-514b-49dc-a6b5-e5b8aac7edb1 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:42.863331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:d90125d62dcd5b2c4e43626a14d7011eb84690903ac715c8d5dbc443adaac5ae

Observation 4b7a78d5-fc87-4552-9dfc-399aa26317b1 · inbound

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation cites this paper.

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:25.159443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:05:57.685728Z digest=sha256:3fb3893f371186048b55e5db67ac66c19e091e690a3f923fe0e7609d39345929

Observation d8ecb8ee-7117-4d70-ba5a-7ac8bfe58c4c · inbound

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability cites this paper.

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.809439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T10:00:30.904247Z digest=sha256:f45c232cd127ec5193a732371e9bcb768b2c7f26b2d817d0363b9ff29820fab1

Observation a49d69b0-35de-48a1-93c7-5d5d442eaa76 · inbound

Beyond Single-Policy: Evaluating Composed Organization-Specific Policy Alignment in LLM Chatbots cites this paper.

Beyond Single-Policy: Evaluating Composed Organization-Specific Policy Alignment in LLM Chatbots AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:41:40.158587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T05:41:16.033862Z digest=sha256:ccfcd19e9661ed1532c2b7bd30f4eab65dcdbd924a1459673a00990cb15cd374

Observation 821c38c6-b39a-4790-b855-e6e2e115e6e1 · inbound

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations cites this paper.

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:47:26.405643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T18:37:24.395892Z digest=sha256:520e0c9d6866771379158ab47885abd7075500ddb18abda3f5b63ec9244d3442

Observation cd4ad914-2d01-4bd9-8315-73e47c8324af · inbound

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis cites this paper.

Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:07:30.328565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:48:54.802860Z digest=sha256:d7378ae0bb45c48b34f293de5b024441d09d517558945463b11d60259fc92ff1

Observation b466eda9-ecf2-4ae4-9d68-1757f587d565 · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:34.777211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:ea3388ec51cff6cb841127e27fea64f50a23cbe3f7e018f6138a8d13e5a50877

Observation 3eccc86e-7846-4ba5-a478-6d7a2ec44d1e · inbound

Efficient Safety Benchmarking via Item Response Theory cites this paper.

Efficient Safety Benchmarking via Item Response Theory AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:55:48.974164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T15:51:00.484343Z digest=sha256:636fa71a07fe3be4321031f41c8eed499459b737c0d63dd0e958c04cb7ac70d9

Observation 7c3bdb28-b686-421f-ac4d-5bb66a412136 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:52.928929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:52.928929Z digest=sha256:aea327a4ed7c72e4f015266f4783b56b9bdf9c1f8093a53ddcc6bae2d6584e62

Observation 76cd7933-fe44-421e-b840-c1cef0f98e17 · inbound

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models cites this paper.

AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:30:01.720511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:30:01.720511Z digest=sha256:d7eff812cb47673844340a6c0929768ab4736722c98317d4489f48d0bd22a603

Observation 2a2bfb26-0fbe-4c53-955a-2c7da7224d28 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:22.163420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:22.163420Z digest=sha256:acca232332dc10ffb7a0d5a9be504ac7329dfd3396006a00cabacaafbda5f1a0