Pith. sign in

Paper Citation Record · LEDGER

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2503.05731.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.05731 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:29:09.565742Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a6ef3c8f-3054-4d30-97e9-7e93f3ea1736 · inbound

How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape cites this paper.

How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:10:26.071510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T23:08:21.946111Z digest=sha256:8d31b7ff51b0a3fec3f7b1792481e02a0d65c5e0436d5f245fe4ac39a3ba7d3a

Observation cec9f3ff-7c78-4e8f-a92e-0da682b92f07 · inbound

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety cites this paper.

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:01:06.452715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:17:43.880661Z digest=sha256:ca19b73e20b37083d7fb98ebaeb9868fc11a908d67735f8d5c05c1f3d7cda501

Observation f01b7fbf-26c3-4dd2-8114-46bfdbe97cfe · inbound

Retrieval with Multiple Query Vectors through Anomalous Pattern Detection cites this paper.

Retrieval with Multiple Query Vectors through Anomalous Pattern Detection AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:21:04.904857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T14:59:11.547082Z digest=sha256:fbc911c380d242b09cfe5411a12e90efade26e6fa5889b74922fd0f9e08ec794

Observation 4801a16e-23cd-46e8-ad57-4cee300fb83b · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:13.644833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:48ff09d0413dc6d4ca44a2347d566f8e945573bf36eb8e57a1f32d89ac1160e0

Observation fa202f1e-c2a3-40a0-973d-319de57059ec · inbound

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels cites this paper.

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T21:39:24.661259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T12:07:02.778631Z digest=sha256:3777aa361a3d6552473d55c3d1bfad0067453dab9d46e333ead3d8051fb481d0

Observation e66af813-f6b0-4ab6-a8bc-9b87f23dbd9a · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.376279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:50:17.399580Z digest=sha256:9f5eb800e3c2527a283e631c04df26a638bec9a1cfda9fa9ba4f59dc22f19382

Observation db573a90-5090-4e7c-9177-03035981bd0a · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:29.004498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:20:32.494840Z digest=sha256:9b9fadc3dc723266038f29fe98b177fedfe633e97088897eadb6922409480a99

Observation 28073fe5-0c47-42d2-b035-1b59a0baab60 · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:08:50.584214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:97a1c8c0354d620f56a81e6535a001aa0ad8dd2b8be908e5ab64c701057d7ed6

Observation bf96173f-7fa0-423c-8fe4-199467748f0d · inbound

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety cites this paper.

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.532077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T10:38:45.689156Z digest=sha256:5b50f351d209aba6187df52920790d33f0297bc9978fe2ba80947328d91d03fb

Observation 0ac190e0-7e1c-45ad-b02c-b6af840a6912 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.986970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:9e3c29652d9f3116ccba10e1d570f110999cb8df07b810075571d692e4e3c615

Observation 5f66702f-98dd-4a31-bfdb-06bb349e4fb6 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.210043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:1515d58637ef87d77f1e8e7cdf35905a9b3bd2f35ebd598855133320ebc1c711

Observation 3e03a6a5-bc49-4674-8f64-267cfb3e490b · inbound

No Safe Dose: How Training Data Drives Unsafe Image Generation cites this paper.

No Safe Dose: How Training Data Drives Unsafe Image Generation AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.542541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:41:53.768791Z digest=sha256:5677c7c24dd008dffe97ad6b92ee29566ebbe7d087d6d9e556beea5d238c705c

Observation 0133e07d-be5e-4f4d-b821-ca38b69ea7a4 · inbound

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories cites this paper.

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.791513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:33:26.600072Z digest=sha256:bfa62dd0ae0b14e7526f259f62b96bde1a9c50b39a4c55ead074e03859bc493a

Observation 0cb02992-0286-46d6-a0a2-f189af52e861 · inbound

Next-Billion AI Index: The compass for AI utility and adoption in the global majority cites this paper.

Next-Billion AI Index: The compass for AI utility and adoption in the global majority AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:35.732962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T19:41:14.664207Z digest=sha256:48fa594a5a169f01b1b0434cfbd204319686bd691265a2e6e3a1062d8b414a3c

Observation 0bed949c-135b-46a5-be37-15c7151858c8 · inbound

Can Data Work be Reparative? cites this paper.

Can Data Work be Reparative? AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:21:30.122134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T15:05:14.689451Z digest=sha256:88c049f140a943fbccbf378678f2ec447534b481bc179fb1756d26bc57456fe8

Observation 187eaf4c-0767-4ac9-9c9c-5e6dd4e50397 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.632648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:fa4d8dcf6510eb259c48b062833fee2fe5943d7e28eddc869658dc381319f3a6

Observation 22711662-d4ec-4f49-b298-e6c63afaf4fa · inbound

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming cites this paper.

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:09:34.828204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:11:40.088809Z digest=sha256:e0c2d18ff28209d44700bbd8b0d1bff69021ab4be0f3f3651de6b1290d6fcf0a

Observation de6bfeee-e473-4212-a9fb-9f11af126d6c · inbound

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing cites this paper.

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:18.830699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T06:31:55.631719Z digest=sha256:0e9399a3024532c1fb85f719c75348f1bb1f23764dbc33f820fefc85a9bd57ee

Observation b023883e-b65e-41ae-91fc-f0cccb70b193 · inbound

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability cites this paper.

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 124

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T13:54:58.599385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T13:50:54.083165Z digest=sha256:005eb4befaa6a3072190f2e056d98c02a0b66f87e4c06afcf19d26cd3ef8cab8

Observation 3dbe5de8-58a5-47f0-ab86-8c15e9dc5dd0 · inbound

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem cites this paper.

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T10:29:09.565742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:29:09.565742Z digest=sha256:a38cfc61ae4b9af06fab7681510c7802fe5cfdcd829ffaf33e45b65c114c01a0