Pith. sign in

Paper Citation Record · LEDGER

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

As of 11 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2605.00689.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00689 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T19:14:31.076067Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact4
  • verified fuzzy42
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbf7c867-4046-4f18-b971-9058db8276fa · outbound

This paper cites Gpt5 system card.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Gpt5 system card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.643832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:523047989f9ff4cc8b8f0b42ceab4e63008ab5973270f0924338c53db569f85e

Observation bfe9f411-a038-4311-bd79-dc0990d226c4 · outbound

This paper cites Gemini 3 pro model card.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Gemini 3 pro model card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.652690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:0ea32d3c726d06c62b4dc26a5b6b471649880c29502c4474e58c42ab539dc11c

Observation c84a14e1-1a4f-44bb-8b48-fe2b0919d5cc · outbound

This paper cites Meta llama guard 2.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Meta llama guard 2

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.664154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:fe2b33da7a1793299720f1c5f7395c9058987ad2840751b1305ef034def4b896

Observation bccae67a-138c-4612-b84e-c10de03efbc0 · outbound

This paper cites Aegis2.0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Aegis2.0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.683922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:b795e00e55618627213350ca9ec5b539ea223985f4e58ff3041610e64699d2b1

Observation 7692d6f4-10db-4e08-9b35-7c583873b478 · outbound

This paper cites InThe Twelfth International Con- ference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models InThe Twelfth International Con- ference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:04.896363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:4124b80a64fe3cc87c52c544d28820580cbb5d380f0096cd31c793011aeb1a28

Observation 9572f68d-7634-4fff-a245-4aea309d24bb · outbound

This paper cites Polyguard: A multilingual safety moderation tool for 17 lan- guages.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Polyguard: A multilingual safety moderation tool for 17 lan- guages

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.710886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:1d86162139e17e572abdeb933adadebf70f395d1a7bceac9767906acadf9e309

Observation 6ee575a7-71b0-483c-b980-ac83ad56f12c · outbound

This paper cites Qwen3Guard Technical Report.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Qwen3Guard Technical Report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:33:37.835217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:0c04289c4f8371b83c657c77722a20e3765ac054e988516b9831be55152b7f7c

Observation f5de8075-f32e-408f-b0b5-fc452f05210f · outbound

This paper cites Eu ai act.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Eu ai act

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.628986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:f1798d9b02821076be1f29d31220dd385251ee0ab877fd844cecc552950f7444

Observation a6439295-2e7f-4e83-b61e-b7a056f425b5 · outbound

This paper cites Joshi, R.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Joshi, R

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:05.088390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d73c7ea50a3e5dca785deb6ab25900306223069dc8cac26137545b22fb47289f

Observation ae71efff-2331-4a56-b63e-696b5c59875b · outbound

This paper cites Rtp-lx: Can llms evaluate toxicity in multilingual scenarios? InAAAI.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Rtp-lx: Can llms evaluate toxicity in multilingual scenarios? InAAAI

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.636572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:0ce4309aa767bf6062eb95cf8002378073f137442aff9deed947195c1bbf97ca

Observation cec8f723-e7af-41c0-b763-21d565f26210 · outbound

This paper cites Polyglotoxicityprompts: Multilingual evaluation of neural toxic degeneration in large language models.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Polyglotoxicityprompts: Multilingual evaluation of neural toxic degeneration in large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.621629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:4b1d23cd496601647c692fbddfc586ecb008c667a10c5f7362fd7c47414d1418

Observation 26bc4f88-ae2f-4c1e-b1af-9079f8a5bb03 · outbound

This paper cites Multilingual blending: Llm safety alignment evaluation with language mixture.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Multilingual blending: Llm safety alignment evaluation with language mixture

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.504981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:cfba1c3791ec26d5248231f1e516f8356f19057d208c578d68e469fe4af3c470

Observation 4f8522d9-d01a-46b1-a329-8a81b76581fd · outbound

This paper cites LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:47:05.810661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d98c9284155b311bfa8cdca47c19ee6d7e1385263a40a2246a9903e2c9b84259

Observation 8c1f389d-771b-4170-8fc6-6ccdb87f8028 · outbound

This paper cites All languages matter: On the multilingual safety of llms.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models All languages matter: On the multilingual safety of llms

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.606193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d9ec0703349143b1fed31b4d6da2876183a37d6004639668c5148c3e07fe6104

Observation 7878118f-30f7-4d2a-aca3-9d8aca38be7a · outbound

This paper cites Multilingual jailbreak chal- lenges in large language models.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Multilingual jailbreak chal- lenges in large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.613890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:63fc438f60c9b554079604e016d925ec7248121f47cddb912fc0ca3236b5ed88

Observation ced81b66-1487-40d7-9629-f050518aa71f · outbound

This paper cites Meta llama guard 3.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Meta llama guard 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.617633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:5dcb3680fb916cbed1b52d6409a31908daada4f9228c2fb223737a02b55b6808

Observation 19672b90-168b-42ff-918d-18b776a2780a · outbound

This paper cites Llama guard 4 model card.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Llama guard 4 model card

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.609789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:9648c474cfc68cd54a0a3a23cb1ccc33bd6fdee767c0bc34d5d77383be35119d

Observation 659c54ec-38c1-42b1-bbbd-a90536dad5fb · outbound

This paper cites Mrguard: A multilingual reasoning guardrail for universal llm safety.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Mrguard: A multilingual reasoning guardrail for universal llm safety

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.632567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d8ee329825e23046cebd5804080b402ef3ebdc556c92cc6aadbfec99f1b84d0f

Observation 37cd6d0c-67a9-40c2-8824-09eaea70f79f · outbound

This paper cites Introducing gpt-oss-safeguard.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Introducing gpt-oss-safeguard

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.702683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:6cf2ac3319de6867e3afd948613e186e335fd9b89599e2705788d1740f02385b

Observation db883f73-9541-4f28-a634-053a2f351c8b · outbound

This paper cites Polyguard: Massive multi-domain safety policy-grounded guardrail dataset.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Polyguard: Massive multi-domain safety policy-grounded guardrail dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.544108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:795b56a638534088bee82719c96f154560bff8e3669b5f26eae803776a696a81

Observation 6d596321-94d4-448e-ae54-50f522c2ea21 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Jailbreaking black box large language models in twenty queries

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.458961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:6e0344e23f0bc0c4ccd3e1b7a60b4964ffe795867e796e894e3d0af8e678ee0d

Observation f082fc59-1ef2-4896-805a-14fe3f1c6c7e · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.455348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:f333bebab6f7d5121d1d27dde0d16b95c7116773c631396abc3585408c50bec3

Observation 8d4f7f54-3764-41dc-b1d5-acfcc480a5d2 · outbound

This paper cites Qwen3 Technical Report.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Qwen3 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:47:06.021702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:60d1d4a3836ee21713690d54c48d5534b90ad767690137a2f10121f264b43067

Observation b4eaeafd-c112-4d66-93c0-d209f23046db · outbound

This paper cites System card: Claude sonnet 4.6.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models System card: Claude sonnet 4.6

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.587268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:5a7e8c09b468bf1c81efa6631e00ffc19c326ef6732bf8a46e0646af57f2572c

Observation 628a4d3b-c71b-46a9-b4c9-e0fcfe05bfc5 · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Qwen3.5: Accelerating productivity with native multimodal agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.447303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:ac5971c0f2cf73619f8198878cea4008a23ace9071af3d3541799bdc15aae98f

Observation d1b16d28-a25f-4f01-8cba-f51c0e6182ff · outbound

This paper cites Grok 4 model card.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Grok 4 model card

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.579994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:5669208203b976ac88f964c006130d17a37ae15eb6e56c0f498367c683b93f72

Observation a0271195-a98f-4632-883f-3d590a767cc0 · outbound

This paper cites DeepSeek-V3 Technical Report.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models DeepSeek-V3 Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:47:05.421350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:ff581fb82d3793c32031c877d75d0250c72a6f58dee5fca688a00dd5fd89d0c1

Observation b10241d9-a8e8-46fd-a168-138525f3c609 · outbound

This paper cites Fast- dllm v2: Efficient block-diffusion llm.arXiv preprint arXiv:2509.26328.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Fast- dllm v2: Efficient block-diffusion llm.arXiv preprint arXiv:2509.26328

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:47:05.991372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:64d218f8d362d460a427918fc802eb8edbc3eecc1d2045a4d10a9fbef7f58199

Observation bc024da3-f2c1-4a64-b3b4-21d9994da8fa · outbound

This paper cites Code-switching red-teaming: Llm evaluation for safety and multilingual understanding.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Code-switching red-teaming: Llm evaluation for safety and multilingual understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.493906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:ba279457e0fec5a88e3db8079dbb6aace95cce8607cba346ce17b84febad475d

Observation 1dec8f30-51c7-4845-b6e2-70e0df6badc5 · outbound

This paper cites omni-moderation.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models omni-moderation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.490562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:0effb5c02c47afbb4dbf262d7d1b00c6ab512a711e53631ddd2042a61b49e9fa

Observation 3d01ea8c-5204-41b5-b1b2-e08d88d33c6c · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.497659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:34466d56fce66aeee688e7e56b30c4f312a25ed34b20970493b4de40a720da54

Observation 8ee356fc-1611-4f79-9fdf-52f30bbcd08a · outbound

This paper cites Realtoxici- typrompts: Evaluating neural toxic degeneration in language models.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Realtoxici- typrompts: Evaluating neural toxic degeneration in language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.508543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:b99ffc401745b40811ca36b516ccf355116f7156e8cb091feb67d582576ddc42

Observation aad9430b-15f3-4bd2-bdc3-bfffcefe94fc · outbound

This paper cites - Break them into atomic (single-action or single-concern) rules.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models - Break them into atomic (single-action or single-concern) rules

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.660672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:49f300956c89e68fb95ce161756b20a55a1aefcb46bf1d0472431b17dedbaa90

Observation 7ad28e86-6967-4c15-b25b-b8e0e0962415 · outbound

This paper cites - Combine them into a single unified rule that preserves all important details.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models - Combine them into a single unified rule that preserves all important details

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.583650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:53d199f3a699d3e4f9a9d338ab65ca69741153e32f321bef7602c78072cc7eae

Observation 4266e94b-7606-444b-9563-52876a93b3a9 · outbound

This paper cites - Each category should capture a distinct type of safety concern relevant to the behavior on Regulation.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models - Each category should capture a distinct type of safety concern relevant to the behavior on Regulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.591358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:7350873bcc1097d919325ad77bb15e923a1a15ecac393b7aa89e5469006d5331

Observation 2402fc02-3b51-4107-b93f-8881814d0364 · outbound

This paper cites - Ensure each rule is concise, precise, and consistently formatted.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models - Ensure each rule is concise, precise, and consistently formatted

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.451475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:0a81e58194d35441d4dc8df395fe82625db42810e5d20c2ac8d6537a3f41b437

Observation f068b681-c184-437e-b022-97345b882e1f · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.569428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:c3ac85a0e8a7ac847a3109e8f0e30f8064355d816e61f98c9ef0fdb68d3e7ff3

Observation a056c8ca-a3e8-4970-b19d-3964691fe874 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.687535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:461e8686a2315438be2379c922a03f43ea1098b778047eb1b7c37da78419a759

Observation 49c51550-ba3f-4348-9507-ebfa92499690 · outbound

This paper cites Use continuous numbering across categories.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Use continuous numbering across categories

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.738959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:4989a9b641cb45bc0a85574586ec67bcc94a3098f8ef159a9c482a617b1eeead

Observation 697582fe-9cb1-434b-8c99-d08fe8187e89 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.566119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:f18280901b7a32bc780b8513fcaba87b57737d7625017f7a8427831454c94ecd

Observation f0bf520b-02c4-43ef-a35f-6260279de9c8 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.573024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:ed823d621c05bc06cfe8a34f8ead702b51c1a4ddc8cb6c86b765724c0f8553a5

Observation 84659945-b024-4a7e-8b57-1dc9cb2443bc · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.640270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:9d449bf1e66315fc5330463885bd403a717fe6821648130f46582b49ab966b82

Observation 6dd8fb92-782d-4c48-b6b6-ae78cd11e144 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.625407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:6346be8cfcaeb12ca2677b84dd960aa8dcebd44f74a8aedb6f6cb2c366d11a2b

Observation fb04dc98-f14b-4075-9391-e1613b11e94c · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.699008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:cf409bdfaab1e80cb38529c7bbb83e245751dfa181fc0c863e41b29923cbc171

Observation 81edc4b7-52f6-43d5-82b5-ce3a7c274b2b · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.675987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:4844f632e7abac035c502c4cf7f65e610bc2532122f2581da8f9a564163524c5

Observation 3febc151-d21e-441f-a437-1ba0b597330a · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.555137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:60f391a6b7f4862bb66d460ad043b463a22e407298a9d866c809f6bcc35aece7

Observation 82987978-848d-4fa5-8a76-4180b0586458 · outbound

This paper cites Remember: Create prompts that would realistically test an AI system’s ability to detect and refuse requests that violate this specific rule from the {Regulation}.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Remember: Create prompts that would realistically test an AI system’s ability to detect and refuse requests that violate this specific rule from the {Regulation}

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.558887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:55a964661a257335d34f1d7dd290f0ff52a20f5e0d68003d7f85291c7db9afd0

Observation dd30d00d-82eb-42c2-bdef-d1f5644ce49b · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.562484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:be923774273ca0bf9ed2918aeaf820af39dcd7715387be13529a19195f9e0087

Observation bfffe5e6-6efb-4928-b1ea-fedd6285b3bd · outbound

This paper cites This includes idiomatic phrasing, tone (formal/informal/polite/honorifics), and references to local legal, social, or institutional context.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models This includes idiomatic phrasing, tone (formal/informal/polite/honorifics), and references to local legal, social, or institutional context

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.595419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:1364d1bd0f3f6341d0154f693556b4b32e9ff56f665bd16f79b160d8c34ffc7c

Observation 0d8533fe-05d8-46eb-a8ca-044325895368 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.529716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:5bd961472964a17816b57c74bce01ef605f817d51a2c80c52304566be427f84d

Observation 18dab9f0-2964-43f3-a133-09e47fc68b23 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.533054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:1024fd47ea34659cafcb6dee24b3f2a4b740c2e33f72b217eb05019194c9357b

Observation f9be1d3b-dd29-4f24-b0c7-783bbf5c86a8 · outbound

This paper cites Output only the refined unsafe instance in {language}.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Output only the refined unsafe instance in {language}

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.576452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:9522ce8f662ae5e1397cd61ba11f6851c5061edae4d8a3182496162d9bab14c2

Observation 5b59e4d2-acb7-4cf3-8f86-f7991db30e5d · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.599092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d1a23fbc7de09a910b435d390e7346b1bb34caf5a7a852f2abf2c7d3a79e9a7d

Observation 813d4814-ef4c-42db-9602-b0aea031235f · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.691116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:fc0b93eb1c12985f2dd367b3149ea8e17143ad8cf87a0f9b3e15abb6c8808847

Observation db64166f-a4ec-4ad4-9858-1b4219bc8855 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.474962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:342907dc9f50fc5fd3f5b70aec17aaef1f197e1eac0862c9ecfc7297e68fbe4a

Observation d0456f31-6312-410f-8d77-93f3347b01f8 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.672029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:7a1db974317dcd26441cd7ea890348922b2f404bd248b166d8761798d29a106f

Observation bff727b9-04c9-487f-a5f6-ffcd2a37cb15 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.466816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:2a3f8b7227ce66795a19e151aa4061bbd384af921681537c50d68136e0926b34

Observation bb2bc63d-cb7e-451c-ab13-97be0b41c72d · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.667940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d7c86564f2e7ca5aa7fbc3ea807d45938732f6b0e456a1968a59c8cfaa18ae49

Observation 67d779f2-c3fb-40e6-91c4-1887cc217e85 · outbound

This paper cites This includes idiomatic phrasing, tone (formal/informal/polite/honorifics), and references to local legal, social, or institutional context.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models This includes idiomatic phrasing, tone (formal/informal/polite/honorifics), and references to local legal, social, or institutional context

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.551780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:bb21ca752d8d111122f512c5a35d0f5cb097fe7ad57ef782748f8ec8ec2f282d

Observation 93b137a4-7aac-4c5d-bd90-c52231b0aa40 · outbound

This paper cites You must think, reason, and generate entirely in {Language}.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models You must think, reason, and generate entirely in {Language}

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.471003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:3d9497a18451b9a804776ca04302528773c0a8f7842177ad566c27e64c6aa526

Observation 07d5964b-6445-444a-b270-885a35718c00 · outbound

This paper cites This includes idiomatic phrasing, tone (formal/informal/polite/honorifics), and references to local legal, social, or institutional context.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models This includes idiomatic phrasing, tone (formal/informal/polite/honorifics), and references to local legal, social, or institutional context

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.656899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:5fcce14ce19489eda3ebd5ee333dae4fdca4fe8a2e2ee0d588a0262670ff311f

Observation a657935c-9946-4e20-a860-5ba9c65a78f8 · outbound

This paper cites I will consolidate my prompt by ...’, ‘prompt’ : ‘[new prompt P]’.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models I will consolidate my prompt by ...’, ‘prompt’ : ‘[new prompt P]’

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.648923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:a68202159d09fe089791dbd5a35e7feeed37ca8f045aefe5c7fea9a56f06dd31

Observation bc5694ae-9808-4de9-96d2-48352c9af8d1 · outbound

This paper cites Figure 15: Prompt for Attack-Enhanced Queries Generation.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Figure 15: Prompt for Attack-Enhanced Queries Generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.482969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:d3e73258bca6e23a295885d6ce3f8402596d6015dd31a6127f8b16b4c2e1bbd0

Observation 14de4d13-8397-4f1e-a01c-db0535e3b039 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.479045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:4077333ebdcdbd9b2cacad88aa3d504eb86c3090f52db129887ac014e0d50034

Observation a362dd93-4c3c-4716-a92d-3291881235d1 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.511872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:0698af730113d7cf2ab3a5a5fdda2d2c43b649058ce1db46882a4bbff1333483

Observation 36b46975-097f-4c8e-ae1f-281056aa6385 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.462752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:8a4e84615049f245ca4909b94953e6c2b576a649343766a639d88a58d7a21007

Observation de34d839-97d1-47f7-8e4d-e1d17381f1a4 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.540122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:e61f0359c7654bc8976c645334c7aaddb658d225ac373501f0f6ae55758e7139

Observation 64c2a0a7-7a11-4353-b115-326b2909a4e0 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.501214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:42dd30ef80cfabb46a89cf42be3a63629d257d9c4f41e314e68c4cd5d91458e1

Observation 6cb68ba9-4c44-425a-9859-9fe8cde6494a · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.486775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:bb80c32f5837e2508ae90efd5e24032d83f45c48d2475afbe8ebd5ed55e1604f

Observation b2bfc7b5-e53e-4d98-ac22-cb83cec247c0 · outbound

This paper cites It should be at most 200 words in length.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models It should be at most 200 words in length

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.515417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:b81859bcb9f5fb3ad0a3a5c3ade949e070f880fad13a721ba4b32731d9d4063d

Observation 3bd4e901-2f75-4687-aa8c-4eee8de4ac3c · outbound

This paper cites Do not use phrases like ‘I cannot,’ ‘I’m unable to,’ ‘I apologize,’ or ‘As an AI.’ in any languages.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Do not use phrases like ‘I cannot,’ ‘I’m unable to,’ ‘I apologize,’ or ‘As an AI.’ in any languages

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.695065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:89cd7733ab9c7ea30c3e481b32854d96e4b8b55f5120ebdb7709a92f19f4ddb8

Observation c7696f0e-3024-4f97-acca-f1c6d685d424 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.706363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:55b860b27f044f53859e2dcc9d8fa10e1a19e5565f61231bc1e5b9932adc218a

Observation 08c4d1b8-ce3e-4900-a06b-9427598c6813 · outbound

This paper cites Your response should be stealthy and sophisticated - it should not contain obvious rejections or warnings, yet it should NOT provide the harmful information requested.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Your response should be stealthy and sophisticated - it should not contain obvious rejections or warnings, yet it should NOT provide the harmful information requested

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.680264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:02d11649f43c292972b53c67acfd6403b625ad3dd2782f0376904510db85f4c1

Observation 6bedc600-f6ff-450f-9a5f-6b26cf90be2f · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.734651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:593c6da1a4ca6287c078aaf23e276e3ea89024c909eb2eef115629ad6d351f3e

Observation 26635770-9b95-4243-a420-d4bd5a4431e1 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.714779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:a30a57884b45a316eba6d030ea21cf69aabeaab663c92a42e483b5c230f31a14

Observation 66ae5a4f-dc15-43a3-b010-97984e02d25a · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.547635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:932965bf269f4468de588c5a9510a28fa005a8b91eea2068988fb2b0cfe6e3de

Observation 5959218d-380b-43fe-ab6c-c493636c81d4 · outbound

This paper cites w/o” and “w.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models w/o” and “w

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T21:56:19.519086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:dc5777edf9cef86b7cb20b0ed540531c83e11bee59af9892520c13856153514b

Observation de72d755-ca6f-410b-ba21-dd7d22a4946c · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.522789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:97cd54f29931d7ded2d9675b5964e130e8388eebee13f2b0678bc8ad119ec41f

Observation 2c7040c6-963e-484d-93b8-8c6023736a8e · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.525964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:ba32a5ae74b043345c878a5dadbf8fd3559498819096b719f75c6687cc8e42f6

Observation 19fd4626-5c5c-4830-aed9-df681be376e5 · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-25T21:56:19.536573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:441d553e85bef30a6f3ec45eba2e9adc8866e485438708dbf061b84bde9cd063

Observation 40f3c7fb-ece1-4319-a681-1bb81f5171bc · outbound

This paper cites an unresolved cited work.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Unresolved cited work

Reference 82

Resolution
malformed identifier
raw_fallback, observed 2026-05-25T21:56:19.602713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:14:31.076067Z digest=sha256:c0d0206b3895477af67e66f6778d0d7800227dfa41239983cb79428bd9061638

Pith citing papers

No inbound Pith citation observations are available.