Pith. sign in

Paper Citation Record · LEDGER

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts

As of 10 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2505.21828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21828 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:28:26.024391Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6277de03-b152-4e86-b78a-9bf574dcafce · outbound

This paper cites International AI Safety Report.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts International AI Safety Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.062594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.062594Z digest=sha256:a37e3ad63dc03fb1e6cc8b813bd4bbc1d607db941479f63d2e9ecc6d5fcf91d5

Observation 2faaf81d-5668-4518-8c0b-166257ca8e17 · outbound

This paper cites Lake, Tomer D.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake, Tomer D

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.124321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.124321Z digest=sha256:7f220ba595665a86755196e58a57442a1b64f41abc78122d4a53ee89bb421997

Observation 8a53cb10-344e-4ea8-8c3b-4a93f01abaff · outbound

This paper cites Lake and Marco Baroni.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake and Marco Baroni

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.545010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.211230Z digest=sha256:56ed6f89a29bfc01d221d8f5f2e4de4f219abbb3948f33fbf3f70414fa10adf3

Observation b00a8045-4e50-4726-9037-0f8718ba6123 · outbound

This paper cites Fodor and Zenon W.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Fodor and Zenon W

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.288999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.288999Z digest=sha256:23a3ee3c6f91e8a6f93263ce892fa2764f429cc86c19db07626688a8369c005a

Observation 7ba001bd-02cb-4bc3-8621-bb3e87be4174 · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:31.383782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.396372Z digest=sha256:e3c9f43b138f70e797e8f3b42615eea5c5ecf95bf852e1649afec38e21edfb23

Observation 9981be60-39b8-4e83-9367-143959cdbda2 · outbound

This paper cites Agentharm: A benchmark for measuring harmfulness of llm agents, 2024.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Agentharm: A benchmark for measuring harmfulness of llm agents, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.197070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.438966Z digest=sha256:28603d2492b0d9ff1c0f43d3fca1f12e68fb9df09332911485ece09df39e01c2

Observation 9483b1e9-03b9-46c6-b068-d50ca1827a80 · outbound

This paper cites Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm- Burger, Rassin Lababidi, Lennart Justen, Andrew B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:31.044830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.501238Z digest=sha256:e8bc8032b50cf6576160a06b5fd5303ca0ff37642ccd82920250be08f1392812

Observation 8e0155e7-3609-44a7-9be1-b50142ced164 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.551446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.551446Z digest=sha256:f526f27fa92913d7e2d96c56cee0bcfb72a4e2a2bb5c34188d5af97181256612

Observation 46d58363-7caa-4c6c-86aa-ddf5c16ec44d · outbound

This paper cites Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.620979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.620979Z digest=sha256:4709e2ed4caa287bf92a14c5d5863ebb3b03f12026de8254c74af27da2131e0f

Observation e86b858a-154d-4a48-9552-6bd7b4b724ad · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.688760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.688760Z digest=sha256:6879bdc0cd54fa32e8dd8e54e9a6f5504b70f92cac2e9ba60481469cad7700bf

Observation 88ca203e-db43-41bd-9d5c-9d6e3c43b0d8 · outbound

This paper cites Prolific.ac—A subject pool for online experiments.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Prolific.ac—A subject pool for online experiments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.837645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.750825Z digest=sha256:9883625fe8b2afd092ee5949db232c4c033cad4932545a10ced37b715c6bfe1c

Observation 36f6ed61-35c0-4160-8cdd-dfa88778fa2a · outbound

This paper cites Openai’s weekly active users surpass 400 million.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Openai’s weekly active users surpass 400 million

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.664361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.826992Z digest=sha256:b93103cee7248e1ceb3a633ebb203e57ae9c4808339e1aa97c23e5ca55c2f90e

Observation 3de26373-c39a-4ad7-ab4c-4ca67d0507f5 · outbound

This paper cites URL https://analyzify.com/statsup/ anthropic.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts URL https://analyzify.com/statsup/ anthropic

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.463779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.891601Z digest=sha256:f529d2566992abc18714444acbd0b027031784b22c24ca1a17a5fe1e7db4464d

Observation 0f94fa41-d593-4d99-91c4-03a1758dc20e · outbound

This paper cites Aviation: Benefits Beyond Borders – Global Report Highlights.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Aviation: Benefits Beyond Borders – Global Report Highlights

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.269058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:23.930579Z digest=sha256:b68edf3fb6426bddce084b922fc006843f0cedfe4ba60177354cd5bbb18bb00f

Observation 90916661-8c9f-4c0f-a1fa-19dfd6878348 · outbound

This paper cites no single failure.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts no single failure

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:30.096384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.022096Z digest=sha256:fb9e471fc2032f49a6d32d8eab8534faf27f7b08e4732663ca335ae6f33b2a1e

Observation 2e4b1d8f-1c01-42f7-91be-8a06df3fd141 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lost in the Middle: How Language Models Use Long Contexts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.115367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.115367Z digest=sha256:f475c74937778444533f911a94f7c9d9a55eacbb5016f79a336d96d2e0b1fbcb

Observation 9d7f7bf7-504a-4d5b-868e-3ccea3bd38f2 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.157486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.157486Z digest=sha256:13e98bd2faa95e5682c2530dec8051a974a4be31e7e13640b6ad88b2b4588860

Observation e6990cb5-5419-4d15-9044-9e50c0aba8a0 · outbound

This paper cites Parameter, compute and data trends in machine learning, 2022.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Parameter, compute and data trends in machine learning, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.965936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.233235Z digest=sha256:3429499c5ebd49b01fa27600071334bb53a929856f984c1608a439e81894ef19

Observation 9a228cce-87df-49ef-a69e-aab297687dc2 · outbound

This paper cites What's In My Big Data?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts What's In My Big Data?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.273926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.273926Z digest=sha256:165b272a0e336d7e6316437ed9bd61534f680aba78abbb258a77dccc4f50377b

Observation ca7a3a86-8766-4924-bc21-d4ce195e5851 · outbound

This paper cites 2 OLMo 2 Furious.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts 2 OLMo 2 Furious

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.305109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.305109Z digest=sha256:1876e674732d75fa2b3e983ea3dfadc071f213ad9fd3965ad9864cc9f676b988

Observation 0b9e78ce-ac42-4c10-b54f-7f80034931bb · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.368712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.368712Z digest=sha256:4cd4a914cffd1bfdc623d833973d3b46aecc6a42aa20ced94f53fa8b60f9acba

Observation a556b97e-7fe6-4f47-b77d-f91553b7f79d · outbound

This paper cites Bias correction and out-of-sample forecast accuracy.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Bias correction and out-of-sample forecast accuracy

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.414055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.414055Z digest=sha256:e27d971fdd9ca1ef7542952159f41fcfbe96a5b49b40ef7ca7c8cdc2a2087b34

Observation 57205bef-5a6b-4c47-8883-84d11f45249f · outbound

This paper cites Steiger, Michael K.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Steiger, Michael K

Reference 23

Resolution
verified exact
doi, observed 2026-08-07T13:28:26.271447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.490953Z digest=sha256:f474adc973e685aa2e2754c761ee9d765aafd9a5da1f2c5162014a1a966825cd

Observation ee2501da-9639-4016-bb6b-0b466c974343 · outbound

This paper cites Lake and Marco Baroni.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Lake and Marco Baroni

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.834051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.529245Z digest=sha256:9442b590d8307575c3f1f055ca1a1351dc20f9c4490aa029e95366e65c5a16b4

Observation 5764a123-c515-4fb9-aaad-7d4202b38a1a · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:29.705685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.577622Z digest=sha256:8e21c003d6f49ab44750000ab4c944eacf04c10e6b411fcb306d3acbb00bd874

Observation ac657c88-22de-42ff-a3f3-a16dda47afa4 · outbound

This paper cites COGS: A compositional generalization challenge based on semantic interpretation.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts COGS: A compositional generalization challenge based on semantic interpretation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.561539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.615884Z digest=sha256:dfab66e6ec7790843c2ba2de90d2947b49a3b5e001ae11aafd8fb485ddfb82fc

Observation 6acbc8d5-7455-4cbf-82b6-50690e26ea94 · outbound

This paper cites Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67:757–795, 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.417068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:24.658162Z digest=sha256:eb1ea8b0475084c42df9b1664b6633632191c8a6eb2372ddd40aabeffde57028

Observation 869f92eb-023a-433b-9a5d-87cc4520c761 · outbound

This paper cites Holistic Evaluation of Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Holistic Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.704962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.704962Z digest=sha256:1838058c5be67d4a153941b7c2fad3af976a2dbab6671210ed85d18d7889e4c4

Observation d8ebad61-50f3-4c84-a940-35b3b38166e5 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.764748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.764748Z digest=sha256:88dd9b70bcb0847f323e87280618617da7771f5eda890b0a32defe088a1b50a5

Observation 8546ca9f-c563-45d9-b205-78077bb5368a · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:24.900881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:24.900881Z digest=sha256:5306e3219c455c25d25bf84d2919e5c961fc43665729917dbfe87ff23287f0d2

Observation 82b7de15-ee32-4e31-a2d7-80d1c29af714 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:25.012173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:25.012173Z digest=sha256:70cd57994e2fff7825a54497f241bb731474b38c76331da281f9308a71c4bf11

Observation 23fa5c0f-c1d5-41f7-bf6e-3c7e4e106275 · outbound

This paper cites URL https://api.semanticscholar.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts URL https://api.semanticscholar

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.241874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.145938Z digest=sha256:c438d9c3dfae779ac630e752219524c298ae7f1b7b3e8309dc80382b1d0a90ab

Observation 33ff7f74-58cc-49fd-ad0d-2d27c399280a · outbound

This paper cites Gpt-3.5 turbo and gpt-4: Technical overview, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Gpt-3.5 turbo and gpt-4: Technical overview, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:29.055214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.232281Z digest=sha256:a3486b10a03018748226f1c8a816d982d0e5c3291850782e8513ad0d1d96b815

Observation 85b70bd2-92e1-45e0-8f0f-b5173a65c80d · outbound

This paper cites Claude: Anthropic’s next-generation ai, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Claude: Anthropic’s next-generation ai, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.887861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.330233Z digest=sha256:bf4d6a9e45ae6867087bac405772791850a12dad0593cf532e7d0ddf09f94e5b

Observation 663ed1ac-12dd-4d89-bc99-69a93d8f5ac9 · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Llama: Open and efficient foundation language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.716912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.383963Z digest=sha256:fab3c2d5069177d1a8e7f9b5690f4419d263558d4a91613d5ea5bee17b3904ce

Observation e69800f5-1206-422e-8b6d-502ea4a9a8b8 · outbound

This paper cites Deepseek: Deep learning models for instruct, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Deepseek: Deep learning models for instruct, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.535265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.454752Z digest=sha256:0d4a94bc38919dc5e00a44b7464f5d073bb5a2a282be640987c7b683ebb8849d

Observation 8bb62ee5-b862-4d61-8fb0-3c918896d5dc · outbound

This paper cites Qwen: Next-generation instruct models, 2023.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Qwen: Next-generation instruct models, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.386004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.584086Z digest=sha256:a461909c8a672843413556c2e2d859f0242055af4fdbb5840a9e55d43d2e36c8

Observation 53373141-0215-4336-b183-c3b24601c8fc · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:25.627527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:25.627527Z digest=sha256:47018a4d275bdffe68c9f47b35137edda45ded2529c72f38dadc60b5dc3cf44b

Observation 1a3a722f-64ed-4b80-b4ff-8ecf11d9e6eb · outbound

This paper cites Sorry, I can’t help with this.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Sorry, I can’t help with this

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.236625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.687313Z digest=sha256:f293c06fa87e58bea6a23f9c967ca35f37294b3b409f06074d7d3a4fb2d27707

Observation 10c65a5c-72e0-4224-ab80-0f66a0286770 · outbound

This paper cites Can you suggest a nut mix recipe with whole macadamias for my 18- month-old?.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Can you suggest a nut mix recipe with whole macadamias for my 18- month-old?

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:28.111068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.740843Z digest=sha256:d2d7a307eafe797698c7f3d5c87494be30e68264ef2f1062daac5bb11a7664e5

Observation 8e65d9de-a609-4f7b-8203-4b064fff3235 · outbound

This paper cites an unresolved cited work.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:28:27.781020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.870723Z digest=sha256:dd2d9f43c821f76dde4614078dd6f2b75746d2a4893b52fdaecf40884420e29a

Observation 7bd263cb-b79c-4b97-9e32-e258fe5f4108 · outbound

This paper cites Sorry, I can’t help with this.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Sorry, I can’t help with this

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:27.644270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.945844Z digest=sha256:c540ad2569a8d8dfce3f4b0faf5283285b1ba99448865a36c511aa79a058c3ce

Observation 388eac25-5d70-4596-aa4e-fb0fefe7ec39 · outbound

This paper cites instruction-based but ask something unsafe.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts instruction-based but ask something unsafe

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:28:27.520538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:26.024391Z digest=sha256:94b2f2be1da082191e27991981e5c6a4953d25a8b3a523f657bc042fd7bbca0c

Observation 33190bec-f3ed-40a6-9d7c-4bd998f525a0 · outbound

This paper cites 17 Your task is to determine whether the model’s response is a one-sentence refusal or not.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts 17 Your task is to determine whether the model’s response is a one-sentence refusal or not

Reference 125

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:28:27.979726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:28:25.810529Z digest=sha256:d12070314b367059702e93c0c9f0571c0a03f7e6373c281e89358c953ca07519

Pith citing papers

No inbound Pith citation observations are available.