Pith. sign in

Paper Citation Record · LEDGER

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

As of 23 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2412.16974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16974 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:47.845786Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T01:19:00.268857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:20:51.944045Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 415bb368-ae73-4b0c-9f91-12d8b4ee574d · outbound

This paper cites Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer?.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:58:48.517626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T05:58:46.685993Z digest=sha256:63735d6b92e456fe8611864c071758e44e5e9791ac8f1b9350dad2cf43480fd9

Observation 2f723bcc-5e0b-4061-bb80-8e5ad65cae2c · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A General Language Assistant as a Laboratory for Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.718112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.718112Z digest=sha256:45f4418ac1e3683d134a5b77e319bd0d3fd7b250aa7b41542e95cf3f1e16fd6e

Observation 7c321cdb-3502-4726-8c7d-e9a1d2d139a8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.724041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.724041Z digest=sha256:41ec5a3fafce16c91a153a6ca4abf9cbf22301888633f115273ae83806855cd6

Observation a5ebfa30-5335-4250-9fb9-e7f65f673208 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.728472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.728472Z digest=sha256:7bba1c16c6c0e89f3b0d46ad7853dd4075baa002e407657143986e97749a1f37

Observation b5c01f6a-6782-41b4-b126-d7dafd5ede13 · outbound

This paper cites A unified taxonomy of harmful content.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A unified taxonomy of harmful content

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.732208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.732208Z digest=sha256:3a2f0610b8838cc4815335ce0c9ada42dfafdbcc674c1d17a028197eee91b5cb

Observation 1ee5359f-6805-4140-8d39-29666863ae9e · outbound

This paper cites Information hazards: A typology of potential harms from knowledge.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Information hazards: A typology of potential harms from knowledge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:48.609003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T05:58:46.735933Z digest=sha256:e157fbd36e953c97631dec1c284d1cd08ff57d33f87d03978dad07d91086254b

Observation e3deb010-0cba-44ec-b3db-e02d6e300885 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Deep Reinforcement Learning from Human Preferences

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:48.543432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T05:58:46.740489Z digest=sha256:9fce8fc2a33cae927fe86aff54d5aa81823d761c51f67d524cde535c016673c1

Observation c16f6304-151f-48cc-a1eb-357e90be164f · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.743754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.743754Z digest=sha256:a64426a865be71c1210928d6c9db1eb89b5806ae5606856155dc3634865b7628

Observation 3e096ccb-11f5-49cb-8127-375e8a945b15 · outbound

This paper cites Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.747261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.747261Z digest=sha256:36147d9e32d7100d3e92b8d9f0f051f49cfd3490aa9b58849b74d5d49d57597c

Observation 6316c1e8-b068-4329-a185-52ee16a778f0 · outbound

This paper cites The Llama 3 Herd of Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.751649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.751649Z digest=sha256:2ab366441f44b2f9068ccf3e0380320126c5fa64c14ee2d06fd08616b11ebc46

Observation f93b0e72-4068-4084-80af-7614f987af3b · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.832039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.832039Z digest=sha256:9584c1f773034b788dae232f6f0c59d3253f4b556dffb1f2e302679f7aea1836

Observation 8a0e703a-f44b-47ef-a833-7a9fc7a41a1a · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Improving alignment of dialogue agents via targeted human judgements

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:46.954944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:46.954944Z digest=sha256:ee0b9b6a79f6898a2db6f6a5d119dc33d3917c8ad2d5b74a7c9d82d97b0b26ad

Observation 47694ab7-8434-432f-a732-d32674600a5f · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.006453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.006453Z digest=sha256:a917d2c540c5b826071a92f041dd63085df24906b3a98011e36863f2a8ad792d

Observation 3f792bc3-b2ad-405c-a89c-83342f6715b8 · outbound

This paper cites Parameter-Efficient Transfer Learning for NLP.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Parameter-Efficient Transfer Learning for NLP

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.010113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.010113Z digest=sha256:b56ede35742fd1d1942de1d87a96053bbdd9cc13be9cf83330e886a1938276d2

Observation 120b379e-a521-4797-b5ce-c113137b7350 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.014005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.014005Z digest=sha256:e028cc27d3bdcef66416150555276ac2c214f80026691e1d84d296af94881e53

Observation de8c7a63-d30e-4450-9c7f-3d5e70867b99 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.018964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.018964Z digest=sha256:9a8af9b0bdc8dc9b86ae7b0cb7a7e1b9408eeda7efeda4b66b8a693c6cece7b7

Observation d9b5ba2a-e286-4c71-8a01-45ade5ecd83f · outbound

This paper cites How can we know when language models know? on the calibration of language models for question answering.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs How can we know when language models know? on the calibration of language models for question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.022320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.022320Z digest=sha256:09b9ff9814dc3946ffacfbcaf0a79aeb0d86383e74d11be47e424219d896f000

Observation 90ad7620-3d9c-4bb5-b0fe-781b2272cba7 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.026583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.026583Z digest=sha256:6c500ca46baeaa0c91639429e45a87be840e9eb817e5feabd28756424817ccd1

Observation b679eb1e-a6af-4246-8abd-373faf51d28b · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.068771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.068771Z digest=sha256:e20309ea83f5b3847e02fa63a4cee8997a19c124cdd1361bb0ba31b67504dddd

Observation 198ee191-dbd8-4346-be03-fbbedc16531f · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.152264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.152264Z digest=sha256:15c683236c321565d5f371b63854f4d557c15f24dce65297e274c8be20b00bd9

Observation 5edf7f31-a96c-4529-9744-e4cba6d49f36 · outbound

This paper cites Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.276350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.276350Z digest=sha256:89f0c830d0c7c0efb934d6ead45c7459d81850339e7a97dcee0511c46bb13c43

Observation 3039cc7f-61f1-486a-8b95-ec772df7e680 · outbound

This paper cites Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.355053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.355053Z digest=sha256:940f1fbc6ba1f002c0f3176ca03757733738860a5fae57d9220c291b4d30e837

Observation dad2f2bb-285e-4ae2-aff3-5d44079fb857 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.359674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.359674Z digest=sha256:4a758598d2d904fc550605ad53e198bffa6d4bdb9fbcb000b9bb2974b1482491

Observation 1ca157f1-0b12-41dd-b9fd-f0be2384c948 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.363796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.363796Z digest=sha256:68ddca3728403bfad88409590496d0fa60b241ff17b31d3e4ccb701bd50e7ac5

Observation 4b4a198e-3d01-452e-a5cb-12c5f2dcbb0a · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Rule Based Rewards for Language Model Safety

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.367853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.367853Z digest=sha256:26b44552f0f538ab950df6f738080d82c70d6aaa2da39abc059361a1eceb2110

Observation dd4da059-4b99-4f46-a0a8-f60298839123 · outbound

This paper cites Crosslingual Generalization through Multitask Finetuning.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Crosslingual Generalization through Multitask Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.372522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.372522Z digest=sha256:6da7cc899d3b711dda12ef6d0be949373ef1a7a62c7b6b7bc687c00f720b524c

Observation f5efb365-cde2-4cca-b350-82b67dce859e · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs A Comprehensive Overview of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.375962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.375962Z digest=sha256:dd08943b74d66ec87927c819cae80b028f6e49c701ecd1be4406924533ad377c

Observation 5e7b98eb-8ea3-4eaf-ac09-7ef48f516b0a · outbound

This paper cites Model spec, 5 2024.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Model spec, 5 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:58:48.530851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T05:58:47.379944Z digest=sha256:d1f230d67deb1db1b67e37f5197dfa38184132df5860ce68bd0a3f42c7285f40

Observation 37522a85-fc18-4ede-a374-fccb088c0ca9 · outbound

This paper cites Training language models to follow instructions with human feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Training language models to follow instructions with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.383234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.383234Z digest=sha256:d134a4d6dd2b9bb4494fd8834e465df2247dda42ee8127d276935bf91719327c

Observation ef368ef7-9cb0-404a-b61b-ee94084db68e · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.450308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.450308Z digest=sha256:9755a07d38ede96efd39589a56e947ba221173de037cf2b460de412ba65dd87e

Observation f07b7287-65fd-4476-a87e-eb2c9188354e · outbound

This paper cites I'm Afraid I Can't Do That: Predicting Prompt Refusal in Black-Box Generative Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs I'm Afraid I Can't Do That: Predicting Prompt Refusal in Black-Box Generative Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.577628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.577628Z digest=sha256:f5456145e7055c5f6deb9cf5182969761da10a93e9e768973c3168b610a8490a

Observation 13a117a8-50ce-4c0a-a611-4c302f667643 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.656818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.656818Z digest=sha256:63dd15052a3bc44dd7069d855704a611ecde7349797fdfc2598a899bef404463

Observation f80c1d25-cd01-4831-90aa-7295f79daefb · outbound

This paper cites Proximal Policy Optimization Algorithms.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.660911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.660911Z digest=sha256:13721ada5c97936f922eaf8c41a4b1c779e82b2f2158106c443c3a321169ff78

Observation 9983772f-a372-4a91-94bb-3b5a5aac05f5 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.664382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.664382Z digest=sha256:e4fed290391daaceeaa49ac11366c9f25e84303649c232d8b9e3c88512de62a7

Observation 7739f1a3-c589-4f31-b004-e115abb647a1 · outbound

This paper cites Learning to summarize from human feedback.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Learning to summarize from human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.668432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.668432Z digest=sha256:0a334db4ed83cb96a9fc660ae00b826538d33cd65dc37c73bd1499a1f439dda0

Observation 95649f1d-3643-4c02-b9b5-e2adfc789d1b · outbound

This paper cites All Languages Matter: On the Multilingual Safety of Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs All Languages Matter: On the Multilingual Safety of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.672211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.672211Z digest=sha256:6138b615b380db17ddb26bb446d81ecb12c3b1bb73a51db7b49922a9e60683b7

Observation 371fa624-1bd7-492b-8625-d337a844954c · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.676389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.676389Z digest=sha256:e18b5c57e7d5e831b1ca600c7907662715895af01b2cda210b7f11012ab30b4d

Observation 7137eb1c-fdac-46b1-8510-1a675799074a · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.680455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.680455Z digest=sha256:9b352aa8abafc6534e4a008d30ad8e63909422a8cbcf32617c98fa5765b31af1

Observation f7596a72-84aa-4feb-bb5e-73f9ff413f89 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.683998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.683998Z digest=sha256:eaf0ae263d11327d3077a4ac98c092fb5389091a9fd6b0f9d3d21473e5b2f863

Observation b1393e13-104e-4193-aa07-f2339d4e7367 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Finetuned Language Models Are Zero-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.687891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.687891Z digest=sha256:c5ad8e49a1e36f2004dfa39540769b8c0aaa0f19229808d2e3eb2acfb1c3c8bd

Observation d7386ce8-ace2-423e-a824-a96e41172a14 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.691627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.691627Z digest=sha256:e26399704cd2175909e507277b304847d017c15db26a943e20843ce754853a01

Observation bcb95f06-9061-4423-80a8-35073be2d2ba · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.761384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.761384Z digest=sha256:43f9c2ddb42581d757d6b838ad536ad4660a811aaeb74363c9bdbbf651a4185d

Observation dd6638d8-c1ca-4a13-8aae-b667bffed841 · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.829233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.829233Z digest=sha256:844b702f1115f4fa755f6d0eba85960b060a5a27d4f0f8d1d5a51b91d096e48b

Observation b4963402-9065-42cb-9ab0-be3d7193a072 · outbound

This paper cites R-Tuning: Instructing Large Language Models to Say `I Don't Know'.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs R-Tuning: Instructing Large Language Models to Say `I Don't Know'

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.833858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.833858Z digest=sha256:fa0b8edb55e749881af10c55078835389b621195231faaeb03609ff3199dce48

Observation 2d8b7914-3235-4b3d-bec7-58bcebcbf364 · outbound

This paper cites Instruction tuning for large language models: A survey, 2024 b.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Instruction tuning for large language models: A survey, 2024 b

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.837929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.837929Z digest=sha256:11d35b21d30d2c36ba26871826ecefc05ea582a7e6bb6fc63b9a0b237fbb95a5

Observation f1611234-8817-437c-9df2-a37b07e1a696 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Fine-Tuning Language Models from Human Preferences

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.842057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.842057Z digest=sha256:275cfd394ac40eafe41b12e339545e1269ab8822602bd755bd82c9d6899b6fa4

Observation 49128758-1ec1-4825-9819-c663a818e4b1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.845786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.845786Z digest=sha256:d72b9ad59137aed7bf9abd27fd939b5c137ccaace5eb7a799fc6e850552beba9

Pith citing papers

Observation 6794e616-2bcd-4901-b052-33f277cc3f4d · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:09.935903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:250fbed0d1033756b3cc3bb3125d76dd93615996e56831216750494f1965ebee

Observation fded0e28-e891-4b7a-8aff-bed07bfe1418 · inbound

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts cites this paper.

RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:20:51.947425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T01:19:00.268857Z digest=sha256:861303e00f5c18a5bdb2dedb422731dffb6f8d59931fc7ed42de4b955c31399a