Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:58.481821Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2504.16120.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:58.481821Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a36507a3-bce5-4285-9dd6-c4f21cec32a5 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content HateBERT: Retraining BERT for Abusive Language Detection in English
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320e32b3-3abb-4d98-bc45-50e738c89da4 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Generalizable implicit hate speech detection using contrastive learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 336e58c1-2cfe-4d1c-bf3d-e07727761498 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Toxicity Detection with Generative Prompt-based Inference
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60bfc48d-daeb-4c23-ad37-d75bc4ece791 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Interpretable Unified Language Checking
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942e92d9-3e44-4cee-9252-678314e1b8dd · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Efficient toxic content detection by bootstrapping and distilling large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f16f588-b224-4399-9ac9-37a9e14e03e0 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Autorag-hp: Automatic online hyper-parameter tuning for retrieval-augmented generation, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d7aff224-a7f6-4137-bc57-64d541e29bb4 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Enhancing rag-retrieval to improve llms robustness and resilience to hallucinations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6e635869-fc1b-4e51-9553-e43277539d36 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a57a7fb-b546-4e15-8b5f-fdb0cdd21f80 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Principle-driven self-alignment of language models from scratch with minimal human supervision
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e9d6f8f8-ada5-48c8-a1c8-67027b085e5d · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Large Language Models Can Self-Improve
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12df2720-929c-498a-81d4-be70bc9e27c2 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Defending chatgpt against jailbreak attack via self-reminders
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328eb7db-aa73-4019-a1bb-2600609ff86d · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Learning and Forgetting Unsafe Examples in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d057b93d-707c-4f3e-83be-681ddef8f05f · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4efc1b79-49e9-43bd-983c-8110d327362b · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content N-Critics: Self-Refinement of Large Language Models with Ensemble of Critics
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5dafa0f-1739-4a33-8471-f128150f9b1e · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Self-refine: Iterative refinement with self-feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c951aff8-9d27-4c74-b26c-0e1abec7eba9 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Learning From Mistakes Makes LLM Better Reasoner
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e95c6c9-21bd-4feb-aa81-0e9731522847 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content On the Intersection of Self-Correction and Trust in Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bdca6dd3-a4ea-485b-8a01-8c25ffec1435 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Self-correcting LLM-controlled Diffusion Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62211133-2455-4a2c-8fb3-50b068d890d2 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a84671f-7d73-4341-97fa-7a785fc02f8a · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92cceca0-5f6a-4e54-9951-7ad63ac85a96 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Large Language Models Cannot Self-Correct Reasoning Yet
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f6f01b-3dcc-4f2e-8df2-e8216a911a56 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a9de15-d9ee-471a-b5a3-04297d85cda3 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ebb4933-f62b-4d22-8a2f-3e5ab959c992 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Model editing as a robust and denoised variant of dpo: A case study on toxicity
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0024866d-96ea-4474-9adb-00402ab5ab4a · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Gpt-4 technical report, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 413fb6e2-2ebb-419e-a05c-91442c24af35 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Dai, and Orhan Firat et al
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f7858d28-88a2-4eba-b66d-39c88436ce30 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Mistral 7B
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a28d03b-e9d8-44c7-bc94-f7f4a124058a · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Gemma: Open Models Based on Gemini Research and Technology
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688c89b4-8cd8-42db-bd51-d31a57780f21 · outbound
A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.