Pith. sign in

Paper Citation Record · LEDGER

Probing AI Safety with Source Code

As of 21 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 0 inbound Pith citation observations for arXiv:2506.20471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20471 v1

Coverage vector

measured 4 of 4 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:52:33.805092Z

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

4 of 4 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2dee118d-3ce8-494a-9682-5ea3b603bfaf · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Probing AI Safety with Source Code Fine-Tuning Language Models from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:33.805092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:52:33.805092Z digest=sha256:a832aad41167345cf57fef73eb8ede22fe4cafc6a3b25d6a8b6e5f0beb086e45

Observation 8fc4d1c2-04b1-4cb5-8827-6806b1a5db60 · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

Probing AI Safety with Source Code StereoSet: Measuring stereotypical bias in pretrained language models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:33.800578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:52:33.800578Z digest=sha256:891a5cbda182e96e25a4a34a32b5ed9e6aa5c13978c0f70fd3e7cd7c214fed2a

Observation 9cba7311-bda5-4d6f-b1e8-e0453a1f6fdc · outbound

This paper cites GPT-4 Technical Report.

Probing AI Safety with Source Code GPT-4 Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:33.790582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:52:33.790582Z digest=sha256:bf08defaebd2197466375d7845db4b0127f846a5bfafafe8f8f3daf13253ce30

Observation 008df72d-9ed6-4bc0-bf0e-59aa3955ca21 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Probing AI Safety with Source Code RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:33.795772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:52:33.795772Z digest=sha256:c75ce6f3d48cfcfaa4aec2c9288595a2a45188379e9d4108d20b7ca60ffd6519

Pith citing papers

No inbound Pith citation observations are available.