Pith. sign in

Paper Citation Record · LEDGER

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

As of 18 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2501.17749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17749 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:37:17.038006Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:13:50.508020Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.350151Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 00626506-bf78-4424-b696-1cffda6c7dca · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.943712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.943712Z digest=sha256:c16f424ca93b2dc87b4508bb2d253f4eac30ea3098ba6b476a446ece2e2913d4

Observation 2b96bf4d-ed93-4bed-b67b-095f5ad09327 · outbound

This paper cites S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.948537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.948537Z digest=sha256:7b2378e2baa9940912bb7bb4aeb5d5ce67dc796756176d60fe3c1269423df0e4

Observation 10b0be7e-f833-4089-a919-4ac79cdbfcff · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SafetyBench: Evaluating the Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.952819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.952819Z digest=sha256:e061b342850ae0dd4f7bab204263b9c78481de29df5f7a39a5f760cb99c7f34b

Observation bb7698cd-626c-41ce-8617-d4761030ec04 · outbound

This paper cites CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.956518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.956518Z digest=sha256:5c4010ef80c318f4ff88eb86fe8b1e733193a11e08c46ece36e54339340a7f87

Observation 842d4427-4229-4d9f-b061-38f3e911ef4c · outbound

This paper cites SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.960430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.960430Z digest=sha256:1f73b0eaed775d3f306761436e0c115c2aa30cc2593cff884fb64fe98e713234

Observation d0ff1b51-4537-40f7-b0e4-8ba34959e8ea · outbound

This paper cites LongSafety: Enhance Safety for Long-Context LLMs.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation LongSafety: Enhance Safety for Long-Context LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.964203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.964203Z digest=sha256:f417397cde7fea9b366a87f8f8d46ae1e4214c8ecd532e617695ce1de9a46b7f

Observation bb969384-9c6c-4882-b7fe-83ea091a5669 · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.968365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.968365Z digest=sha256:7da01ef7bb785a37e311dcd75459b7d1b1738258d794550379d3b8f2b6a45737

Observation a893c6c9-40be-435a-b83f-d26c4f08386d · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset,.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Beavertails: Towards improved safety alignment of LLM via a human-preference dataset,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.314896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:16.971913Z digest=sha256:5835f78b52d8756b76ffd7588680ec01f11a251f47cf62a5cdf8ccbaaa1d3c55

Observation 47aba5ba-9f30-4c3b-9b95-6928a2e9bb0b · outbound

This paper cites SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.975584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.975584Z digest=sha256:7e7ea724dac74861122a0e83c69fdbf92ccbe4d855ae8e4e4699a87d1cca47cb

Observation cccad398-d7c5-4490-8bf7-dd5ffd2b1e11 · outbound

This paper cites Astral: Automated safety testing of large language models,.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Astral: Automated safety testing of large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.302892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:16.979482Z digest=sha256:458a509148a82e7680ee8f9fe010692907de30da2eabc1175919551705355081

Observation d79e7cbe-e727-45df-bafb-bce88417f312 · outbound

This paper cites A survey on metamorphic testing,.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation A survey on metamorphic testing,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.291051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:16.983004Z digest=sha256:8dbd6b762a1a03ca71d0d2349ea0beed91f3b8a052af3d75fe4280e76f74f003

Observation 04dee86e-138f-4109-a9fe-b5931220627e · outbound

This paper cites Guardrails for trust, safet y, and ethical development and deployment of large language models (llm),.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Guardrails for trust, safet y, and ethical development and deployment of large language models (llm),

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.280008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:16.986274Z digest=sha256:8c98bd348fe808b81cbf41341b18610bcee745112d06e97fb36fc9032659a698

Observation d8295da2-1dd8-43a7-b7fa-3d9105a14fa2 · outbound

This paper cites European Commission AI Act.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation European Commission AI Act

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.267036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:16.990397Z digest=sha256:f09f67dc6560f88056de2e3274100c0266b9d26958f7ebed441cba0f64721733

Observation f86c79da-9169-4d0b-8b5b-d75b3355f507 · outbound

This paper cites Artificial Intelligence Act (Regulation (EU) 2024/16 89), Official Journal version of 13 June 2024.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Artificial Intelligence Act (Regulation (EU) 2024/16 89), Official Journal version of 13 June 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.255058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:16.993873Z digest=sha256:6808d697de1f9c5581f35afea596506297eabe4985cd6841e31408fca52cbb2b

Observation 1cb77ed6-af48-4b3f-a363-3dafa0f2183a · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.998077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.998077Z digest=sha256:58c5b163f1fa8dee8f292b6840b7e240d9a30d2dab67a09562dedb9ccd1f0391

Observation 5ab876b4-fd61-4419-aeff-ba967d05e0b7 · outbound

This paper cites ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.002476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.002476Z digest=sha256:aec2ec87dd4379f5068356bc5776e73328ae68636eb68801ba05c69f71cd4976

Observation 32549227-2140-4849-890a-5b7dd32a6123 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation A StrongREJECT for Empty Jailbreaks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.006505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.006505Z digest=sha256:6b2689d7d10c4036247d5ab8796097f531ada02eb21c18de322716782ea05605

Observation 0797e6cc-8eec-43f3-af8c-bedab0117b73 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.010084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.010084Z digest=sha256:e9f884582e8a061302e2343e6493896bb257ebf4fde4fee7df27067d6a878610

Observation 90df0369-a45a-42dc-a089-0ad59a44f7b3 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.014284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.014284Z digest=sha256:7eb4206f3d3866c7741bc5fc60a5c9685807766dfa0bbe5fde79b26c91435f4e

Observation afba7695-fc6d-448e-bfff-2b153789b5f7 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.018055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.018055Z digest=sha256:779fce2fc9bb5f49bad180e99c2bb471896459e1d6a2d393d15e309ed53d31cb

Observation 38e85f31-01f8-42c7-bad5-766efcf9985a · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.021691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.021691Z digest=sha256:526ecf2d5b45c0f4d0d465832a31e412576128fff19fe73271bd806c30e2c957

Observation 1fd88780-7f4e-4dc4-a69b-db7ef0a9d757 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.025612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.025612Z digest=sha256:e02a6a59f5382384e9c6df8c6b0060f980489435c9fce7d787eb7faa9b0cb25d

Observation befa1f1a-7b65-4894-9e48-d125dd22d8da · outbound

This paper cites Jailbroken: H ow does llm safety training fail?,.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Jailbroken: H ow does llm safety training fail?,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:37:17.243309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T04:37:17.029662Z digest=sha256:a94538da7a19cdfd6a7030c7bf431084378d781deea96b32d5fa0864e18988f7

Observation f68c3631-350a-45ac-90f4-496ba9779108 · outbound

This paper cites WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.033434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.033434Z digest=sha256:c72a01795b864fc7f9c16b92e6c9d7d1143a006c107524598b7cb89c0823e55b

Observation 7d4fb587-da85-4b7b-ac9f-b386a5e3ed18 · outbound

This paper cites Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:17.038006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:17.038006Z digest=sha256:bb2a40f7967b762d0519b0044848ae9be7db987a348061e2937e4d4d8cc8dded

Pith citing papers

Observation f91ec8fc-2a41-4e0c-9663-6682edc638e3 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.353112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:a548068c52022ca420c5db5bb847a6a0f5075afa0731cc78ace566f9299badf2

Observation 15ecdde5-f430-4616-8d8d-7541c6213c5d · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.508020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.508020Z digest=sha256:b559588ce06f4b854a754ea6ce293b04bff2670e2d07563b6eb91b237c7defce