Pith. sign in

Paper Citation Record · LEDGER

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense

As of 12 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2501.00517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00517 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:53:20.963955Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31561fc9-05c1-44cb-bacf-fbb7550e4640 · outbound

This paper cites Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.344092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.344092Z digest=sha256:9f3a1a5a02c7d3f33294a7b8e9eaab7aee80a9cc74c1c973ffb933f4951f6cad

Observation 087a0014-2e97-4466-9c8b-600f6881ab33 · outbound

This paper cites Can LLM-Generated Misinformation Be Detected?.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Can LLM-Generated Misinformation Be Detected?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.372871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.372871Z digest=sha256:10877f0ca7e719002c5cd517920c0979041507d439eceacd2328508d6327818b

Observation 9c9485fd-3d9b-48a2-a535-f7d52dd85d64 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.428520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.428520Z digest=sha256:b966696695ddd0f13dd47318eb3440dee09ebd2e6a14c455643c6647fe73faee

Observation bdc615e5-ee8f-44f9-88ff-58dd2205d023 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Scaling Instruction-Finetuned Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.480547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.480547Z digest=sha256:52b0afccbd87e026cb22afa79fa8cb1eeb1d798932c181dbf29eff18fd9cb806

Observation 2d85b8b3-0060-4a14-8675-58054c37313f · outbound

This paper cites Training language models to follow instructions with human feedback.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.557650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.557650Z digest=sha256:affc0d6abf6807a53277d5b5db5c4defd1de2c25e2d06c35ec4f6834132f8329

Observation fe6fb275-3db4-472c-97dd-84e84362f4b5 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.585262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.585262Z digest=sha256:766b7442c071907c1cdbe553fcbb256c46ae966a0293fd6db2dc263dff13667a

Observation 375f6167-abdc-4ad2-a10b-031127dcf45c · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.610290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.610290Z digest=sha256:e3db8ec7d9b99bb49ac9a2f156d0a5a2e2511312ff332f468a0f4e5388a2a85f

Observation 77e9c73c-6acf-4760-b655-cab31d36e36a · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense KTO: Model Alignment as Prospect Theoretic Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.621497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.621497Z digest=sha256:e63e2b7a405acac65b794189d9e66bb018f934c5d6059d6cf04ed8e01274232d

Observation fb1d6030-b4c2-4c56-9435-800cc824a30b · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human - preference dataset.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Beavertails: Towards improved safety alignment of llm via a human - preference dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:21.679484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.633584Z digest=sha256:dcea287a737b22a8ba4a20f68cf603ea08208376e67a82dec18e64af9bd7da9c

Observation aa68c2c4-eda7-44fe-a71c-452ae2bd883d · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.636788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.636788Z digest=sha256:21be0195bfcdb403eda5d4600f980b921d274ecabd3423952536b340a7c9ffed

Observation f0db6692-c76b-48b3-b8b8-14609638632d · outbound

This paper cites SecAlign: Defending Against Prompt Injection with Preference Optimization.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense SecAlign: Defending Against Prompt Injection with Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.639962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.639962Z digest=sha256:c85b8fef54183a47c61630140ba0f7955de90972b9f0687dede82b288b1732cb

Observation c5d680cc-0042-4fe3-b989-cce3c0ef94c0 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.644003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.644003Z digest=sha256:7cb49cf45823204788df21fc112600d2ad9bfb699cbca787c843221bc21598cb

Observation d671b371-d266-4ce3-87f8-5be94a9224b0 · outbound

This paper cites Superficial Safety Alignment Hypothesis.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Superficial Safety Alignment Hypothesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.650695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.650695Z digest=sha256:de05d8764280398df463278c59d2c053dc1ebf532a0beba00d2fe766960b529b

Observation c3bbda7f-1a48-422f-97ce-e35bce0d84c7 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:21.668623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.655282Z digest=sha256:7004652746351c059d0d476da9d0ede19f9466a590faff6adc5e154a3a9f0a36

Observation e796599c-3564-4d6a-96ca-4edc5fbbd1cd · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Multilingual Jailbreak Challenges in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.662431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.662431Z digest=sha256:5913d69ce2b65b16eebb882589722d43baf392b2a5ec28621106918ad7e121a9

Observation a82aabb7-4b40-4daa-a556-e8e0d0903a6b · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense SafetyBench: Evaluating the Safety of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.710337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.710337Z digest=sha256:ebdcf4876f7a2bab1ae974a929ba1b475aaaffac9b19fbf63f5cb883e268fdf1

Observation 502968ec-c692-4812-b17d-d6c4964ee2d8 · outbound

This paper cites CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.778379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.778379Z digest=sha256:bdfa5b151adf76a0d449d1435b6d4e43969c21532b88595d7a528abcb46c68ab

Observation 3c2d03b5-39a7-4daf-9977-694f7bab023b · outbound

This paper cites S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.842188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.842188Z digest=sha256:36a864c6f744811d2dde84a64042f76d73a3ea71290bdda9843abe4b0397989d

Observation 4e73d1ec-f0c0-4127-af19-9d94cf0b1394 · outbound

This paper cites A Post-Training Enhanced Optimization Approach for Small Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense A Post-Training Enhanced Optimization Approach for Small Language Models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:53:21.614718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.867904Z digest=sha256:a0968c581c818c430a6cdce040466bd0600d461363a4d321b3307cfc06db7c16

Observation 4a17f408-296d-49d9-a27a-5a28bccd50cc · outbound

This paper cites Safety Assessment of Chinese Large Language Models.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Safety Assessment of Chinese Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.901818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.901818Z digest=sha256:940fe3adfab1125e410196a6019eb92254009759e6bf60bef8e49df44173db7f

Observation b273681b-e41e-46a8-9ac5-55a50aa3bb35 · outbound

This paper cites an unresolved cited work.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:21.546684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.922825Z digest=sha256:3f69548677e4e759ac9dd8817e47f5be71c447a2979c9055bab9a3cdcf987bce

Observation 3e6070bb-a32a-4651-8914-a3630ceefd27 · outbound

This paper cites an unresolved cited work.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:21.344711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.940426Z digest=sha256:056ffd3de7ee2c3909ea49adf57221baa658b2255fe9e9c7a934769971f6cbdb

Observation d97dc54b-3a49-48ad-9948-f7eb9dc9cae8 · outbound

This paper cites an unresolved cited work.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:21.267410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.953353Z digest=sha256:b41e0b2af133b3e3c91667e74940b5f61e4945758e07a26d23722f18c17643f2

Observation e3800634-b248-4896-93cb-531948b027f4 · outbound

This paper cites an unresolved cited work.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:21.256475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.956873Z digest=sha256:468fa855b2ead3bd00d24447780a0282d4766bc185f42622b4a8685da6c6de94

Observation c0877385-0913-4394-9198-bce4417cd12c · outbound

This paper cites an unresolved cited work.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:21.246408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.960308Z digest=sha256:e898802b75ddceca7d707f83271ce1be3e063139b5122b637e4bb88f55300266

Observation 29c8aaf3-22fd-4f92-ac75-196ad80bdcca · outbound

This paper cites an unresolved cited work.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:53:21.236077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T22:53:20.963955Z digest=sha256:f1f90ae047c28c0dfcff2e2d4b33499eade833a4ff763d7b567aa963b2da000d

Pith citing papers

No inbound Pith citation observations are available.