Pith. sign in

Paper Citation Record · LEDGER

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

As of 22 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2509.06807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06807 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:09:41.587858Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved58
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9dc5398e-f585-4da2-9d83-c20d7a04e93f · outbound

This paper cites Mogu: A framework for enhancing safety of llms while preserving their usability,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Mogu: A framework for enhancing safety of llms while preserving their usability,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.077038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.412785Z digest=sha256:596ded3f338b6e4121926f36df2d5cad7031bfe1a412027efa132a0ec954ceaa

Observation 4bb66a28-3120-4825-b8da-f5254d3cd532 · outbound

This paper cites Gpt-4 technical report,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Gpt-4 technical report,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.418878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.418878Z digest=sha256:5ba6f80d58d17e48ae455424ec0bc5c0a326819c1be850f8bfe3b09fa0f9c6fe

Observation a1558a58-a354-443b-b40a-f503ba477256 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.421041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.421041Z digest=sha256:1eee3d3ea87ec12fca89c0281d1045b89bb36c08d769269bd4524052c0ef2717

Observation edc29ab5-a9a5-4367-a3c0-04ee523c00e2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.423390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.423390Z digest=sha256:d0f1dbae8d59c93da0be991c2bfcfa33e99d86b0d008742aa91c2a4bc25923f3

Observation 7c0d83d0-d2b6-471f-a4a5-fae1eb0f28cd · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.426641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.426641Z digest=sha256:788193a116446cf18f963e98ac561a51f4d9412dac01401e61acbcd38b559c4e

Observation dc764439-8a44-4e88-95b7-ec1d378ff20f · outbound

This paper cites FLIRT: Feedback Loop In-context Red Teaming.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security FLIRT: Feedback Loop In-context Red Teaming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.429359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.429359Z digest=sha256:a8266004a45fe7cc3a93893641aad758c70f3a7ba5d5dd3e3c55dac8266bf43c

Observation 403519d2-2a8b-4bc2-8ae3-e9693caa9657 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.432374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.432374Z digest=sha256:d21aa488e701c1f938356dc17dc5d5ac8e78d09a5e06e06ff8953bcb94046e2a

Observation ead34801-dcd9-434d-b7cf-10580f06e3e0 · outbound

This paper cites A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.435223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.435223Z digest=sha256:0adfbb2cd2b65dd1ad8713261de1cbe0819319dc3452ca3971da21bb42b57b1f

Observation 12cedbfb-1706-4921-8499-8150a0757aaa · outbound

This paper cites Lima: Less is more for alignment,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lima: Less is more for alignment,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.063715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.437938Z digest=sha256:b14efd48c7c10aac991a1c162cdd88fd1c38036e948dca46e191029dc220705a

Observation 02db291a-5a5d-4e42-af66-a501dd18c588 · outbound

This paper cites Training language models to follow instructions with human feedback,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training language models to follow instructions with human feedback,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.440287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.440287Z digest=sha256:424f10aba968c580a3a26c0a3d7cc3010960a0f704d7e90e69e7a2202e9f8884

Observation fb86b84a-ac4c-4939-9349-a3745d597338 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.442984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.442984Z digest=sha256:b1f51af4c102139d45db22b3681886a2d56cf1bfe8c0e1207cb2e259642aca69

Observation 262bc832-9a17-474f-8b41-3324f2f1f943 · outbound

This paper cites A comprehensive study of jailbreak attack versus defense for large language models,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A comprehensive study of jailbreak attack versus defense for large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.050203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.445835Z digest=sha256:dd17e788c296e44aac8b2be0c828ced287ddb92283a4b7869bfa827ef3a0a69b

Observation 8f4adb80-c946-4fdf-a9cc-2de908f44f53 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.448210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.448210Z digest=sha256:63174a846c88f123a15c291218c26b1999404c363267423a698b44b0a71d2973

Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.450991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.450991Z digest=sha256:7618068f30e53181482d58b9c2ee77dfde9f13ecccd53591784a2314b76aa54f

Observation 6c917c07-f06e-4e09-9d7a-21b822e1d425 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.453737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.453737Z digest=sha256:033c5019a3469c050d994f658be2ccc822af092ea8e60c4bc1537d1e98ed25dd

Observation af156b49-2e8e-48a1-b665-2b82b8b86dd3 · outbound

This paper cites Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.456061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.456061Z digest=sha256:1673ea3ccbe6cd935745e45854f2e9b3429009e9f91b2858eb2b0c477173c399

Observation 3d039875-38b8-4456-ac5f-702912bac0dd · outbound

This paper cites Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.896274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.458494Z digest=sha256:cb065b17c4df358c5f0e0e56904ec307aa3500b631fc60ab8af37350b25c1eaf

Observation 12afa9f2-8495-4ea5-a717-ffd7c39d6a48 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.460875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.460875Z digest=sha256:29aa4b4dc5132488cb8f4c9daea4f57063f917d0236fd6ee288645dc3b0f8f76

Observation 0a575274-5666-4764-9913-0df741e67ef7 · outbound

This paper cites Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.463474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.463474Z digest=sha256:40f51826bbb2af00c2d74a29b09d2484755387a4ab8c59bc3f4c998ababeec62

Observation 98df20ed-9ab8-4277-9bf6-30125ef8485f · outbound

This paper cites The Llama 3 Herd of Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.465759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.465759Z digest=sha256:7f357e953da08ec281bdf40f1b0ad43547a63f616cc3d635d1c666739d8631fc

Observation f161f21a-9365-4361-ab7b-725b7b14fb53 · outbound

This paper cites A holistic approach to undesired content detection in the real world,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A holistic approach to undesired content detection in the real world,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.040961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.468117Z digest=sha256:1a9cf0c0ce9a00553fc48175939ad5679d671e12163d4c956f13a9eb2decd4d6

Observation 881ed3e0-8875-48cf-a485-a295d796bf27 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.470662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.470662Z digest=sha256:9372dfa340fc1be5ac27dfa1968b1556141da2a3fdfee8e31870a08cf33f80c1

Observation 1eebe795-ea34-4255-a580-d7a0b4b0d66b · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Certifying LLM Safety against Adversarial Prompting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.473339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.473339Z digest=sha256:23ddf4171720ca0bd92631dfda14d11bf585b70f6d9c870c1f8a1db33586dfad

Observation 8b5417fd-5bef-4bc5-8652-0547f8590e12 · outbound

This paper cites Lightweight safety guardrails using fine-tuned bert embeddings,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lightweight safety guardrails using fine-tuned bert embeddings,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.475614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.475614Z digest=sha256:ab65f25d36029fc9868e42bf5564ae3afc471a59d61aa5ce88ca18cb0fccd6d8

Observation 0986a88f-38f8-4ea1-9def-a33b45ba276e · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.478401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.478401Z digest=sha256:4b1ec72abafb4a8f4aa8d9739f84d7d42955580c904140b015fce752cb8f56a5

Observation e3c0882b-84ae-4945-a51d-c351a28d9757 · outbound

This paper cites Navigating the OverKill in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Navigating the OverKill in Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.840571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.480864Z digest=sha256:0e0ad64ae3699501689bbf62f77e8b563c3d74f4c2f81cd9e04d63dcba4ce17d

Observation 10a3c887-315a-40de-a43f-b04723e1dbe2 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.483523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.483523Z digest=sha256:af41f861c5025fce3c36b25d5dc0bc46be6986e1df7190f2d901f3ac61ed412d

Observation 8cfb6400-5c9e-4148-bcae-beb5aafdf18e · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.485922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.485922Z digest=sha256:51dd1cc395202ba07ccbb98aeef7f1eaa8ef36ef1e0f457bbe9249205c6634b4

Observation 3b83ca1f-1e85-426b-b604-e817fe22386e · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.488370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.488370Z digest=sha256:816eeeab23463079ab5d65ea956eeb6654fefc4a82eb84280e426153b36eedcb

Observation ef3557ae-8039-45ce-8b44-83a2979f2965 · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.491116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.491116Z digest=sha256:efa0c5e362fb0eb68571ce3e24906b9335cae7eca3d8987a25b5ff480a6d9454

Observation d0abc586-af39-432a-8053-36a2ccc3571d · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Survey on Mixture of Experts in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.493535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.493535Z digest=sha256:1f39cdd53f56d569f2e31ebb1e095306d978c8389205f22f1c69ba27c841ca6e

Observation 4632282d-a94f-4fc2-b241-76f34871da81 · outbound

This paper cites Uni-moe: Scaling unified multimodal llms with mixture of experts,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Uni-moe: Scaling unified multimodal llms with mixture of experts,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.495912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.495912Z digest=sha256:a7993d988bef730c0ce20cac598bb92331ac792de5d371fbc71f4da397f57721

Observation 8371582d-84ec-4ec6-b30c-437d0b14014e · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.498475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.498475Z digest=sha256:3a1871c797696ee05a3a5985dfa89594718e6a58c56e2535b9d9abf1e5071216

Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · outbound

This paper cites How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.500858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.500858Z digest=sha256:6eeca4ee7dd61eaf472af0406da28f44a575b2e98e85bf63f431f223741cd8c5

Observation d9becc34-a79c-413b-b080-229bd19f3777 · outbound

This paper cites Red Teaming Language Models with Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models with Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.503307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.503307Z digest=sha256:170ebb57f9e10af3bae37a003b983d194b3c2e53febf3cbd21ec2ba00b47eae2

Observation 426a32ea-c602-491d-9fc6-da5b8952a489 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.505782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.505782Z digest=sha256:833fbf558bd4767080220106cfe6d416ec5f914243264206d6d9e3cce27643ff

Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.508015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.508015Z digest=sha256:4f467e4fb554cd3ce4a8eb388fd5240abfe9ab053a915aaac7c181154406c814

Observation 9fe51abf-e933-4e20-802a-28412b67252e · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.510729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.510729Z digest=sha256:55446b812136c723f4325cbdaaa09dc035c88b8515b775b46ab28f10ad5cddaa

Observation 64bd5f15-fa7e-4f0d-b2ed-e26b626adbd3 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.513177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.513177Z digest=sha256:09e498ed6eabb3305052fa864aced322ee3e8a1f454357c7a5553fb0e4a4613f

Observation c6ba4a17-9d7e-4a19-8b37-598ca4bf46c7 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.515431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.515431Z digest=sha256:79b63489bd62c21d040fc723d3584b05d3c6d26b2521586b2175df541de787c8

Observation fef8e249-e825-45c4-9063-a5adb4df700e · outbound

This paper cites Jailbroken: How does llm safety training fail?.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbroken: How does llm safety training fail?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.517713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.517713Z digest=sha256:973155885289cba2d7a9cb9a448ff21beba249d57e4fadbe95080efb5868bb32

Observation fe644a2a-601f-4f18-b136-82dca3400e09 · outbound

This paper cites Automatically auditing large language models via discrete optimization,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Automatically auditing large language models via discrete optimization,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.023516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.520081Z digest=sha256:6f48920448eda6450d78e685c70db865dfc0b39033d3613b1e1709a2b8989154

Observation a6efadde-989c-4147-8d82-5b3e79debe8b · outbound

This paper cites Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.522170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.522170Z digest=sha256:6894888a4359b956cf708f96d3b7ac54a9615b615ab94050b6e8aed9ee12ffd8

Observation efbff674-42ba-4754-b406-ece032bb95c6 · outbound

This paper cites Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.524212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.524212Z digest=sha256:f1c2bd6479709a8bf5bc82d73872af4065668196cd93b5840e49ed36ec39892d

Observation 8277e1ce-0848-4875-b947-5de95d990291 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.526476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.526476Z digest=sha256:c4f097a00db3b3af6b0cfa2a26291e7f52b1872504cefe2ea896b6bbe84ca3a6

Observation 7fed6c88-2376-428a-bba4-3f70b0047335 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.528615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.528615Z digest=sha256:40097bfce10892ded4678a01381b3291e10835a0f4eaa3a962ea85352925d440

Observation 60840032-8432-4455-8d31-8136cb96982e · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.530996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.530996Z digest=sha256:8d808c7d3766df9f012314e36e915c62cecd152a555513b406d4036a94940706

Observation 0746e8d6-ca2d-490b-bf08-d3c599108cad · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.533142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.533142Z digest=sha256:573010c5a5f5d021ab7c74cf4adab8689c47c773225858441b8b4bd84be12af3

Observation 1a343dbb-5ea4-4ee0-8f64-0f9446a2b4ad · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminder,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Defending chatgpt against jailbreak attack via self-reminder,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.535213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.535213Z digest=sha256:cadfb8986f14dbeff315ddf5b3805e9324d4b2b6fb04339bc92ba4eefa89779f

Observation 6b551ba0-8dcc-4009-9481-9a872dae0065 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.537604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.537604Z digest=sha256:d7068f86ab2329d4043d7d43292852fdeaa92e1970e505306ec4c64abed7567d

Observation 41b7dcb4-1a29-42f4-825f-7e6bcf4fb978 · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.540879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.540879Z digest=sha256:a2544e7c9f121416e7137e3c8fdc7d7d8eb2e897eccf3a8c24a9f0ca14596c48

Observation 27d14d86-01b1-421c-a413-6a07a1aa0db2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA: Low-Rank Adaptation of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.543359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.543359Z digest=sha256:31679e101d6a400cf4d832260560f51161ec9dd5c7ea30ce816bafbfb339fca7

Observation 16c5c855-997c-46f1-a7ea-85d38f328d70 · outbound

This paper cites Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.545861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.545861Z digest=sha256:822bafde31d1ebc35950d79405123594418e399268fbefdba4ce902b51ba4b1a

Observation 10d452d0-ad48-457a-9e2f-2a8516dd405b · outbound

This paper cites A Closer Look into Mixture-of-Experts in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Closer Look into Mixture-of-Experts in Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.548125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.548125Z digest=sha256:9b81640523c8cca39ee95f770c6f5016449eb1871ebc40a808a72427284978be

Observation 5381ff8e-c8a9-4738-95c8-afc06896dcd8 · outbound

This paper cites Adaptive Attention Span in Transformers.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Adaptive Attention Span in Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.550995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.550995Z digest=sha256:36f59a7de5c6cc64dd0f86e0fe04b27910410b8ed708e6e4fd24fc919ecd4d1e

Observation 3e54ab47-115d-4f55-8f27-43baeb65ec61 · outbound

This paper cites Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.679537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.554098Z digest=sha256:3cfdc0e45b6e82a3ac666c90e5b9aea302ee14c5d41d0f96345017f66c3f322a

Observation 236bfd5a-4043-4a69-9b94-c05979f94c3b · outbound

This paper cites The hidden risks of large reasoning models: A safety assessment of r1,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The hidden risks of large reasoning models: A safety assessment of r1,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.556460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.556460Z digest=sha256:e2e6288fa0b56666ca633324a43d0b4d316c644207dba18be33b01799f5d36a4

Observation 3817d6ae-1a90-4122-ad7d-1712b98f17b0 · outbound

This paper cites Falcon-40B: an open large language model with state-of-the-art performance,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Falcon-40B: an open large language model with state-of-the-art performance,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.004374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.558613Z digest=sha256:0f031d625e605a6a8da916d2f8027f78fe07204b3f643ac787e2918d80a1421c

Observation 1e150241-98f4-4055-aadf-150db08c3e80 · outbound

This paper cites Qwen2 technical report,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2 technical report,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:41.995121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.560941Z digest=sha256:2a10b6191e57bc0518cb34a847ace5ada0fafe3f781d5e463df984c72326a692

Observation 72b6ea39-32e3-4be0-9da5-bd150d90686a · outbound

This paper cites Qwen2.5: A party of foundation models,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2.5: A party of foundation models,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.563287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.563287Z digest=sha256:2c327adcd31e383641bc9cde0cb32081f8c39be1ea3fdf20dcc212dbe760971f

Observation f203041f-f3e2-43da-91ea-3b60b6da5816 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.565548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.565548Z digest=sha256:bab724147458650ee78ec7c292a9b88a2fc456b0874b46c3aac33c71a341e624

Observation ccb3ab6d-60ab-4875-8b4e-ea684f463a66 · outbound

This paper cites DeepSeek-V3 Technical Report.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security DeepSeek-V3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.568088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.568088Z digest=sha256:ebaf99a40494781b983819a060dcdba8f4b3ee18648764570518d88ff486f8ae

Observation f663f21d-6fa0-4780-abeb-cc0ad5f7768c · outbound

This paper cites The unlocking spell on base llms: Rethinking alignment via in-context learning,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The unlocking spell on base llms: Rethinking alignment via in-context learning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:41.982843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T23:09:41.570710Z digest=sha256:bb1c6c6216d56e8606ed805fefe8a928b3f3ffd156036c26dabf65f7c55318c0

Observation f4736afa-cf44-4f03-8fef-46af2cc792f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training Verifiers to Solve Math Word Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.573265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.573265Z digest=sha256:b863f905d3210b9bb5026e95795335ca1c757c009e48ed3ae83bb030855942ee

Observation 89dce95f-0d2c-4792-9e61-17f46f8792bb · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.575944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.575944Z digest=sha256:01bdfc5c2749e00e6dc84ce75ac0c72a44a450de430ee773d213f1f9f036b67b

Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.578345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.578345Z digest=sha256:c6b38e26cf796e07a51355d885aa5d63969f375cc51966649ffd9b2f266d1617

Observation 4af7982f-a890-4203-9d1f-94fefb2fb084 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.580806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.580806Z digest=sha256:518ffc645e310527f1887c14b4eaa2707b9069b8bccd1d23715a2355b96ef223

Observation fc050fd4-0ae3-4005-a987-602c34e7b2ed · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language models are super mario: Absorbing abilities from homologous models as a free lunch,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.583242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.583242Z digest=sha256:22b9ab4e7edae34f77c5019b734adf969a5b61a56316ca05db9a60f46db04e60

Observation 9694cb0d-1a26-4cb1-8a28-613fb8689937 · outbound

This paper cites The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.585521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.585521Z digest=sha256:bb5e1d36ced1e0e9b444a1e7c1252672ec042ca7ccc97b4f4ee72d4e219fce90

Observation b591270b-0b7d-4011-845c-880f0d93f679 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-04T23:09:41.587858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.587858Z digest=sha256:09691b975c3cc7bfa2ce26174fb8b8e678ecf3534c9cace4c745f99c0a0176da

Pith citing papers

No inbound Pith citation observations are available.