Pith. sign in

Paper Citation Record · LEDGER

Strategic Deflection: Defending LLMs from Logit Manipulation

As of 9 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2507.22160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22160 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:05:44.625951Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:24.267269Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 21742424-0228-49b6-b944-22f39c585070 · outbound

This paper cites Training language models to follow instructions with human feedback.

Strategic Deflection: Defending LLMs from Logit Manipulation Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.381035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.425087Z digest=sha256:4fc086974f0df085016bd542020235323bd373598f88df013116d1183ec3f084

Observation 16a914b2-ff48-422c-91e7-ace163ab7bcf · outbound

This paper cites Deep reinforcement learning from human preferences.

Strategic Deflection: Defending LLMs from Logit Manipulation Deep reinforcement learning from human preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.360768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.433864Z digest=sha256:753b6193c4dd3f28a92d548c2566e061b4b216c24974dbdb4cb94f97bf91539b

Observation dfa9d83a-feca-46da-a357-460ec701c679 · outbound

This paper cites Mistral 7B.

Strategic Deflection: Defending LLMs from Logit Manipulation Mistral 7B

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.442565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.442565Z digest=sha256:2740763c25322786756b76ae1945aee88af5cbcee9237b38b5e94ca6dcc57634

Observation f09ff4ca-dc03-46de-b451-fd2cba7fc169 · outbound

This paper cites The llama 3 herd of models.

Strategic Deflection: Defending LLMs from Logit Manipulation The llama 3 herd of models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.454732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.454732Z digest=sha256:eea27b528b1d2f996f13a77bd6333daa9b195afb80d62ce9a927f1b12cd70827

Observation 8b56acc9-7aa6-4136-ba2a-9940a912e427 · outbound

This paper cites On large language models’ resilience to coercive interrogation.

Strategic Deflection: Defending LLMs from Logit Manipulation On large language models’ resilience to coercive interrogation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.320699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.462785Z digest=sha256:1c97077c972cea9a28fc679c0da0832cc37bd1d52958b048ceac830e645a02b3

Observation 6ab8a4e0-999b-4b93-a1ca-820f02528e21 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Strategic Deflection: Defending LLMs from Logit Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.471626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.471626Z digest=sha256:375f1ea09f6e3cabc3c3583e6e015425a2ed9da5846a290230a7655b226d607a

Observation 5d7d4768-5c38-442e-900d-7453e798e9ea · outbound

This paper cites Jailbreak open-sourced large language models via enforced decoding.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreak open-sourced large language models via enforced decoding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.297838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.477735Z digest=sha256:c68eda32f6fb925df186a3898173f09528c53b19c7d6faccdb775a96179f6818

Observation 99440afc-c6b0-4c6b-9eb2-c10f3cdb3c0e · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Strategic Deflection: Defending LLMs from Logit Manipulation Safety alignment should be made more than just a few tokens deep

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.275737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.486012Z digest=sha256:030912c04e4e4e8ef42253a53189dc5a4c01e1eac2ca8d9e91e17d5d6678cdc7

Observation 180b2f18-aaf7-4ad0-86ef-3d9682ca6f8c · outbound

This paper cites Jailbroken: How does llm safety training fail? In Advances in Neural Information Processing Systems , volume 36, pages 80079–80110, 2023.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbroken: How does llm safety training fail? In Advances in Neural Information Processing Systems , volume 36, pages 80079–80110, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.251296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.491524Z digest=sha256:a675bad7106c846ce0ef01ee69879e069ee8abe6922bdc629a936698bc285eaf

Observation e838db4e-e370-4e8c-b125-56c08475e087 · outbound

This paper cites do anything now.

Strategic Deflection: Defending LLMs from Logit Manipulation do anything now

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.216533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.498317Z digest=sha256:c87c921c4d82be6d43b69fc33a1885b25c1f3c6ead4fedca42318a4d78fcf8e3

Observation 4fd0b2e3-1ba8-43d1-9d0e-8a75fcc6e530 · outbound

This paper cites A Cross-Language Investigation into Jailbreak Attacks in Large Language Models.

Strategic Deflection: Defending LLMs from Logit Manipulation A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.503676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.503676Z digest=sha256:731729ad1e52783a2403c0538febb41ba780f67f087dfef0a8127d4720c88955

Observation 6b12d06e-437a-4d7b-9924-6203a60c4777 · outbound

This paper cites Protecting your llms with information bottleneck.

Strategic Deflection: Defending LLMs from Logit Manipulation Protecting your llms with information bottleneck

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.192046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.509524Z digest=sha256:f48f6529555e1f354edb068942f91249632d3993b40481dd5bfef791e36d42a9

Observation 959e6f3d-fa51-4b4f-ab98-39129bcc81b5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Strategic Deflection: Defending LLMs from Logit Manipulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.514125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.514125Z digest=sha256:9d2f62a42cac1695968d411a758f98ae573f8d3d1aff549ab9601984f3325ea9

Observation 7f96471a-2d22-4819-acaf-8b43d1aebc20 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreaking black box large language models in twenty queries

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.169706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.520627Z digest=sha256:1138bb2036b168d89a6e0c25db8f766dffd9387b67264944e63a51ebb633b94d

Observation ec7747ca-aa95-40b9-829b-cb04fb8f2627 · outbound

This paper cites Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models.

Strategic Deflection: Defending LLMs from Logit Manipulation Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.145466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.526173Z digest=sha256:94bd835ba8cb5570feb7bd9259fcdf0f423f6b9760a6b78da255b50673c8ede0

Observation df81673c-0842-46ba-859d-37721d4b7f7d · outbound

This paper cites Catastrophic jailbreak of open-source LLMs via exploiting generation.

Strategic Deflection: Defending LLMs from Logit Manipulation Catastrophic jailbreak of open-source LLMs via exploiting generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.120803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.531779Z digest=sha256:c0473474d89f73fccae222479cd363e4bafe54c292ea40c34ed26639a1a7f8af

Observation 3c873539-fcdf-4f0e-ac3e-142987dd9777 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Strategic Deflection: Defending LLMs from Logit Manipulation Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.536146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.536146Z digest=sha256:0323327f93a3024a143662b8c756b42d91a37bbb4ddb5f7ba938b5826a29fead

Observation 15f793b9-b949-4dce-9e30-e3e82aa31e24 · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Strategic Deflection: Defending LLMs from Logit Manipulation Certifying LLM Safety against Adversarial Prompting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.541606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.541606Z digest=sha256:28fce00782edc3f7fe25ab4736f4610bd5bf6e2db21a1ea382ce7f5212636dd8

Observation fb9a6b82-8580-41da-95fa-7ed979b2a608 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Strategic Deflection: Defending LLMs from Logit Manipulation Detecting Language Model Attacks with Perplexity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.546775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.546775Z digest=sha256:72680e9db4b04c1932429da94a36beb7f1d51512277219acb0512cd2f4e03b32

Observation 7f9a2905-5fa5-454b-8b4f-e8392af0d7b5 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Strategic Deflection: Defending LLMs from Logit Manipulation Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.552474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.552474Z digest=sha256:8d9788740a0e2256418c4b58b25178d08408b6d002acd5c15ee1414d671efb17

Observation 5f17508b-2095-457f-8ce1-88baba66d487 · outbound

This paper cites Robust safety classifier against jailbreaking attacks: Adversarial prompt shield.

Strategic Deflection: Defending LLMs from Logit Manipulation Robust safety classifier against jailbreaking attacks: Adversarial prompt shield

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.097945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.558156Z digest=sha256:5c9ae60f854f01290183cfa8b841f8409f820055af76a19f2a9e9b3491ea1385

Observation cf1f2e26-6ed2-4ed9-ae91-db14b03b78ea · outbound

This paper cites Jailbreaking leading safety-aligned LLMs with simple adaptive attacks.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreaking leading safety-aligned LLMs with simple adaptive attacks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.078524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.562678Z digest=sha256:ae0b26a775503cc96abadd7a15e4429edc35468fbf63f68927e90e63a1069035

Observation 96d42cbc-56fc-4e49-9716-d763f85de590 · outbound

This paper cites Contrastive preference optimization: pushing the boundaries of llm performance in machine translation.

Strategic Deflection: Defending LLMs from Logit Manipulation Contrastive preference optimization: pushing the boundaries of llm performance in machine translation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.058799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.567565Z digest=sha256:542e6344d37e1e4bf86456959243a80128e99b4b8cb2d9588e1f6131583de930

Observation 5a10df76-0587-463d-ae29-e55144495ca9 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Strategic Deflection: Defending LLMs from Logit Manipulation Direct preference optimization: Your language model is secretly a reward model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.038739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.573509Z digest=sha256:cca32352f1dc6a33db82dd81afd0275d2aead27d89268ecdd2e346d3cb315cf5

Observation 30688163-6060-4951-9142-cd1519a81b63 · outbound

This paper cites GPT-4o System Card.

Strategic Deflection: Defending LLMs from Logit Manipulation GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.578742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.578742Z digest=sha256:dab6bed45cd9e4b564e2ac099a714c009af1d828eb887aad459d238244430662

Observation 0ccbce3f-2f9f-4247-abac-6a0be005e2f2 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Strategic Deflection: Defending LLMs from Logit Manipulation LoRA: Low-rank adaptation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.019654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.583833Z digest=sha256:9a3e53d776c0cbf2b2057396af5bbee1fdff5cb96054f42871d4940bff976e07

Observation bc351a75-c6fb-427e-acbc-49c9ac3f7208 · outbound

This paper cites Trl: Transformer reinforcement learning.

Strategic Deflection: Defending LLMs from Logit Manipulation Trl: Transformer reinforcement learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.589072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.589072Z digest=sha256:f5423600bf3449bf8d672fd1ad42ae9dc98e9468b213e3195985594ded4b0e59

Observation 873c020d-26d4-407f-8a25-49f083c58d78 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Strategic Deflection: Defending LLMs from Logit Manipulation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.597243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.597243Z digest=sha256:7494f76383044c191b9f05d2b3db4af0c5cbc9741a07daff5d40c4d61f1c1f68

Observation 0ae32380-0592-472e-9f90-fdee11c65965 · outbound

This paper cites tinybench- marks: evaluating llms with fewer examples.

Strategic Deflection: Defending LLMs from Logit Manipulation tinybench- marks: evaluating llms with fewer examples

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.987019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.603212Z digest=sha256:4c653aaecaad8187c85426259a9fe90079da39da269798add9aa1aa35859cd58

Observation 7681c4da-35cc-4c17-88e6-1bca9d12be53 · outbound

This paper cites Measuring massive multitask language understanding.

Strategic Deflection: Defending LLMs from Logit Manipulation Measuring massive multitask language understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.966908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.608545Z digest=sha256:03f9744bc4a5001284d799a7f7b889c89c2fdafbe6f3edf730a92d73e55f97af

Observation f5ba09d9-18eb-47c9-8748-12d938f62721 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019.

Strategic Deflection: Defending LLMs from Logit Manipulation Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.948768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.614731Z digest=sha256:d043ff9b5de1c6212425261a741316684ac3376f4a5cd446e69c755286183432

Observation e1449f35-7554-4725-8fee-7f70f47afc52 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Strategic Deflection: Defending LLMs from Logit Manipulation Truthfulqa: Measuring how models mimic human falsehoods

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.928307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T12:05:44.620968Z digest=sha256:56aeb56ac4516a53ea5e8ef392b2c35b615b1945b8ec18b7bb7f4c4ff4ff7303

Observation a8c10808-bfa8-4ae3-b582-4d13098932b0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Strategic Deflection: Defending LLMs from Logit Manipulation Training Verifiers to Solve Math Word Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.625951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.625951Z digest=sha256:2ceaf9ea7cd02b2d6b94cab610c94eefef876a9fe8ed521deeac78a67a872cec

Pith citing papers

Observation 6555fb52-5a81-4d7c-a31d-47c9eba147db · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Strategic Deflection: Defending LLMs from Logit Manipulation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-03T14:29:07.776313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-03T14:24:24.267269Z digest=sha256:fdcb097d5b22e9c81f57226daa0f60de279b6ebbcab0a54bc73bf52745f4fc0f