Pith. sign in

Paper Citation Record · LEDGER

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.24232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24232 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:34:37.431217Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e35ecabd-0d1e-4c7b-9611-f606becda309 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.447358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.447358Z digest=sha256:444d309c3f99fb3ae66782364144414f8fbfb4604398e59a78c6fc89f510c25c

Observation c1d163ed-ece8-4eee-8c4a-5f1de2657dcd · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.513188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.513188Z digest=sha256:2cc7d58ed924a03475231b91b772eff37af792dbcaff12a46703223f07418071

Observation 6b6dafde-1e18-4ff5-a51d-c0f0326916f1 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.562775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.562775Z digest=sha256:b565a1b1f872c7d6ca2cd988ff0b537d72275e6cf71a8af336e3bb6f5013a7eb

Observation ab5dc5e0-dbe8-408b-ad5b-74309992baaa · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.729370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.729370Z digest=sha256:02d7fff1b1ca88dbcb1dcb3bac3c02a95ff4e078af2c6cfb49d919cc22e5ea9e

Observation 7d16e534-ae4a-4e03-a9e5-8de253fa91a5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:32.878990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:32.878990Z digest=sha256:e73b0f8309cdec9b25e1bc61d16c714d83605bc1bae093edda25431531feb379

Observation 228d84c1-4589-483e-83c3-3e103d475d26 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.055462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.055462Z digest=sha256:11e60294ad42721ef8f861c469b951b8b8661beb33f5bdafd60965525e573733

Observation 4c6fcf8d-c3d7-485e-b0e3-0974b692ccc8 · outbound

This paper cites Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.183946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.183946Z digest=sha256:0722889bc35bedb9dfcd1a6badda22fe7bccf23d2664eec51ac59693ddc6016a

Observation 7ca8e42f-9646-4f40-b65c-8ee8cde55c58 · outbound

This paper cites Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.321878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.321878Z digest=sha256:77f2dea2b1499f8716b15a950c82ea464e0eb542060e15327ac1759c760b1d4a

Observation bf596b9f-910d-46d3-b4a9-aa9a91c2ae6b · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models The Internal State of an LLM Knows When It's Lying

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.468297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.468297Z digest=sha256:504f03621a27c891f58891d754ccc2bbef73b39dc4a4de47b39b41860e528292

Observation 8e9f5913-57a1-4726-9a22-4937b27042c4 · outbound

This paper cites Llm factoscope: Uncovering llms’ factual discernment through measuring inner states.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Llm factoscope: Uncovering llms’ factual discernment through measuring inner states

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.576187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:34:33.585739Z digest=sha256:5f7c1fe89f12e8cfcc9fb328d7aa9cae2828c84fe89dc54e7e772d2add61d3cc

Observation 0bf08dfa-e27a-450f-a48a-f511d3f98adb · outbound

This paper cites Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.685570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.685570Z digest=sha256:0e4c1c28def9571d153ea0432e0cfd30d5a186aaa3d6e6fc8c36291dcb84dfbd

Observation 99ffbb33-564d-47a1-b37d-2b06830a0d9e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:33.845038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:33.845038Z digest=sha256:0cdc09ca63e239eb58cb512c40cdf31be21be6c7e80888bc72a31b679d76be42

Observation 25e2e540-3aba-44b3-850c-3469a159a4b6 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.003149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.003149Z digest=sha256:93dcdad15da6110d002ba0555bad3255663608483b0b8d7af2b4a735a7c6fb8d

Observation a0cad85c-55d5-4f37-9815-ac974c527827 · outbound

This paper cites Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.arXiv preprint arXiv:2402.03299, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.arXiv preprint arXiv:2402.03299, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.167588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.167588Z digest=sha256:07072627cfd4e544594f5844c0279ab1aadb49390758ae1956f101768db9c546

Observation 79c19ef7-639f-4dca-99db-c8d3a743b0e5 · outbound

This paper cites Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.241608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.241608Z digest=sha256:5ed3809fbe6f64a512782998bde9806fafbfc239cebeb335246257bad9d7a2d5

Observation 29787c52-bcc5-40eb-8475-36a31bd4a65a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.347192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.347192Z digest=sha256:ca439c5ad6756937a5bf30afd439e90e366aac37f7bef66daf47ee707fdca95e

Observation 06467daa-7a79-41c9-a2a9-30d088d57aaa · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.434538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.434538Z digest=sha256:2bc38fceccd9cc1f185fc2e5e780150cf6d668f7e7961e62ce8531d08c5bc132

Observation b71f90d3-ed83-411a-953f-8f89b9ec73c1 · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.521629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.521629Z digest=sha256:0c222114ef3eeb0240acdf553c54ef982540e862502c5b72086fe7c3eb4984fc

Observation 75a00167-735f-4d5b-84bf-38140ee83424 · outbound

This paper cites In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.621558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.621558Z digest=sha256:5fb7a0c33195da129ecc78d2bb486e4cde1e2fff3c5f91bc9cc65d176181945d

Observation 47e5f63f-0824-4a7c-a89a-75e1252df0cb · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbroken: How Does LLM Safety Training Fail?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.714953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.714953Z digest=sha256:929501fdd37b16d499cad7d1334756022beb86d2a73eaaae962b7dbec7198f30

Observation cc291607-87bc-4601-8fff-56431417bdf9 · outbound

This paper cites Query-Based Adversarial Prompt Generation.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Query-Based Adversarial Prompt Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.809699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.809699Z digest=sha256:c89f15481e0607a0b064d8f12b0413f4e96bda1c7584a24607f888a5da37f6d4

Observation b24cd802-df9b-4580-8ccb-2f148221cb3a · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:34.930392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:34.930392Z digest=sha256:57e7d756f2011dacad628e3a801ae7c18153239625a2c0c1a01421e4f4f87634

Observation bdcec99b-775e-4fe4-8083-8ac83e1925fc · outbound

This paper cites DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.024627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.024627Z digest=sha256:9751731a5982490cbaf997a3d70b0a919ba708155669c77eb2d0ecbcb849b71f

Observation 25ce4ff6-5ffe-4b86-864f-aa98df872c3b · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.156845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.156845Z digest=sha256:39efe5e9c2c8ed79dc03a9cb746b0b4d57be88e420892cb4d57ddee5dd5b8bf2

Observation c52ceef5-c819-4a08-b794-1da06089cc8c · outbound

This paper cites Jailbreaking proprietary large language models using word substitution cipher.arXiv preprint arXiv:2402.10601, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreaking proprietary large language models using word substitution cipher.arXiv preprint arXiv:2402.10601, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.268065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.268065Z digest=sha256:40246835f3d21179f7d3f9bbb02f7c29cd3d7ca3de2c683d251a4d297cbae391

Observation 308a610b-bc2e-431e-b044-508bc991be46 · outbound

This paper cites Are aligned neural networks adversarially aligned?Advances in Neural Information Processing Systems, 36, 2024.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Are aligned neural networks adversarially aligned?Advances in Neural Information Processing Systems, 36, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.415641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.415641Z digest=sha256:cc6acd7636bbaef627030e607ea86ac2309b263dda0b78d7f9acfbc105309951

Observation a9546250-14fd-4632-86b1-3034e8ad0d68 · outbound

This paper cites On Evaluating Adversarial Robustness of Large Vision-Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models On Evaluating Adversarial Robustness of Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.513830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.513830Z digest=sha256:4be5b6428a24ed09fbf743e98d8704b184d5faab1ed92e1c71cb2bf67077f8b4

Observation cfa67e39-094b-494c-8967-ee361b543bc3 · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.624649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.624649Z digest=sha256:cb73652d96d4c8d63c312048f60515b57d429b1ed64e467271929925a4d8975a

Observation 75ed4a4d-394e-4782-845e-8d1011d2d211 · outbound

This paper cites On the adversarial robustness of multi-modal founda- tion models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models On the adversarial robustness of multi-modal founda- tion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.409736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:34:35.729482Z digest=sha256:eb12e5217276b39fc861c874025bd526135fd10ddff817ab98a54a22f74cde6e

Observation 4fbd0ab7-905c-4e47-86f0-224db5bb024e · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.811622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.811622Z digest=sha256:5724b85d0abdc946bdd0d2fedca90b5b66211dbe50061d04ed572fd600330430

Observation 8d5d5faa-34ed-4891-aed0-77149d82e322 · outbound

This paper cites InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.908111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.908111Z digest=sha256:a5924e1e26f7e40bb7c6a8a2c73893ac093a5781de69ca72b0c66cd6ac266bd9

Observation 88b0f8bd-44d8-4046-9f84-f6627fb9abde · outbound

This paper cites Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:35.999122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:35.999122Z digest=sha256:826c57e52b8e56442b4090fd62061c7a8ce6a89f4e8502e2b4794c5950bd8289

Observation c1cd4eaf-7b93-47b5-9414-62d703309478 · outbound

This paper cites Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Adaptive Activation Steering: A Tuning-Free LLM Truthfulness Improvement Method for Diverse Hallucinations Categories

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.079992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.079992Z digest=sha256:ad08a27e7f71bfc4334a4393a1ece6af852a99c22484f563c23b476c96827897

Observation 3f94b444-3932-4159-a2b9-fb24652fb22b · outbound

This paper cites INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.185423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.185423Z digest=sha256:d5ccefaba1ac5267db2241710f22222531ebc91c6308ed17e12d80a0f4d9dea4

Observation 3baef64a-4f67-42b3-a59d-ec7f2c243386 · outbound

This paper cites Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.282420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.282420Z digest=sha256:29e7a3302d8ea3ac89f445c8bdca75568cdfc651ccc79de9f0f4361957f0e697

Observation 0fde98d1-3178-4553-82be-f66021aa2fbc · outbound

This paper cites LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.391398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.391398Z digest=sha256:8ffe65a5cb4598f462346f29185d04ce04dbdc4a66f08a052d6a15dcfeb9e03c

Observation 91ee52b9-20e9-4240-a91b-7ce7c1e36166 · outbound

This paper cites Are sixteen heads really better than one? Advances in neural information processing systems, 32, 2019.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Are sixteen heads really better than one? Advances in neural information processing systems, 32, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.498025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.498025Z digest=sha256:10566e1dacf30c661f1e04ae3b7d72cc8131803cabd6ad326ddb7c1ff5575026

Observation 38e9fd75-f90e-4821-880c-ac10222c7841 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models What Does BERT Look At? An Analysis of BERT's Attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.609925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.609925Z digest=sha256:10a08d74ab99f419054f97ba783c71ef66b35b225cd9c91f9264559aed8727a4

Observation 0b2ecfab-c7d8-4841-aad9-c93d781795bd · outbound

This paper cites Improved baselines with visual instruction tuning.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Improved baselines with visual instruction tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.681348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.681348Z digest=sha256:996b28477e9138f4cc8a06b021aeb452aee79dcfa3a43cc9bd2f8a0eb8d4cf9b

Observation 85533a89-f4f5-42e4-a519-bbd9dd3cbd31 · outbound

This paper cites AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.773583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.773583Z digest=sha256:ab0d134230b19762533833decd4b460e473e2108cfd8ad9e12b8e1a842e2d736

Observation 94a69ac0-6493-44e9-9945-eece68ad9f8d · outbound

This paper cites AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.872796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.872796Z digest=sha256:b7f37992f4814c6cd70b2164423cd2255439879a7a95773c7dbd9e1bd8758845

Observation 9337c7f5-1e48-4495-81cd-ac7b7dc8f23f · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:36.950087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:36.950087Z digest=sha256:021303fa6aad0865560516f211f1c940322a68d44a2e772bbba423f7b71222b1

Observation 30cdca5a-afda-4c9b-a8a7-42cc3f014507 · outbound

This paper cites Safebench: A benchmarking platform for safety evaluation of autonomous vehicles.Advances in Neural Information Processing Systems, 35:25667–25682, 2022.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Safebench: A benchmarking platform for safety evaluation of autonomous vehicles.Advances in Neural Information Processing Systems, 35:25667–25682, 2022

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.232440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:34:37.017349Z digest=sha256:f47dd250aa55aaf393ca6526e913fbecc2d4999bb533813b9a14556e939dcd05

Observation 197b9f6a-0da3-4cf5-a240-b7adcd798e32 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:37.104583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:37.104583Z digest=sha256:989d9d2fecd6c10fc61ed36ee5dc4064641df79d08aeebeedcabab6a4cbb77b7

Observation 6db2aa4b-4670-4fc3-9425-f3fce365f5cb · outbound

This paper cites Jailbreak Large Vision-Language Models Through Multi-Modal Linkage.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:37.173951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:34:37.173951Z digest=sha256:4eae10ee0e45c48c389080ea7e99758e6512697b83f8f1dea07116d9be6adc13

Observation 02a122a8-9246-4bef-8e66-453709d59ad7 · outbound

This paper cites If detected, immediately stop processing the instruction.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models If detected, immediately stop processing the instruction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:39.057567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:34:37.247713Z digest=sha256:224267b527bf04b3f7b029d21a4ad15c2b93a88819a235520a4629e5e4329e14

Observation d3afef16-ddb8-443a-81ab-4a4ebfc164fe · outbound

This paper cites an unresolved cited work.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:34:38.841755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:34:37.335057Z digest=sha256:5e990a8530cf7b911a436a9cebc113f657084f954335d13b99b438cc3a203891

Observation 53ba6730-2b5e-4422-8c0f-47914173ae80 · outbound

This paper cites I am sorry.

From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models I am sorry

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:34:38.663788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:34:37.431217Z digest=sha256:11a9614cca180ead62ce3c69af6809afc0f8fc877295d745404052a94dc4aeb5

Pith citing papers

No inbound Pith citation observations are available.