Pith. sign in

Paper Citation Record · LEDGER

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2507.04673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04673 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:46:27.712204Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbe1ae7c-e1cd-4754-b783-12ba81f019ae · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.135856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.135856Z digest=sha256:92df0315b05e664db87060c0171c8057e119afc790584e3025dc7bb9cd3c52ec

Observation 68b78950-31ea-4ae7-a11a-c446900e07a2 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.203611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.203611Z digest=sha256:5a834f2c3ec3975446cec324442e97d1141616b68adf109b9339d170f594586c

Observation 02c6bfce-8387-4b88-961f-51cdbc6e695a · outbound

This paper cites Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.332759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.332759Z digest=sha256:3b6ade1f6e557068e5768ab09468d37e5239e71a51c7b70918365dad6c04c1bd

Observation 42e7b7f6-a4ff-454f-98b2-b7b0f6a236fd · outbound

This paper cites & Lowe, R.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message & Lowe, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:30.312037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:25.420968Z digest=sha256:066542890edc69f7414be07feb2b7cb7315f0a1ada73fb6b16afb2dbbca5961d

Observation 39ef1738-b7ea-4156-ba56-a47d600122a4 · outbound

This paper cites Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.503493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.503493Z digest=sha256:f25f4503beff510bdb41f2bc0a6382a616bded301060375dde4af177e18352dd

Observation a68bd5e6-34fd-4093-8a56-43f078dbbba9 · outbound

This paper cites Red Teaming Language Models with Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Red Teaming Language Models with Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.612647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.612647Z digest=sha256:86f76e3a5f723c27f6ebf2be2a38641d6b6bf039dd3813b1fed7d288387e0081

Observation 6433a5ab-e4ec-4c88-93f2-acf3754ec031 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.761176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.761176Z digest=sha256:d9ef81fa8b347da8c959070f15b9891689d3aff9e26ac7f069b3dd0aa1fd39d5

Observation 3d0799c7-1f3f-4ff7-b50c-591d077226c6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.870469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.870469Z digest=sha256:484dbf4a0c7a82d12191f652b249e13efa9ae54e2988cb8db92096e3f381cf6f

Observation 7fd45238-5c2f-435e-842e-0f93b792de74 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:46:30.169777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:25.984097Z digest=sha256:256ac7b287de62b50d8dd2dd182337b398bce51c32edea80d866bde529cc21a7

Observation c72c4f5c-94a3-4041-b26b-747ee4884a48 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:46:29.946114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:26.136525Z digest=sha256:5eb04bd08574ed1981efbaf36c2fc1a4b170bd0571a6ccd9658067594b754d57

Observation 129cc2a2-71a6-4f6e-9f65-da89169c3590 · outbound

This paper cites Multimodal Pragmatic Jailbreak on Text-to-image Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Multimodal Pragmatic Jailbreak on Text-to-image Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.221727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.221727Z digest=sha256:14b1125858129ea9ed529f71745b885aa09f66310f82abef10c2712e62b71e2a

Observation 86029344-f94f-4cc8-958d-392e967e577e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.337360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.337360Z digest=sha256:4a7e76617ccec75fb958537d28329f387f75517eae6eab6ed8a2cd71d277c28b

Observation 5b2c57e0-11c6-4d76-903a-19dec5ccff73 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:29.747993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:26.401264Z digest=sha256:06f443c24e399f87a338865df2bef2a2ecdf12227911dc90df3e49d00dea255b

Observation 0d2a6696-a20a-4316-9611-5342157ccb23 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Jailbroken: How Does LLM Safety Training Fail?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.491295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.491295Z digest=sha256:0a5aa4b5cc075b97b301587d2bdf762716e98c9de2e0de28996cae7d6ed1583c

Observation e6418170-4cf4-49d2-b7f2-4d1e2b9b094d · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.685840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.685840Z digest=sha256:d50692bf1a5815e07d2c1623e4cab002490f93eb7bc76719874c583f95862e4a

Observation 2b6982bd-78ac-4b73-aa55-315d166a41e4 · outbound

This paper cites Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:46:28.715500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:26.766578Z digest=sha256:5d58ae0a2bfb336f20030d0c20d474f54a06108da6aaa63c730b83f60ef6369a

Observation e4dedf08-ce85-4eb0-92da-562fef96fbe0 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:46:29.028272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:26.858023Z digest=sha256:f79552178ede3faa11ddbe1fbc2f147d65c25ccd45ddc6ef97dc310f9aaab471

Observation 5128f266-0187-4ed2-974e-3b94f5a72697 · outbound

This paper cites GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.940175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.940175Z digest=sha256:ebb495809cfdcfd4ddadc80b8c1d8fc8a7437b296192ebf0a7c01f82a8d6977d

Observation 0a5cbd56-88a6-4c88-8df7-b6e574ac8c03 · outbound

This paper cites Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.999796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.999796Z digest=sha256:2207db047d0e49c823857edf4d4bb7038898f994a830bba33c6a468b0d4f18c1

Observation a6b983d6-49ab-4cdd-967c-1d567a4bd97d · outbound

This paper cites HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.054948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.054948Z digest=sha256:49f3f9fdfee02ebde037a714ff782667953bc0dfd0daa76f9c4ab3f1b5dd75b7

Observation 3e668489-c586-4ce6-84b2-c3e279662485 · outbound

This paper cites SneakyPrompt: Jailbreaking Text-to-image Generative Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message SneakyPrompt: Jailbreaking Text-to-image Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.138588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.138588Z digest=sha256:609e5f70a76a6f6973f532bd4735fecc8c4c7a72310db81670bf339f4b76ae0d

Observation 634b7aab-52c0-4de4-91af-93bce1a07b10 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:46:28.438341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:27.219717Z digest=sha256:27dd3de693df2725095a7ad5ed19aa9b2b2b785bd7dc8ac922253a80b2290cb5

Observation b2067cff-6d94-47f2-8e03-bcdfb55f1ec1 · outbound

This paper cites Gradient-based Jailbreak Images for Multimodal Fusion Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Gradient-based Jailbreak Images for Multimodal Fusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.317130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.317130Z digest=sha256:b1361ce096632e633be513a7cae3e36cd0994f2602af3e32b1cca47b617f6194

Observation 6a6c5220-9606-4a96-be45-78998596fd50 · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.368912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.368912Z digest=sha256:9e45dac7a24efeb6376ff5458571026f28d5d54bd94eaa6031d5a1a7af6f5a3d

Observation f4e24e59-c187-4a25-b386-c1458f598470 · outbound

This paper cites Red-Teaming LLM Multi-Agent Systems via Communication Attacks.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Red-Teaming LLM Multi-Agent Systems via Communication Attacks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.451948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.451948Z digest=sha256:ef2eded7e9739705b25dd76126ed274f0f58ef8fc8efa204e94a348b0e334598

Observation ed12d4c6-a0e6-4378-b560-30a9911f5ae0 · outbound

This paper cites GPT-4 Technical Report.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.536570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.536570Z digest=sha256:9795e3ae08d60da043438fd2349bef4fce235da2de4329728a0c39526fb2ee1e

Observation 3e0cb174-ef09-44f7-a443-6fab71e75112 · outbound

This paper cites Consideration Set Sampling to Analyze Undecided Respondents.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Consideration Set Sampling to Analyze Undecided Respondents

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:46:27.921011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:27.623901Z digest=sha256:a2b5bf895d44e2f38fb7197fce8c724166737edcda80425ffa511d7ae52995e6

Observation 0cf1581b-227e-41b4-9945-b6c7ee04bc4a · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:46:29.529563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:46:27.712204Z digest=sha256:ef2be66ffcd676364a68642f54372e1cabb41d9c7c21c346766333f1b15dfdf0

Pith citing papers

No inbound Pith citation observations are available.