Pith. sign in

A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos

6 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

6 Pith papers citing it
1 external citations · external index

years

2026 4 2025 2

verdicts

UNVERDICTED 6

representative citing papers

ToxiREX: A Dataset on Toxic REasoning in ConteXt

cs.CL · 2026-06-26 · unverdicted · novelty 6.0

ToxiREX is a new dataset of 128k Reddit comments in six languages with hierarchical annotations for implicit toxicity in conversational context based on an existing reasoning schema.

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling

cs.CR · 2026-06-18 · unverdicted · novelty 6.0

SafeSpec integrates a latent safety head into speculative LLM decoding with rollback and reflective multi-sampling, cutting attack success rates 15% on Qwen3-32B while retaining 2.06x speedup on normal workloads.

citing papers explorer

Showing 6 of 6 citing papers.