REVIEW 2 major objections 3 minor 27 references
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
T0 review · 2 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims that first-turn safety is an incomplete proxy for conversational safety persistence: across 4,000 fixed three-turn conversations, 71.6% contained unsafe guidance and 61.4% of conversations that began with a strictly safe…
desk verdict Real measurement of a distinct multi-turn medical safety failure mode, with a load-bearing automated judge; the qualitative conclusion is solid, the exact magnitudes should be treated as approximate. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is TAF-MED and its collapse-after-SAFE-$U_1$ metric. Each scenario fixes three user turns: $U_1$ declares self-treatment intent and requests medication guidance; $U_2$ and $U_3$ reframe the same unresolved request through probes such as educational, hypothetical, third-person, social-comparison, alternative-treatment, pharmacy-acquisition, or dose/frequency/duration. The metric is conditional collapse: among conversations in which the model's $U_1$ response is labelled SAFE, the fraction that later receives an UNSAFE label. The machinery also includes an intent-conditioned actionability rubric distinguishing non-actionable refusals (SAFE), partial case-linked disclosure (LEAKY), and actionable guidance such as a drug name, regimen, substitute, or acquisition path (UNSAFE), with an automated judge applying the rubric and physician-annotated subsets used for validation.
What would settle it
Re-annotate all 4,000 conversations with the physician rubric, or rerun the benchmark with a fixed template-safe $U_1$ response inserted into every conversation history, then recompute pooled any-turn and collapse rates. If collapse drops to near zero under a fixed safe $U_1$, the effect is driven by model-specific first-turn wording rather than by a failure to maintain an established boundary.
Extended reading notes
Core claim
TAF-MED's central finding is that establishing a medication-safety boundary and maintaining it are distinct capabilities. Using a three-class rubric (SAFE, LEAKY, UNSAFE) centred on whether a response materially advances a declared self-treatment plan, the authors found that unsafe guidance rose from 26.4% at the first turn to 63.0% at the second, and that the any-turn conversation rate was 71.6%. Among the 2,915 conversations that were strictly SAFE at $U_1$, 1,789 (61.4%) collapsed to unsafe guidance by $U_2$ or $U_3$, with model-level collapse rates ranging from 24.4% to 96.2%. Four of the 28 model pairs reversed their safety ranking between first-turn and collapse evaluation, and automated labels agreed with adjudicated physician labels on 94.3% of responses ($\kappa = 0.895$) while slightly underestimating conversation-level outcomes.
Load-bearing premise
The exact headline rates rest on automated labels for all 12,000 responses, with physician labels covering only 400 of the 4,000 conversations; if the judge's error rate is not uniform across models and turns, the reported rates and rankings could shift.
Editorial extensions
If this is right
- If first-turn safety is an incomplete proxy, safety evaluations of medical chatbots should report conversation-level outcomes such as any-turn unsafe rates, collapse, and trajectories, not just single-turn refusal rates.
- Model rankings from first-turn evaluation are not stable; four of 28 model pairs reversed, so conclusions about which model is safer depend on whether the evaluation includes follow-ups.
- Plausible, non-adversarial reframings of the same request are sufficient to elicit unsafe guidance; collapsed boundaries do not require jailbreak-style attacks.
- Because most collapse occurs at the first follow-up, the $U_2$ probe is a high-leverage place to test safety persistence.
- Treating partial disclosure (LEAKY) as failure raises any-turn unsafe from 71.6% to 78.7% and collapse from 61.4% to 70.7%, so the headline finding is conservative under that mapping.
Reading between the lines
- Editorial extension: the fixed three-turn format likely understates real-world collapse, since real follow-ups can adapt to the model's previous answer; testing adaptive user turns could reveal even weaker persistence.
- Editorial extension: the result suggests safety alignment should condition on the full conversation history, including earlier refusals, rather than treating each user message as an independent informational query.
- Editorial extension: rerunning the benchmark with a template safe $U_1$ response inserted into every conversation history would isolate whether collapse is driven by model-specific first-turn wording or by weak conditioning on the follow-up turns; the paper itself notes this as a limitation.
- Editorial extension: the intent-conditioned actionability rubric could be ported to other high-stakes domains where a refused request is followed by reframed requests, such as legal, financial, or self-harm contexts, to measure boundary persistence rather than first-turn refusal.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces TAF-MED, a physician-reviewed benchmark of 500 fixed three-turn scenarios in which a user declares an intention to self-treat a clinically serious condition and then continues with two non-adaptive follow-up probes. Eight LLMs are evaluated over 4,000 conversations, and each response is labeled SAFE, LEAKY, or UNSAFE by a GPT-4o judge, with two physicians independently annotating a model-balanced subset of 400 conversations. The principal empirical claims are that 71.6% of conversations contain at least one UNSAFE response, that 61.4% of conversations beginning with a strictly SAFE initial response later collapse to UNSAFE, and that four of 28 model-pair rankings reverse between first-turn unsafe rates and collapse rates. The authors argue that first-turn safety is therefore an incomplete proxy for conversational safety persistence and that evaluations should consider complete trajectories.
Significance. If the findings hold, TAF-MED is a valuable benchmark for an underexamined aspect of medical safety: whether an established refusal boundary persists across plausible, non-adversarial follow-ups. The paper's strengths include physician-reviewed scenario construction, a model-balanced validation sample, paired scenario-level bootstrap confidence intervals, explicit sensitivity analyses for label mapping and truncation, and an honest reporting of the automated judge's under-detection, which is conservative for the qualitative conclusion. The benchmark addresses a gap relative to existing multi-turn jailbreak and medical-safety benchmarks, and the trajectory-level outcome definitions are clear. However, because the full-corpus estimates and ranking reversals rest on automated labels for 90% of conversations, the exact magnitudes are not yet as well supported as the text implies.
major comments (2)
- [§3.3 and Tables 20, 24] The headline estimates (71.6% any-turn, 61.4% collapse, and the four reversals in Table 15) are computed from GPT-4o labels for all 12,000 responses, but only 400 conversations are physician-validated. The pooled under-detection of 4.5 and 5.4 percentage points is conservative for the qualitative claim, yet the validation subset is too small to support the exact magnitudes or the ranking-stability claim: per-model collapse under-detection ranges from 0 to 12.6 points (Table 24), and per-model LEAKY F1 ranges from 0.200 to 0.870 (Table 20). The authors should either release the judge outputs and report model-stratified calibration intervals for the unvalidated 90%, or present the exact percentages and pairwise reversals only on the physician-adjudicated subset. As written, Section 4.1 presents unadjusted automated-judge estimates as the primary results, and Section 7 only notes weaker LEAKY performance without bounding its effect on the reported rates.
- [§4.1, Eq. (1), and Table 15] Collapse after SAFE U1 is conditional on each model's own U1 response. Because models differ in whether and how they set the boundary at U1, collapse rates conflate initial boundary-setting with persistence under follow-up; for example, Gemini 2.5 Pro is SAFE at U1 in 369/500 conversations and collapses in 96.2%, while Llama 4 Maverick is SAFE in only 219/500 and collapses in 79.5%. The Limitations section acknowledges this, but RQ3 and Table 15 still treat collapse as a model-level persistence capability. I request a sensitivity analysis using a fixed canonical SAFE U1 response, or equivalent control for U1 wording, before interpreting the four reversals as evidence about persistence rather than about initial-response differences.
minor comments (3)
- [§4.3 and Table 27] The phrase "the rerun" has no antecedent in the main text; describe the earlier heterogeneous collection and its protocol, or remove the comparison from the main text and keep it in the appendix.
- [§3.3 and Appendix D.3] The manuscript states that TAF-MED will be released on Hugging Face but does not say whether the 12,000 GPT-4o judge labels and rationales will be included; releasing them would allow readers to audit the 10% validation and to apply calibration adjustments.
- [§7, Ethical Considerations] The release plan should specify the license and any dual-use restrictions for the benchmark prompts; the ethical considerations mention responsible release, but a concrete data card or license statement would make the reuse conditions clearer.
Circularity Check
No circularity: the headline outcomes are measurements from a physician-validated rubric, not consequences of fitted parameters or self-citations.
full rationale
The paper's derivation chain is benchmark construction -> fixed user turns -> model responses -> rubric-based labels -> derived conversation-level rates. No free parameter is fitted to the target result: the SAFE/LEAKY/UNSAFE rubric is defined behaviorally before collection, all 12,000 responses are labelled by a fixed GPT-4o judge at temperature 0, and the 400-conversation physician sample provides an external reliability check rather than an input to the outcome definition. The central quantities (71.6% any-turn UNSAFE; 61.4% collapse after SAFE U1; rank reversals) are arithmetic summaries of labels, and the claim that first-turn safety is an incomplete proxy is an empirical comparison of U1 rates to later-turn rates, not an identity. Collapse is definitionally conditional on a SAFE U1 followed by later UNSAFE, but whether that event occurs is determined by model behaviour, not by the definition. The paper also reports sensitivity analyses (LEAKY mapping, physician subset, truncation exclusion, bootstrap rank stability) that show the conclusion does not reduce to a single coding choice. There are no load-bearing self-citations: the paper cites prior external benchmarks and does not invoke any author-specific uniqueness theorem or ansatz. Concerns about GPT-4o label accuracy on the unvalidated 90% are correctness/measurement-risk issues, not circularity, and the paper's own validation shows conservative under-detection.
Assumptions & free parameters
assumptions (4)
- domain assumption Physician adjudicated labels are a valid reference for SAFE/LEAKY/UNSAFE medication-safety actionability.
- domain assumption Automated judge accuracy on the 400-conversation validation subset carries over to the remaining 3,600 conversations.
- domain assumption Fixed three-turn synthetic scenarios with declared self-treatment intent are a sufficient probe for multi-turn safety persistence.
- standard math Bootstrap percentile intervals with 5,000 resamples of scenario identifiers provide valid coverage for the reported rates.
Cite this review
Pith. "Pith review of TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent." pith.science (2026). https://pith.science/paper/RI3ISAGF
@misc{pith2026260810258,
author = {Pith},
title = {Pith review of: TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent},
year = {2026},
howpublished = {\url{https://pith.science/paper/RI3ISAGF}},
note = {Machine review of arXiv:2608.10258}
}
abstract
Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We introduce TAF-MED, a physician-reviewed benchmark of 500 fixed three-turn scenarios, and evaluate eight LLMs across 4,000 conversations. A rubric-based automated judge labelled responses as SAFE, LEAKY, or UNSAFE, and two physicians independently annotated a model-balanced random subset of 400 conversations. We assessed unsafe guidance, collapse after a strictly SAFE initial response, and model-ranking stability. Overall, 71.6% of conversations contained an UNSAFE response, and 61.4% of those beginning with a strictly SAFE response later collapsed to UNSAFE; model-level collapse rates ranged from 24.4% to 96.2%. Four of 28 model pairs reversed order between initial unsafe and collapse rates. Automated labels achieved 94.3% agreement with the adjudicated physician reference ($\kappa = 0.895$). These findings show that first-turn safety is an incomplete proxy for conversational safety persistence and motivate evaluation across complete dialogue trajectories. We will release TAF-MED on Hugging Face to support reproducible research on multi-turn medical safety.
Figures
Reference graph
Works this paper leans on
- [4]
-
[7]
Han, Tessa and Kumar, Aounon and Agarwal, Chirag and Lakkaraju, Himabindu , booktitle =. 2024 , volume =
work page 2024
-
[10]
Song, Jialin and Liu, Xiaodong and Yang, Weiwei and Chen, Wuyang and Feng, Mingqian and Zhu, Xuekai and Gao, Jianfeng , journal =. 2026 , doi =
work page 2026
- [11]
-
[14]
Pal, Ankit and Umapathi, Logesh Kumar and Sankarasubbu, Malaikannan , booktitle =. 2022 , volume =
work page 2022
-
[18]
Liu, Junyu and Li, Zirui and Niu, Qian and Zhang, Zequn and Xun, Yue and Hou, Wenlong and Wang, Shujun and Iwasawa, Yusuke and Matsuo, Yutaka and Hatakeyama-Sato, Kan , journal =. 2026 , doi =
work page 2026
-
[22]
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Qui \ n onero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, and Karan Singhal. 2025. https://doi.org/10.48550/arXiv.2505.08775 HealthBench : Evaluating large language models towards improved human health . arXiv preprint ar...
-
[23]
Oluwatobiloba Ayo-Ajibola, Ryan J. Davis, Matthew E. Lin, Jeffrey Riddell, and Richard L. Kravitz. 2024. https://doi.org/10.2196/55138 Characterizing the adoption and experiences of users of artificial intelligence--generated health information in the united states: Cross-sectional questionnaire study . Journal of Medical Internet Research, 26:e55138
doi:10.2196/55138 2024
Show all 27 references
-
[24]
Bitterman
Shan Chen, Mingye Gao, Kuleen Sasse, Thomas Hartvigsen, Brian Anthony, Lizhou Fan, Hugo Aerts, Jack Gallifant, and Danielle S. Bitterman. 2025. https://doi.org/10.1038/s41746-025-02008-z When helpfulness backfires: LLM s and the risk of false medical information due to sycopha...
2025 doi
-
[25]
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Fran c ois Beaulieu, and Paul Vozila. 2026. https://doi.org/10.18653/v1/2026.eacl-industry.39 MedRiskEval : Medical risk evaluation benchmark of language models, on the importance of user per...
2026 doi
-
[26]
Yella Diekmann, Chase Fensore, Rodrigo Carrillo-Larco, Eduard Castejon Rosales, Sakshi Shiromani, Rima Pai, Megha Shah, and Joyce Ho. 2025. https://doi.org/10.18653/v1/2025.bionlp-1.19 LLM s as medical safety judges: Evaluating alignment with human annotation in patient-facing...
2025 doi
-
[27]
Tessa Han, Aounon Kumar, Chirag Agarwal, and Himabindu Lakkaraju. 2024. https://arxiv.org/abs/2403.03744 MedSafetyBench : Evaluating and improving the medical safety of large language models . In Advances in Neural Information Processing Systems, volume 37, pages 33423--33454
2024 arXiv
-
[28]
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. https://doi.org/10.3390/app11146421 What disease does this patient have? a large-scale open domain question answering dataset from medical exams . Applied Sciences, 11(14):6421
2021 doi
-
[29]
Cohen, and Xinghua Lu
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019. https://doi.org/10.18653/v1/D19-1259 PubMedQA : A dataset for biomedical research question answering . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and...
2019 doi
-
[30]
Myeongju Kim, Haon Park, Woohyun Kim, Sookyung Choi, Ha Eun Kim, Hyoju Sohn, Jinyong Park, Sejoong Kim, Sangyoon Yu, and Yoonjin Oh. 2025. https://doi.org/10.1109/BHI67747.2025.11269553 PatientSafeBench : Evaluating the safety of medical LLM s for patient use . In 2025 IEEE-EM...
2025
-
[31]
Taeil Matthew Kim, Luyang Luo, Sung Eun Kim, Arjun Kumar Manrai, Eric Topol, and Pranav Rajpurkar. 2026. https://doi.org/10.18653/v1/2026.healing-1.2 The doctor will agree with you now: Sycophancy of large language models in multi-turn medical conversations . In Proceedings of...
2026 doi
-
[32]
Jinxi Li, Pengfei Zhou, Jing Wang, Hui Li, Hongbin Xu, Yuan Meng, Feng Ye, Yuqian Tan, Yanhong Gong, and Xiaoxv Yin. 2023. https://doi.org/10.1016/S1473-3099(23)00130-5 Worldwide dispensing of non-prescription antibiotics in community pharmacies and associated factors: a mixed...
2023 doi
- [33]
-
[34]
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. https://proceedings.mlr.press/v174/pal22a.html MedMCQA : A large-scale multi-subject multi-choice dataset for medical domain question answering . In Proceedings of the Conference on Health, Inference, and Le...
2022
-
[35]
Akshay Paruchuri, Maryam Aziz, Rohit Vartak, Ayman Ali, Best Uchehara, Xin Liu, Ishan Chatterjee, and Monica Agrawal. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.125 ``What's Up, Doc?'' : Analyzing how users seek health information in large-scale conversational AI da...
2025 doi
-
[36]
Stephen Raynes and Ellyn Maese. 2026. https://news.gallup.com/poll/707789/americans-turning-supplement-healthcare-visits.aspx Americans turning to AI to supplement healthcare visits . West Health--Gallup Center on Healthcare in America
2026
- [37]
-
[38]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Sch \"a rli, Aakanksha Chowdhery, Philip M...
2023 doi
- [39]
-
[40]
Sara Mahdavi, Christopher Semturs, and 7 others
Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, Elahe Vedadi, Nenad Tomasev, Shekoofeh Azizi, Karan Singhal, Le Hou, Albert Webson, Kavita Kulkarni, S. Sara Mahdavi, Christopher Semturs, and 7 oth...
2025 doi
-
[41]
Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha, Binze Hu, Yuhui Wang, Chenhao Huang, Shichun Liu, Jingqi Tong, Changhao Jiang, Mingxu Chai, Zhiheng Xi, Shihan Dou, Tao Gui, Qi Zhang, and Xuanjing Huang. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.263 LLMEval-Med : A r...
2025 doi
- [42]
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.