REVIEW 3 major objections 6 minor 24 references
Repeated conversational pressure with no new visual evidence destabilizes vision-language model answers, making multi-turn stability a distinct reliability axis from single-turn accuracy.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 15:02 UTC pith:MFMAXDZX
load-bearing objection A useful new multi-turn pressure-test protocol for VLMs with a real missing-control problem, but the core idea and trajectory metrics deserve referee time. the 3 major comments →
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that when the visual input is held fixed and only the follow-up prompt changes, repeated conversational pressure acts as a destabilizer with bounded upside. Across 720 multi-turn runs, final accuracy changes by only about one percentage point, but underneath that stability are asymmetric trajectories: more runs regress from correct to wrong than recover from wrong to correct, and 26% of runs flip at least once, with most first flips on the very first follow-up turn. The effect is strongly model-specific: one model is the most brittle and oscillatory, another is accurate but stubborn and becomes more confident when wrong under contradiction, and a third is stable but toke
What carries the argument
The Just Keep Prompting (JKP) framework: each run starts with one video-question pair, records the model's answer, confidence, and rationale at turn 0, then applies up to ten pressure-bearing follow-ups in one of three strategies—adversarial negation ('No, I disagree, that is not correct'), pure Socratic interrogation ('Are you sure?'), and context-aware Socratic summarization (feeding the model's own summarized rationale back before asking again). The key measurements are the trajectory itself: accuracy change, wrong-to-correct and correct-to-wrong transitions, number of flips, turn of first flip, confidence trajectories, and token use. The mechanism is that these follow-ups carry no new vi
Load-bearing premise
The claim that conversational pressure causes the answer flips rests on the assumption that the model would not flip equally often when simply asked again in a neutral, pressure-free way; because the protocol has no such neutral control and uses nondeterministic sampling, some reported flips could reflect stochastic re-sampling rather than the pressure itself.
What would settle it
Run the same protocol but replace each pressure prompt with a neutral 'Please answer again' (no disagreement or doubt), holding video input and temperature fixed. If the neutral condition produces a comparable or larger flip rate—for example, approaching the 71.2% pure-Socratic flip rate—the central claim that pressure, not sampling, drives instability would be falsified; a pressure-driven account predicts flip rates near zero under neutral re-asking.
If this is right
- Single-turn accuracy is insufficient for evaluating VLMs in interactive deployment; multi-turn stability must be reported separately.
- Because regressions slightly outnumber recoveries and flips often happen on the first follow-up, models should not treat 'asking again' as a free reasoning step.
- Initially correct answers are fragile: pressure can turn an easy interaction question into a wrong final answer, so safeguards are needed against revision without evidence.
- Calibration metrics should be computed separately for correct and incorrect states under pressure, since some models become more confident when wrong under contradiction.
- Physical-affordance questions are both the hardest and the most oscillatory, suggesting robustness work should target physical grounding specifically.
Where Pith is reading between the lines
- A neutral re-ask control—prompting 'Please answer again' with no disagreement—would separate the effect of conversational pressure from ordinary sampling variability at nonzero temperature.
- The bounded-upside pattern suggests a design target for multimodal assistants: asymmetric revision behavior, changing answers when evidence warrants and holding ground when only the tone of the user changes.
- The early-flip pattern hints that these models interpret a first disagreement as social evidence; a related test would measure whether adding explicit reassurance ('I'm not saying you're wrong') removes the flip.
- Because one evaluated model sees fewer frames than the others, a fully fair model-level comparison of brittleness would control frame count and visual sampling; the paper's contrasts are informative but not purely conversational.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces Just Keep Prompting (JKP), a multi-turn evaluation protocol for vision-language models in video question answering. After an initial STAR-benchmark answer, each run applies up to 10 follow-up turns under one of three pressure strategies: adversarial negation, pure Socratic interrogation, and context-aware Socratic summarization. The authors evaluate GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B on a balanced 80-question STAR subset, for 720 total runs. The headline result is that aggregate accuracy changes little (68.8% to 67.8%) while trajectory-level instability is substantial: many runs flip answers, correct answers regress, wrong answers sometimes recover, and models differ sharply in their 'pressure-response profiles' — e.g., GPT-4o is reported as brittle/oscillatory, Qwen as stable but overconfident when wrong, and Gemini as stable but token-expensive. The paper argues that repeated conversational pressure, not new visual evidence, destabilizes VLM answers.
Significance. If the causal interpretation holds, JKP is a useful, low-cost probe of a largely underevaluated axis of VLM behavior: multi-turn epistemic stability under user pressure. The trajectory metrics (flip counts, first-flip timing, confidence trajectories) are transparent, and the evaluation against a fixed external benchmark (STAR) is a direct empirical measurement with no fitted parameters, so circularity is not a concern. The finding that aggregate accuracy hides large turn-to-turn instability is an important methodological point. However, the central causal claim is currently under-supported: the protocol lacks a neutral repetition control, statistical uncertainties are not reported, and a frame-cap mismatch confounds model-level comparisons. The framework is valuable, but the present experiments do not yet establish that the observed flips are caused by conversational pressure rather than by sampling stochasticity or by differences in visual input.
major comments (3)
- [§3.2, Eq. (8), §5.4] Missing neutral-repetition control. The protocol contains only pressure-bearing follow-ups ('No, I disagree', 'Are you sure?', self-summary), and decoding is done at temperature 0.2, which is not deterministic. Without a control condition that re-asks the same question without disagreement or doubt (e.g., 'Please answer again in the same format'), the headline flip rates — such as GPT-4o flipping in 57/80 Socratic runs — and the C→W / W→C classification cannot be attributed to conversational pressure. A neutral baseline would quantify the per-turn flip probability due to sampling noise alone, which is exactly what the pressure interpretation must beat. This is load-bearing for the abstract's causal claim and for the 'pressure-response profile' framing.
- [§5.1–5.4, Tables 2–4] No uncertainty quantification. Each model-strategy cell is based on 80 runs; counts such as 10 vs 12 improvements/regressions and confidence gaps such as 94.68 vs 93.08 are presented as stable behavioral signatures. Report bootstrap confidence intervals, exact binomial tests, or at least per-condition standard errors. Without this, several model-specific and category-specific claims — especially the Feasibility retention ratios in Table 4 and the Qwen wrong-state confidence claim — may be within sampling noise.
- [§4.1, Tables 3 and 5] Frame-cap confound in model-level comparisons. GPT-4o receives at most 30 frames per video, while Qwen and Gemini receive up to 80 frames. The paper's model rankings in accuracy, flip rate, and token burden (e.g., Table 3 and Table 5) are therefore not cleanly attributable to conversational robustness; they may partly reflect differences in available visual evidence. Please equalize frame budgets across models or explicitly present the frame-cap difference as a confound and soften direct model-to-model conclusions.
minor comments (6)
- [Table 1 vs Table 3] For GPT-4o, Table 3 reports T10 accuracy 60.4%, but the average of the three strategy-level T10 values in Table 1 (61.5, 58.8, 62.5) is 60.9%. Please reconcile or explain the discrepancy.
- [§3.3, §5.4] The term 'flip rate' is used both for the fraction of runs with at least one flip and for flips per run. Please define each quantity explicitly and use distinct labels to avoid confusion (e.g., 'runs with ≥1 flip' vs. 'mean flips per run').
- [§4.1] Please clarify whether the temperature of 0.2 applies identically to all three models and backends, and how 'backend-specific video processing' may affect token counts and visual sampling. This is relevant to both the stochasticity concern and the token-burden comparison.
- [§5.5, Table 4] The Feasibility category has only 20 questions per model-strategy; category-level ratios in Table 4 are based on small samples after aggregating across strategies. Please report confidence intervals or at least the raw counts underlying each retention ratio.
- [§5.7] The token-burden comparison is heavily confounded by the different frame caps and by model-specific verbosity. The paper acknowledges this is not a cost analysis, but the 'strong token-efficiency result' should be softened and framed only as a descriptive observation.
- [§4.1] The selection procedure for the 'balanced 80-question STAR subset' is not described. Please report the sampling method or seed so that the subset can be reproduced and the category-level results interpreted appropriately.
Circularity Check
No circularity: JKP reports raw trajectory statistics over the external STAR benchmark; no fitted parameter or self-citation chain is used as a load-bearing derivation.
full rationale
The paper's central quantities — accuracy, flips, turn-of-flip, confidence, token usage — are defined directly from observed model trajectories (Eqs. 1–11) and computed against the external STAR benchmark. There is no step in which an input is defined in terms of the predicted output, no fitted parameter is renamed as a prediction, and no uniqueness claim is imported from the authors' own prior work. The related-work citations (e.g., sycophancy, gaslighting, multi-turn degradation) provide context but are not load-bearing in the derivation of any reported result. The main interpretive weakness — that the absence of a neutral repetition control makes it hard to attribute flips to conversational pressure rather than sampling stochasticity at temperature 0.2 — is a confound and a threat to causal validity, not a circularity: the flip statistics do not reduce by construction to the prompting strategy definitions. The frame-cap difference between models is similarly a fairness/control concern, not definitional circularity. The reported measurements stand or fall on their own experimental terms, so the honest circularity finding is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- STAR 80-question subset selection
- Video frame cap per model =
GPT-4o: 30; Qwen3-VL-30B & Gemini 2.5 Pro: 80
- Decoding temperature =
0.2
- Number of follow-up turns =
10
axioms (4)
- domain assumption Answer changes reflect conversational pressure rather than model sampling stochasticity (no neutral-repeat control).
- domain assumption The 80-question STAR subset is a valid and representative probe of video reasoning.
- domain assumption Auxiliary summarizer (Qwen3-8B) summaries do not inject answer-relevant bias.
- domain assumption The three models were compared on equal footing for model-level claims.
read the original abstract
Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or contradict a model's answer. JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to reassess certainty), and Context-Aware Socratic Summarization (reflecting the model's prior rationale back before asking for reconsideration). We evaluate GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B on a subset of the STAR benchmark across 720 multi-turn runs. Aggregate accuracy changes modestly from Turn 0 to Turn 10, but trajectory-level analysis reveals substantial instability: correct answers regress, wrong answers recover, and many runs exhibit repeated answer flipping. Repeated prompting has bounded upside and often acts as a destabilizer rather than a reasoning aid. The effect is strongly model-dependent: Qwen3-VL-30B achieves the highest final accuracy but becomes confidently wrong under direct contradiction; Gemini 2.5 Pro is comparatively stable but token-expensive; GPT-4o is the most brittle and oscillatory. These findings reveal that multi-turn VLM evaluation captures not just additional reasoning but pressure-response profiles: how models trade off visual grounding, calibration, and conversational compliance under repeated challenge.
Figures
Reference graph
Works this paper leans on
-
[1]
Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736, 2022
Jean-Baptiste Alaymac et al. Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736, 2022
2022
-
[2]
Stephen Casper et al. Open problems and fundamental limitations of reinforcement learning from human feedback.arXiv preprint arXiv:2307.15217, 2023
Pith/arXiv arXiv 2023
-
[3]
Yida Chen et al. Purifying interactions: Socratic prompting for robust llm reasoning.arXiv preprint arXiv:2305.14231, 2023
Pith/arXiv arXiv 2023
-
[4]
Making vision-language models robust to modality bias.Advances in Neural Information Processing Systems, 36, 2023
Yash Goyal et al. Making vision-language models robust to modality bias.Advances in Neural Information Processing Systems, 36, 2023
2023
-
[5]
Suhas Kotha et al. Understanding and mitigating gaslighting in large language models.arXiv preprint arXiv:2311.08332, 2023
Pith/arXiv arXiv 2023
-
[6]
Llms get lost in multi-turn conversation, 2025
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. Llms get lost in multi-turn conversation, 2025. URLhttps://arxiv.org/abs/2505.06120
Pith/arXiv arXiv 2025
-
[7]
Prompt repetition improves non-reasoning llms, 2025
Yaniv Leviathan, Matan Kalman, and Yossi Matias. Prompt repetition improves non-reasoning llms, 2025. URLhttps://arxiv.org/abs/2512.14982
arXiv 2025
-
[8]
Video-language modeling: A comprehensive survey.arXiv preprint arXiv:2305.00222, 2023
Kevin Lin et al. Video-language modeling: A comprehensive survey.arXiv preprint arXiv:2305.00222, 2023
Pith/arXiv arXiv 2023
-
[9]
Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36, 2024
2024
-
[10]
Inverse scaling: When bigger isn’t better.arXiv preprint arXiv:2306.09442, 2023
Ian McKenzie et al. Inverse scaling: When bigger isn’t better.arXiv preprint arXiv:2306.09442, 2023
Pith/arXiv arXiv 2023
-
[11]
Gpt-4 technical report
OpenAI. Gpt-4 technical report. Technical report, OpenAI, 2024. 14
2024
-
[12]
Training language models to follow instructions with human feedback
Long Ouyang et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744, 2022
2022
-
[13]
Sycophancy in large language models: A review.arXiv preprint arXiv:2403.00331, 2024
Yuting Pan et al. Sycophancy in large language models: A review.arXiv preprint arXiv:2403.00331, 2024
Pith/arXiv arXiv 2024
-
[14]
Ethan Perez et al. Discovering language model behaviors with model-written evaluations.arXiv preprint arXiv:2212.09251, 2022
Pith/arXiv arXiv 2022
-
[15]
The glance effect: Over-reliance on linguistic priors in vlms.arXiv preprint arXiv:2401.01234, 2024
Ziming Qiu et al. The glance effect: Over-reliance on linguistic priors in vlms.arXiv preprint arXiv:2401.01234, 2024
Pith/arXiv arXiv 2024
-
[16]
Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024
Rafael Rafailov et al. Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024
2024
-
[17]
Towards understanding sycophancy in language models.arXiv preprint arXiv:2310.13548, 2023
Mrinank Sharma et al. Towards understanding sycophancy in language models.arXiv preprint arXiv:2310.13548, 2023
Pith/arXiv arXiv 2023
-
[18]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. Technical report, Google, 2024
2024
-
[19]
Peng Wang et al. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2024
Pith/arXiv arXiv 2024
-
[20]
Jerry Wei et al. Simple synthetic data reduces sycophancy in large language models.arXiv preprint arXiv:2308.03613, 2023
Pith/arXiv arXiv 2023
-
[21]
Star: A benchmark for situated reasoning in real-world videos.Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021
Bo Wu et al. Star: A benchmark for situated reasoning in real-world videos.Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021
2021
-
[22]
Star v2: Scaling up situated reasoning in real-world videos.arXiv preprint arXiv:2210.02123, 2022
Bo Wu et al. Star v2: Scaling up situated reasoning in real-world videos.arXiv preprint arXiv:2210.02123, 2022
Pith/arXiv arXiv 2022
-
[23]
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng et al. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598, 2022
Pith/arXiv arXiv 2022
-
[24]
Yiyang Zhao et al. Visual gaslighting: Evaluating sycophancy in vision-language models.arXiv preprint arXiv:2402.05202, 2024. 15
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.