{"id":"14d86a0f-0e2e-47a4-ab76-44324e82c3aa","arxiv_id":"2607.22392","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In 8,133 real multi-turn AI conversations, the final user prompt contains only ~36% of the session's unique content vocabulary and reproduces the full detected request-state dimension set in only ~26% of dimension-bearing conversations.","lead":"This paper measures how much of a multi-turn AI conversation's request is captured by the final user prompt, finding that the final prompt typically contains only about a third of the session's unique content vocabulary and often omits explicit requirements stated earlier. It argues that AI-search evaluation should treat the whole session state, not the isolated prompt, as the unit of measurement.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"History-only dimension rates may be inflated by superseded constraints: the rules count any earlier mention even if later negated, conflicting with the construct's 'needed now' definition.","rationale":"The reader's weakest_assumption is the validity of the nine cue families, and my concern is a specific, high-impact failure mode within that assumption: the rules detect prior mentions without regard to whether those mentions are still active, despite the paper's own definition of request state as 'needed to interpret the user’s current turn.' This is a load-bearing concern because the headline 'history-only' metrics directly support the paper's central claim that the final prompt is not a self-contained query. The paper is transparent about this limitation in Section 8, but it does not quantify how often the flagged dimensions are superseded, so the reader cannot assess whether the primary percentages are meaningful. Since the issue is concrete and testable, and the paper's qualitative conclusion would likely survive if the active fraction is high, a conditional recommendation—require the authors to run the annotation-based sensitivity check before final acceptance—is appropriate. This is not a rejection: the paper's other evidence (lexical coverage, final-added dimensions, replication across PRISM) is substantial, and the manuscript explicitly disclaims causal claims. But the internal mismatch between 'needed now' and 'union of ever-mentioned' is a real gap that should be closed before the percentages are cited as established.","tokens_in":9408,"tokens_out":4246,"duration_ms":55107,"concrete_test":"Draw a random sample of ~300 conversations from PRISM (or the commercial corpus if accessible) where the automatic rules flag at least one history-only dimension. For each flagged dimension, provide two annotators the full transcript up to the final turn and ask: (a) Is the dimension explicitly negated, superseded, or declared irrelevant in any later turn before the final? (b) Is the dimension still an active constraint/reference needed to interpret the final turn? Compute the proportion of history-only flags that are 'active' vs 'superseded/irrelevant' with inter-annotator agreement. Recompute the headline history-only percentages after removing superseded flags. If the adjusted rates remain above, say, 40%, the central claim survives; if they drop to near zero, it fails. Also report precision/recall of the nine cue families against this human annotation to validate the instrument gener","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central categorical claim—that 50.3% (commercial) and 44.8% (PRISM) of conversations leave at least one request-state dimension in history—depends on the nine cue families detecting only dimensions that are actually part of the current request state. Section 3.1 defines conversation-conditioned request state as the 'explicit task specification and discourse dependence needed to interpret the user’s current turn.' But the operationalization in Section 5.2 unions all dimensions ever mentioned (S_T = ∪D_i) without filtering for negation, supersession, or irrelevance. Thus a user who says 'under $900' in turn 2 and then 'budget doesn't matter' in turn 4 yields a price dimension that is detected in history and absent from the final prompt, and the conversation is counted as having a history-only dimension even though the dimension is no longer part of the request state. The paper's own Limitations (Section 8) acknowledge that 'a missing dimension in the final prompt can be active, superseded, or irrelevant; transcript-only rules cannot decide which.' Yet the headline metrics treat all such missing dimensions as evidence that the final prompt is not self-contained. If a substantial fraction of history-only flags are superseded constraints, the percentages 50.3% and 44.8% overstate the phenomenon the paper claims to measure. This is not merely a question of precision/recall of the cue families—it is an internal mismatch between the construct definition and its measurement. The reader's weakest_assumption (cue-rule validity) captures this, but the specific supersession mechanism is a concrete, testable threat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that in multi-turn human–LLM conversations the final user prompt is not a self-contained query but a state update over a session-level request state. Using 670 proprietary commercial conversations (discovery + replication) and 7,463 public PRISM conversations, it reports that the final prompt contains a median 35.6–36.4% of unique user-side content vocabulary, leaves at least one detected request-state dimension in history in 50.3% (commercial) and 44.8% (PRISM) of conversations, and reproduces the full detected dimension set in only 26.1–26.2% of dimension-bearing conversations. A length-matched null shows that the low lexical coverage is largely a length effect, so the lexical results are interpreted as information availability. The paper concludes that session-level measurement, not isolated-prompt measurement, should be the default for conversational AI search.","tokens_in":9728,"tokens_out":6066,"duration_ms":73619,"significance":"If the categorical results are valid, the paper makes a useful methodological contribution to conversational search evaluation: it gives an observable, transcript-level construct (request state), a reproducible measurement framework, a discovery-replication design, participant-clustered intervals for PRISM, and an explicit boundary between observational description and causal claims. The use of a length-matched null for the lexical outcome is a nice check, and the paper is transparent about its limitations, including the inability to infer latent intent. The main fragility is the categorical instrument: the nine cue families are reused from prior work without validation, and the cumulative-state operationalization conflicts with the construct definition by counting superseded constraints as active. These issues are load-bearing because every headline categorical percentage depends on them.","major_comments":[{"comment":"The central construct is defined in §3.1 as the explicit task specification 'needed to interpret the user’s current turn,' but S_T is operationalized as the union of all dimensions ever mentioned, with no filtering for negation, supersession, or irrelevance. A user who says 'under $900' in turn 2 and 'budget doesn’t matter' in turn 4 produces a history-only dimension that is no longer part of the request state. The paper’s own Limitations (§8) admit that a missing dimension can be 'active, superseded, or irrelevant; transcript-only rules cannot decide which.' The headline percentages 50.3%, 44.8%, 26.1%, and 26.2% therefore overstate the phenomenon if a substantial share of history-only flags are superseded. This is an internal mismatch between construct and measurement, not merely a precision/recall issue. Please provide a sensitivity that excludes or marks later-negated constraints (e.","section":"§3.1, §5.2, §8"},{"comment":"All categorical results depend on nine transparent cue families (price, location, persona, attribute, time, alternatives, correction, comparison, evidence), yet no human-agreement, precision, or recall validation is reported. The rules are reused frozen from a preceding study, which makes them reproducible but not necessarily valid for the present endpoints. False positives or false negatives in these patterns shift every categorical percentage in Table 3, including the final-completeness and history-only rates. Please report validation on a labeled subsample, or at minimum include the full rule definitions in an appendix so readers can assess specificity, and soften claims accordingly. Transparency about the rules is a strength, but it does not by itself establish that the rules measure the defined construct.","section":"§5.2, Table 3"}],"minor_comments":[{"comment":"The affiliation line contains a typo: 'Tel A viv' should be 'Tel Aviv.'","section":"Author affiliation"},{"comment":"The label 'Final half' in Figure 1 is ambiguous; consider renaming to 'Final ≤ half' to match the text and Table 3.","section":"Fig. 1"},{"comment":"The phrase 'the rules are conservative, so these are demonstrated events, not estimates of all semantic change' is not a formal statistical statement. Consider reporting uncertainty or at least clarifying that 'conservative' refers to the pattern design, not to the handling of supersession.","section":"§6.4"},{"comment":"The denominator for the final-completeness row (n=456 commercial, n=4,534 PRISM) is given in the text but not in the table caption; moving it to the caption would improve readability.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid candidate for publication in a conversational search or evaluation venue, and the empirical base is substantial. The main risk is that the central categorical claims rest on an unvalidated keyword instrument and on a cumulative-state definition that does not handle supersession, despite the limitations being openly acknowledged. I would be willing to support acceptance if the author supplies a validation appendix for the cue families and a sensitivity analysis addressing superseded constraints. The heavy reliance on the author's preceding study is also a concern; if that study is not publicly available, the frozen rules should be included in the supplement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this paper. It makes a real contribution to conversational IR measurement — but the headline categorical numbers are softer than they look because the rules count superseded constraints as 'history-only.' The authors know this, admit it in limitations, and then go ahead and use the numbers anyway.\n\nThe paper's main idea is useful: separate the turn-local prompt from the cumulative request state, and measure how much of the state sits outside the final prompt. The design is careful: discovery-replication cohorts, external PRISM validation, participant-clustered CIs, length-matched nulls, and a battery of sensitivities. The lexical result is honestly interpreted — low coverage is mostly a turn-length artifact, not evidence of semantic drift. The endpoint-added result (about one in five final prompts introduces a new dimension) is a nice counterpoint, showing the endpoint is not just a compressed summary.\n\nThe soft spot is real. The construct is defined as what's needed to interpret the current turn, but the measurement is a union of all dimensions ever mentioned. A user who says 'under $900' and later 'budget doesn't matter' contributes a history-only price dimension even though that constraint is gone. The paper's limitation section says 'a missing dimension can be active, superseded, or irrelevant,' which is exactly the point. The headline percentages (50.3%, 44.8%) are therefore upper bounds on 'potentially missing request state,' not the actual rate. How much they overstate is unknown; the paper doesn't quantify it, and the nine cue families have no human validation. That absence is a bigger deal than a precision/recall footnote because it's load-bearing for the central claim.\n\nThat said, the lexical result and the endpoint-added result stand independently, and the paper's broader point — that replaying an isolated final prompt is not the same as replaying the session — still has legs. I'd guess a human-annotated audit would show the true history-only rate is lower but still substantial.\n\nThis is a paper worth sending to referees. It's clear, transparent, and useful to anyone building evaluation pipelines for conversational AI. If I were editor, I'd ask for a supersession analysis or a manual audit of a sample before accepting, but it deserves the attention.\n\nQuote it if you work in this area; bring it to a reading group.","headline":"A careful, transparent study of how much request state sits outside a final prompt, but its headline 'history-only' numbers are inflated by superseded constraints—still worth citing and sending out with a request for a supersession analysis.","tokens_in":10198,"tokens_out":4929,"would_cite":true,"duration_ms":57056,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The final prompt in a multi-turn AI conversation is not a self-contained query: it typically carries only about a third of the session's user-side content vocabulary, and in roughly half of conversations at least one explicit request dimens","keywords":["conversational AI","request state","multi-turn conversation","prompt evaluation","AI search","lexical coverage","user intent","discourse dependence"],"falsifier":"A human-annotation study on a random sample of several hundred conversations from both corpora, using the same dimension definitions: if annotators' agreement with the rule-based dimension sets is low, or if human-labeled history-only dimension rates fall well below 50%, the claim that final prompts leave explicit request state in history would not survive. Alternatively, showing the final prompt alone to users and asking them to state the full request would settle whether the prompt is self-contained.","tokens_in":9287,"feed_emoji":"💬","tokens_out":5248,"duration_ms":47649,"temperature":0.7,"texified_at":"2026-08-05T21:41:04.951202+00:00","pith_summary":"This paper tries to establish that in multi-turn AI conversations, the final user prompt is not a self-contained query: it usually carries only about a third of the session's user-side content vocabulary, and in roughly half of conversations at least one explicit request dimension (price, location, persona, attribute, time, alternatives, correction, comparison, evidence) exists only in earlier turns. The author replaces latent 'intent' with an observable, transcript-level construct—conversation-conditioned request state—and shows that the final prompt is typically a state update rather than a summary: it often adds a new dimension while leaving another in history. A sympathetic reader cares because evaluation, benchmarking, and analytics that replay isolated prompts are measuring a different object from what a conversational model actually responds to. If the finding holds, session-level measurement, including the history policy, becomes part of the unit of analysis rather than an optional context.","texify_model":"deepseek-v4-flash","texify_usage":{"total_tokens":5871,"prompt_tokens":808,"completion_tokens":5063,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":808,"completion_tokens_details":{"reasoning_tokens":4344}},"feed_headline":"Final prompt keeps only a third of the session's request","feed_subtitle":"Half of AI conversations bury request details in earlier turns; measure the session, not the last prompt.","key_machinery":"The carrying object is the cumulative request state $S_t = \\bigcup D_i$, where $D_i$ is the set of explicit request-state dimensions detected in user turn $i$ by nine frozen, transparent case-insensitive cue families (price, location, persona/use case, attribute, time, alternatives, correction/redirect, comparison/evaluation, explanation/evidence). The key identities are the history-resident set $H_T = S_{T-1} \\setminus D_T$ and the endpoint-added set $A_T = D_T \\setminus S_{T-1}$; when both are nonempty, the final prompt is a delta, not a summary. Lexical coverage $L_T = |V_T|/|V|$ with a length-matched null isolates how much unique vocabulary the final turn makes locally available. These definitions turn an unobserv","core_discovery":"The paper's central claim is that the endpoint of a multi-turn conversation is neither an independent query nor a faithful summary of the request. Across 8,133 real conversations (670 commercial, 7,463 public), the final prompt contains a median 35.6–36.4% of the session's unique user-side content vocabulary, and contains at most half in 68.4–74.3% of conversations. Using nine transparent cue families to detect explicit request-state dimensions, the author finds that at least one dimension is history-resident (present in earlier turns but absent from the final prompt) in 50.3% of commercial and 44.8% of public conversations; the final prompt reproduces the full detected dimension set in only","pith_inferences":["If the cue families are valid, the same session-state measurement could be extended to detect when an earlier dimension is superseded rather than merely missing, using assistant response turns—this would distinguish active constraints from abandoned ones.","A practical testable extension: benchmark suites could report a 'request-state coverage' score for a model's answer, measuring whether constraints carried in history are satisfied in the response, not just whether the final prompt mentions them.","The result implies that user-side state migration (moving from 'I need X' to 'not that one') could be tracked as a structured signal for adaptive systems, but this requires per-turn dimension parsing beyond the current aggregate percentages."],"forward_implications":["Isolated-prompt panels and replay benchmarks measure a different object from what a conversational model sees; they should record turn boundaries, cumulative user state, and the supplied history.","Evaluation designs that hold the final turn and model fixed and compare full interleaved history, user-only history, and isolated final turn are directly motivated; the history policy becomes part of the treatment.","The endpoint adding a new dimension in roughly one fifth of conversations means a final prompt cannot be treated as a stable query even when it looks locally short.","Depth is descriptive but consistent: longer sessions concentrate more request evidence outside the final turn, so fixed-depth cohorts are needed for fair comparison."],"fun_headline_variants":["Last prompt holds just 36% of the request in AI chats","Two in three AI chats bury request info before the final prompt","Measure the session: final prompt covers only a third of the request","AI search: final prompt misses half the request in 68% of chats"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The nine transparent cue families, reused without retuning and without human-agreement or precision/recall validation, correctly detect the explicit request-state dimensions they are claimed to detect; if these keyword patterns are biased, the central percentages (50.3%, 44.8%, 26.1%, 26.2%) could shift.","fun_headline_variants_meta":{"raw":{"variants":["Last prompt holds just 36% of the request in AI chats","Two in three AI chats bury request info before the final prompt","Measure the session: final prompt covers only a third of the request","AI search: final prompt misses half the request in 68% of chats"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000292,"raw_usage":{"total_tokens":1605,"prompt_tokens":874,"completion_tokens":731,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":655}},"tokens_in":618,"tokens_out":731,"duration_ms":8131,"temperature":1.0,"reasoning_tokens":655,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:52:23.503387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-annotation study on a random sample of several hundred conversations from both corpora, using the same dimension definitions: if annotators' agreement with the rule-based dimension sets is low, or if human-labeled history-only dimension rates fall well below 50%, the claim that final prompts leave explicit request state in history would not survive. Alternatively, showing the final prompt alone to users and asking them to state the full request would settle whether the prompt is self-contained.","supporting_citations":[],"review_version":1}