{"id":"54206896-3e1f-4aef-afe4-b8619b0e615b","arxiv_id":"2505.11687","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A proposal for the second Sim4IA workshop, centered on two micro shared tasks that will test how well user simulators can mimic real searcher behavior.","lead":"This paper lays out the plan for the 2025 Sim4IA workshop, including two micro shared tasks where teams design user simulators for ranked-list and conversational search. It explains why the workshop matters: to validate simulation methods and to prepare a larger TREC/CLEF evaluation campaign.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task B's synthetic ground truth contradicts the proposal's 'validate against real users' design principle; fidelity claims for conversational simulation would rest on no human utterances.","rationale":"Good-faith reading: this is a workshop proposal, not a research claim; its central promise is to pilot two shared tasks that validate user simulators and to use what is learned to plan a TREC/CLEF campaign. The most load-bearing condition for that promise is that the evaluation data actually contains real-user behavior to validate against. Task A satisfies this; Task B does not, because its 'actual user utterances' are synthetic. This is an internal inconsistency rather than a disagreement with consensus: §4.1 defines the shared-task concept as fidelity to real users, while §4.3 removes real users from the conversational task. The reader's weakest assumption (unspecified 'simple measures') points to the same evaluation-design area, but is partially mitigated by the paper's explicit statement that measure choice is intentionally left to workshop discussion; the Task B data-source problem is unaddressed. A fix is straightforward: add human conversational data, or explicitly frame Task B as a generator-comparison pilot. With that clarification, the proposal's central claim is defensible; without it, one of the two tasks cannot support the stated validation goal. I therefore recommend conditional acceptance rather than rejection or an unchanged 'unverdictable' status, since the issue is concrete and fixable.","tokens_in":6676,"tokens_out":5091,"duration_ms":51248,"concrete_test":"Replace or supplement Task B's synthetic ground truth with a held-out set of human conversational utterances from a real conversational search log, and recompute the semantic-similarity evaluation on that set. As a control, score a trivial baseline that repeats the previous user utterance and a no-context LLM baseline; if the trivial baseline is within noise of the top submitted systems, or if system rankings change substantially when human ground truth is used, the proposed Task B metric cannot validate fidelity to real users and the TREC/CLEF plan should be revised. If no human set is available, state explicitly in the paper that Task B is a system-comparison pilot, not a validation task.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The proposal's central value claim is that the two micro shared tasks validate user simulators against real-user behavior, 'instead of measuring system effectiveness' (§4.1). Task A is consistent with this: the LongEval/CORE dataset supplies held-out real-user queries, SERPs, and clicks. Task B is not: §4.3 states that the conversational data is 'generated synthetically based on traditional search logs,' so no human utterance exists as ground truth. The §4.1 statement that submitted simulators are evaluated on 'held-out data collected from real users' is therefore false for Task B. Semantic similarity to synthetic utterances can compare simulators with each other or with the generative process that produced the data, but it cannot establish that a simulator 'mimic[s] the interactions of real users with a high degree of fidelity.' Unless a human-utterance validation set is added, or the Task B claim is downgraded to a system-comparison pilot, the workshop cannot deliver the 'validating user simulations' design principle for one of its two tasks. The unspecified 'simple measures' are less problematic because §4.1 explicitly delegates measure choice to workshop discussion; the data-source contradiction is not flagged anywhere.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is the workshop proposal for the second Sim4IA workshop at SIGIR 2025. Its central aim is to test two micro shared tasks for user simulation—Task A on interactions with ranked lists and Task B on conversational utterances—as a stepping stone toward a future TREC/CLEF campaign. The proposal describes the design principle of validating user simulators against held-out real-user data rather than measuring system effectiveness, the use of the SimIIR 3 framework with LLM-based user actions, the timeline for dataset release and submission, and the planned interactive workshop format (keynote, panel, lightning talks, breakout groups). The paper also summarizes the motivation from the previous Sim4IA 2024 workshop and the broader literature on simulation-based evaluation.","tokens_in":7029,"tokens_out":6566,"duration_ms":61716,"significance":"If the proposed design were fully realized, the workshop would provide a tangible infrastructure for community-wide validation of user simulators, addressing a recognized gap in simulation-based IR evaluation. The commitment to distributing a Dockerized SimIIR environment with baseline configurations lowers the barrier to entry, and the use of LongEval/CORE interaction data for Task A grounds one task in real user behavior. The paper is honest about open design questions and positions the workshop as a discussion forum, which is appropriate for a workshop proposal. However, as detailed below, the inconsistency between the stated validation principle and Task B's synthetic ground truth undermines the significance of one of the two tasks, and the underspecified evaluation measures leave the 'validation' deliverable incompletely defined.","major_comments":[{"comment":"The manuscript's central validation claim is contradicted by the Task B data description. §4.1 states that 'The submitted user simulators are evaluated on the basis of held-out data collected from real users,' and the stated design principle is 'validating user simulations instead of measuring system effectiveness.' However, §4.3 specifies that the conversational data for Task B 'is generated synthetically based on traditional search logs,' so no real-user utterance exists as the 'actual user utterance' against which semantic similarity is computed. Semantic similarity to synthetic utterances can measure agreement with the generative process that produced the test data, but it cannot establish that a simulator 'mimic[s] the interactions of real users with a high degree of fidelity.' The paper should either add a real-user conversational dataset (e.g., from TREC iKAT or a similar resource) for Task B evaluation, or explicitly reframe Task B as a system-comparison pilot whose results are not evidence about real-user fidelity. As written, the paper delivers on its stated validation goal for only one of its two tasks.","section":"§4.1 and §4.3"},{"comment":"The evaluation protocol is underspecified in a way that affects the feasibility of the shared tasks within the stated timeline. §4.1 says submitted simulators will be evaluated with 'simple measures' but explicitly defers the choice of measures to workshop discussion, even though Table 1 shows submissions due 27 June and the workshop on 17 July. It is unclear how the organizers can produce the promised evaluation—which presumably feeds the lightning talks and breakout discussions—if the measures are not fixed before submissions are due. Please specify at least a provisional set of candidate measures (e.g., click-through alignment, ranking correlation, utterance-level semantic similarity) or describe a pre-registration procedure, even if the final measures are to be discussed at the workshop. Without this, the micro shared tasks cannot operate as the validation exercises the paper claims them to be.","section":"§4.1 and Table 1"}],"minor_comments":[{"comment":"The day '13rd June' should be '13th June'.","section":"Table 1"},{"comment":"'InProceedings' lacks a space; it should read 'In Proceedings'.","section":"ACM Reference Format"},{"comment":"The text refers to 'SimIIR 3.02' while the citation in [2] is titled 'SimIIR 3'; please align the version string.","section":"§3"},{"comment":"Consider stating a tentative semantic-similarity measure (e.g., embedding cosine or BERTScore) so participants have a concrete target for their submissions.","section":"§4.3"},{"comment":"The phrase 'so that simulations can be based on previous queries and interactions with the result lists' leaves unclear whether document-level content will be provided with the SERP data; specifying the exact data fields would improve reproducibility.","section":"§4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a workshop proposal, and the central issue (Task B ground truth) can be fixed by rephrasing or changing the dataset. The evaluation-metric ambiguity is also addressable with a provisional plan. I would encourage inviting a revision rather than rejecting, provided the organizers clarify these points. The paper's value is organizational and community-building; it does not claim a research result, so the bar for internal consistency should reflect that context."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a workshop proposal, not a scientific result, and the one thing worth remembering is that the evaluation design has a genuine internal contradiction. Section 4.1 states the guiding principle: validate user simulations against held-out data collected from real users. Task A fits that principle—LongEval/CORE supplies real queries, SERPs, and clicks. Task B does not. Section 4.3 says the conversational data is \"generated synthetically based on traditional search logs\" and evaluation is semantic similarity to the \"actual\" utterances. Those actual utterances are synthetic, so no human ground truth exists. You can use Task B to compare simulators with each other or with the generative process that produced the data, but you cannot claim to validate fidelity to real users without an external human-utterance set. The stress-test note is right about this.\n\nWhat the paper does well: it is a coherent, honest proposal for a community-building event. The micro shared tasks are concrete, the timeline is realistic, and the decision to package SimIIR, PyTerrier, and baseline simulators in a Docker environment is considerate of participants. Delegating the choice of evaluation measures to workshop discussion is reasonable at this stage, not a flaw. The related-work coverage is solid, and the self-citations are legitimate—Sim4IA 2024 and SimIIR 3 are the natural foundations.\n\nSoft spots beyond Task B: there is no scientific result, but that is inherent to a workshop proposal, not a defect. The \"simple measures\" are unspecified, but the authors flag that intentionally and invite discussion. One minor issue: Figure 1's \"held-out validation environment\" is drawn generically and does not disclose that Task B's held-out data is synthetic, so the figure and §4.1 overstate the consistency of the two tasks.\n\nWho this is for: IR evaluation researchers, shared-task organizers, and anyone planning a TREC/CLEF simulation track. The community value is real—there is a genuine gap in venues for user-simulation evaluation.\n\nRecommendation: this deserves a serious referee, not a desk reject. A referee should ask the authors to either add a real-user conversational utterance set for Task B or explicitly re-scope Task B as a system-comparison pilot, and adjust §4.1 and Figure 1 accordingly. With that revision, it is a solid workshop proposal.","headline":"A well-scoped workshop proposal whose one real flaw is internal: Task B's synthetic conversational data cannot validate simulators against real users as §4.1 claims, and the paper deserves revision rather than rejection.","tokens_in":7373,"tokens_out":2839,"would_cite":false,"duration_ms":28181,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This workshop proposal argues that user simulators in information access can be validated in shared tasks that score simulated interactions against held-out real-user data, and that doing so on a small scale is the right path to a full…","keywords":["Information Access","Simulation","User Models","Evaluation","Shared Tasks","Conversational Search","Interactive Retrieval","LLM-based Simulation"],"falsifier":"The proposal would fail its own test if the shared tasks run and no submitted configuration scores closer to the held-out real-user data than the provided baseline does, or if organizers cannot agree on a measure because every candidate measure ranks the same configurations differently.","tokens_in":6537,"feed_emoji":"🧪","tokens_out":7435,"duration_ms":72473,"temperature":0.7,"pith_summary":"The paper, a proposal for the second iteration of a workshop on simulations for information access, aims to turn user simulation from a promising idea into a testable evaluation method. It defines two micro shared tasks: one in which participants build a simulator of search-session interactions with ranked result lists, and one in which they predict the next user utterance in a conversational session. Submitted simulator configurations are run in a packaged environment and scored against held-out data from real or synthetic users. If the tasks work, the organizers will use what they learn about validation measures and pitfalls to design a full shared-task campaign at a major evaluation venue.","feed_headline":"Simulated users face real-data validation in two shared tasks","feed_subtitle":"Simulators are scored against held-out interactions, a first step toward a TREC/CLEF campaign.","key_machinery":"The central object is a packaged development-and-evaluation loop: a dockerized environment containing a search system, a baseline user simulator, and data, so that participants change only the simulator configuration and its prompt design. The validation mechanism is held-out data—Task A uses real-user search sessions, Task B uses synthetic conversational exchanges derived from search logs—and the fidelity measures, which the workshop will define. Two properties carry the argument: participants need not build infrastructure, and organizers score against data the participants never see, so the comparison is a fidelity test rather than a system-effectiveness test.","core_discovery":"The paper's central claim is that user simulators can be validated, not just built, and that a shared-task format is a workable way to do that validation. In Task A, participants design a simulator of a search session whose queries, result pages, and clicks come from longitudinal real-user interaction data; in Task B, they predict the next user utterance given a conversational history built from synthetic search logs. Submitted configurations are run in a packaged environment with a baseline simulator and then scored by held-out data. The scoring measures are deliberately not fixed by the paper; they are part of the workshop's agenda. If the two micro tasks succeed, the organizers take the lessons into a full shared-task campaign.","pith_inferences":["The paper leaves implicit that a validated search-session simulator could eventually substitute for live-user studies in interactive retrieval evaluation, lowering cost and improving reproducibility.","Because Task A uses organic user data and Task B uses synthetic data, comparing the two tasks may reveal whether conversational simulators trained on synthetic logs transfer to real user behavior.","A natural extension would be a meta-task in which submitted configurations are scored by whether they generalize across unseen topics and result pages, testing robustness rather than a single held-out split.","The open choice of evaluation measures is both an enabler and a risk: the workshop may spend as much effort agreeing on what fidelity means as on producing leaderboards."],"forward_implications":["If the micro shared tasks yield usable fidelity scores, the organizers can carry a concrete shared-task design into a full evaluation campaign.","Participants can enter through a packaged environment with a baseline simulator, so effort goes into modeling user behavior rather than system engineering.","LLM-based prompt designs become the primary experimental variable for both query formulation and conversational utterance prediction.","The workshop's outcomes—configurations and working notes—are shared openly, giving the community a reusable starting point rather than one-off results.","The evaluation protocol itself is part of the study: because the measures are left open, the workshop is also an experiment in how to measure simulation fidelity."],"supporting_citations":[{"why":"Supplies the packaged simulation framework with both ranked-list and conversational user support, the environment the shared tasks are built on.","marker":"[2]"},{"why":"Sets out the theory that formalized user models make assumptions about user behavior explicit, the intellectual basis for validating simulations.","marker":"[6]"},{"why":"Documents the first workshop's conclusion that shared-task design remains unsolved and a follow-up is needed, motivating these micro tasks.","marker":"[7]"},{"why":"Provides the longitudinal real-user search interaction data that Task A's held-out validation is drawn from.","marker":"[8]"},{"why":"Demonstrates prior use of user simulations in a conversational interactive track, a precedent for Task B's conversational setting.","marker":"[1]"}],"fun_headline_variants":["Simulated users get a reality check","Workshop validates user simulators","Sim4IA: testing simulators against real data","Shared tasks to prove simulator worth","Simulation workshop eyes TREC/CLEF campaign"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that comparisons between simulated and held-out real-user behavior, made with simple measures, can tell which simulator is more faithful, even though the paper leaves the choice of those measures entirely open to workshop discussion.","fun_headline_variants_meta":{"raw":{"variants":["Simulated users get a reality check","Workshop validates user simulators","Sim4IA: testing simulators against real data","Shared tasks to prove simulator worth","Simulation workshop eyes TREC/CLEF campaign"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000112,"raw_usage":{"total_tokens":987,"prompt_tokens":796,"completion_tokens":191,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":412,"completion_tokens_details":{"reasoning_tokens":126}},"tokens_in":412,"tokens_out":191,"duration_ms":2536,"temperature":1.0,"reasoning_tokens":126,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:48:38.723634+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The proposal would fail its own test if the shared tasks run and no submitted configuration scores closer to the held-out real-user data than the provided baseline does, or if organizers cannot agree on a measure because every candidate measure ranks the same configurations differently.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the packaged simulation framework with both ranked-list and conversational user support, the environment the shared tasks are built on."},{"cited_title":"Report on the Workshop on Simulations for Information Access (Sim4IA 2024) at SIGIR 2024","cited_arxiv_id":"2409.18024","evidence_quote":"Documents the first workshop's conclusion that shared-task design remains unsolved and a follow-up is needed, motivating these micro tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the longitudinal real-user search interaction data that Task A's held-out validation is drawn from."}],"review_version":1}