{"id":"f2c9b7ac-6103-459b-886a-2942b0b93473","arxiv_id":"2505.21907","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI copilot preference optimization is organized into a pre-, mid-, and post-interaction taxonomy, with a unified definition of AI copilots.","lead":"This paper surveys how AI copilots detect and adapt to user preferences, organizing existing research into a taxonomy of pre-, mid-, and post-interaction techniques. It aims to give designers a unified framework, but serious citation problems weaken the survey's grounding.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2.3 and Table 6 rest on references that cannot be resolved or do not support the claims made; the survey's literature-grounded taxonomy is therefore unsubstantiated at a load-bearing point.","rationale":"The reader's verdict identifies citation integrity as the weakest assumption, and my independent reading confirms that this is the correct load-bearing concern. For a survey whose stated contribution is a 'literature-grounded' taxonomy, the citations are not decorative; they are the evidence that the proposed categories correspond to existing research. Section 4.2.3 is particularly important because it describes post-interaction feedback as a distinct phase, and it cites Refs [63], [64], and [65] for that phase. All three have arXiv IDs ending in 'XXXX', which by convention are placeholders and cannot be verified. Additionally, Ref [64] appears to duplicate Ref [30] under different authors, and Ref [115], cited for DPOC, is actually an 'In dialogues we learn' paper rather than a DPOC paper. These are not minor typographical issues: they affect the core evidence for the post-interaction portion of the taxonomy and for the DPOC entry in Table 6. The paper may eventually be correctable if the missing references are supplied and the misattributions fixed, but as submitted, the central claim is not supported. I therefore agree with the reader's recommendation to reject as currently written, and I would not change the verdict. I would add that the concrete verification step should focus on resolving the four problematic references and checking whether the cited papers actually support the specific sentences in Section 4.2.3, since that is the section where the taxonomy makes its most distinctive claim.","tokens_in":19946,"tokens_out":2297,"duration_ms":24669,"concrete_test":"Resolve Refs [63], [64], [65], and [115] through arXiv's API, DBLP, and Google Scholar. First, check whether the four 'XXXX' identifiers correspond to real papers and whether the titles, authors, and abstracts match the claims made in Section 4.2.3 and Table 6. Second, specifically compare Ref [64] with Ref [30] to determine whether they are the same paper with different authors. Third, check whether Ref [115] describes DPOC (Direct Preference Optimization with Criterion) or is instead the 'In dialogues we learn' paper. If the placeholder IDs cannot be resolved, or if Ref [115] is not a DPOC paper, then the post-interaction feedback section and the DPOC row of Table 6 lack a verified source, and the survey's literature-grounded claim fails at that point.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim is that its taxonomy is literature-grounded. The most load-bearing failure is in Section 4.2.3 ('Post-Interaction Feedback and Preference Refinement'), where the entire post-interaction phase is supported by references that cannot be resolved or do not match the cited claims. Refs [63], [64], and [65] are listed with arXiv identifiers ending in 'XXXX' ('2310.XXXX', '2306.XXXX', '2308.XXXX'), so their existence and content cannot be verified. Ref [64] is titled identically to Ref [30] ('When to show a suggestion? integrating human feedback in ai-assisted programming') but has a different author list and is used to support a different claim about judging the appropriateness of suggestions. Ref [115] is cited for DPOC (Direct Preference Optimization with Criterion) in Table 6 and Section 4.2.3, but the reference entry is titled 'In dialogues we learn' and appears to duplicate Ref [61] with a different author list; it is not a DPOC paper. Because Section 4.2.3 is the section that differentiates this survey's post-interaction taxonomy from earlier work, these citation failures remove the evidential base for a core contribution. This is not a stylistic issue: for a survey, citation integrity is the evidence, and a taxonomy built on unverifiable or misattributed sources cannot be accepted as literature-grounded.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of preference optimization in AI copilots. It proposes a unified definition of AI copilots, categorizes preference signals into human-driven and LLM-generated sources, organizes detection techniques into pre-, mid-, and post-interaction phases, reviews personalized response generation at prompt, fine-tuning, and architecture levels, and surveys RLHF and DPO alignment methods, including a generalized DPO loss. The main claimed contribution is a literature-grounded taxonomy of preference optimization across the copilot interaction lifecycle.","tokens_in":20241,"tokens_out":7642,"duration_ms":67827,"significance":"If the claims were fully supported, the survey would fill a real gap: it connects user modeling, human-AI interaction, and preference alignment in one framework, and the generalized DPO objective in Section 4.4.2 is a concise unifying device. The comparison tables are helpful for orientation. However, the evidence for the post-interaction stage and several DPO variants rests on citations that are missing, duplicated, or mischaracterized. Because the central claim is literature grounding, these citation problems directly affect the contribution. I credit the authors for a well-organized structure and a useful synthesis of many verifiable sources, but the unresolved references prevent acceptance in the current form.","major_comments":[{"comment":"The DPOC (Direct Preference Optimization with Criterion) entry is unsupported. The text states that preference optimization 'using direct feedback criteria collected after interaction' is described in [61] and then names this method DPOC, and Table 6 lists DPOC[115] with the regularization P(r_a,r_b) = −min(0, log r_a − log r_b). However, [61] and [115] are the same paper, 'In dialogues we learn...' (arXiv:2403.03102), which proposes in-dialogue learning (IDL), not a DPO variant with criterion-based penalties. No source in the manuscript supplies the DPOC formula or its name. This removes the evidence for the 'After the Conversation' DPOC row in Table 3 and the DPOC row in Table 6.","section":"§4.2.3, Table 6"},{"comment":"References [63], [64], and [65] cannot be resolved. Entries [63] and [65] have arXiv identifiers ending in 'XXXX' (2310.XXXX and 2308.XXXX), and [64] has arXiv:2306.XXXX; the existence and content of these works therefore cannot be verified. These references are the sole support for adaptive preference learning strategies in §4.2.2 and for active preference learning and the AFSPP framework in §4.2.3. Without them, the post-interaction refinement discussion is unsubstantiated. The authors must either supply complete, findable references that support each claim or rewrite these paragraphs using verifiable sources.","section":"§4.2.2, §4.2.3, Refs [63]–[65]"},{"comment":"Reference [64] duplicates [30] with a different author list. Both entries are titled 'When to show a suggestion? integrating human feedback in ai-assisted programming'; [30] is the actual AAAI 2024 paper by Mozannar et al., while [64] lists 'Weiyan Xu, Abigail See, et al.' with a placeholder arXiv ID. In §4.2.3, [64] is used to support the claim that suggestion appropriateness is judged from prior user intent and adjustments are made for future turns, which is the topic of [30]. As written, the citation does not refer to a distinct, verifiable source.","section":"Reference [64] vs. [30]"},{"comment":"The Heimdall citation is not verifiable. The reference is authored only as 'Anonymous' with no arXiv identifier or DOI, so the claim in §4.1.1 that 'frameworks such as Heimdall [49] aggregate and anonymize user data securely' cannot be checked. The authors must provide a full citation with author, venue, and identifier, or the claim should be removed.","section":"§4.1.1, Ref [49]"},{"comment":"The IRPO entry does not match the cited paper. The text describes 'Information-Ratio Preference Optimization (IRPO)' with 'auxiliary compression and coverage terms', but reference [113] is 'Iterative Reasoning Preference Optimization' by Pang et al. (NeurIPS 2024), a method for iterative improvement on reasoning tasks. Neither the name nor the regularization described in Table 6 corresponds to [113]. This is a mischaracterization of a listed DPO variant and must be corrected or the entry removed.","section":"§4.4.2, Table 6, IRPO entry"}],"minor_comments":[{"comment":"The heading 'Priliminary and backgrounds' should read 'Preliminary and Background'.","section":"§2 heading"},{"comment":"Several reference entries are incomplete: [59] ('Zhu, Li, Mao, Pandelea, and Cambria'), [60] ('Kong Hao'), and [67] ('Liu Huang, Fu et al.') omit author initials or full names; these should be completed.","section":"References [59], [60], [67]"},{"comment":"The caption says 'Table adapted from [2]', but entries such as DPOC [115] and IRPO [113] do not appear in [2]; the provenance of each row should be clarified.","section":"Table 6 caption"},{"comment":"The notation r_a and r_b in the DPOC regularization term is used without definition; if the row is retained, these variables should be defined.","section":"§4.4.2, Table 6"},{"comment":"The phrase 'fouth edition' in reference [27] should be 'fourth edition'.","section":"Reference [27]"}],"recommendation":"major_revision","confidential_remarks":"The citation problems are serious enough that I would want to see a fully corrected reference list and a revised Section 4.2.3 before judging the taxonomy's grounding. If the placeholder entries cannot be replaced by genuinely supporting sources, I would move to reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a genuinely usable taxonomy for preference optimization in AI copilots, organized into pre-, mid-, and post-interaction phases, and a sensible unified definition of what counts as an AI copilot. That is real organizational work. Tables 2-4 are fine, and the DPO variant table is neatly adapted from [2] with credit.\n\nThe soft spot is not subtle. Section 4.2.3, the one section that makes this survey different from earlier ones, is grounded in references that do not exist as written. [63], [64], and [65] have arXiv IDs ending in 'XXXX.' [49] is 'Anonymous' with a venue that does not resolve. And [115], cited for DPOC in both the text and Table 6, is actually the 'In dialogues we learn' paper, which is already [61] with a different author list. The DPOC method as described has no citable source in this paper. The same section also cites [64] for active preference learning, but [64] is a duplicate title of [30], the actual 'When to show a suggestion?' paper. So the entire post-interaction analysis rests on placeholders and misattributions.\n\nFor a survey, this is a load-bearing failure. The taxonomy may be plausible, but the claim that it is literature-grounded cannot be checked. This is likely bibliography sloppiness rather than fabrication, since the rest of the references look real, but it is disqualifying for the current version.\n\nThe reader's REJECT verdict is fair. I would desk-reject the current version but invite the authors to fix the references, verify each one, and resubmit. The conceptual framework is worth preserving, and with a clean citation base it could be a useful resource for designers and researchers working on personalized copilots. I would not cite it myself until that is done.","headline":"A well-organized survey with a useful pre/mid/post taxonomy, but the post-interaction section rests on placeholder and misattributed references, so the current version cannot be trusted as literature-grounded.","tokens_in":20749,"tokens_out":6191,"would_cite":false,"duration_ms":57171,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey proposes a unified definition of AI copilots and a three-phase taxonomy of preference optimization.","keywords":["AI copilots","preference optimization","personalization","user feedback","taxonomy","human-AI interaction","direct preference optimization","personalized response generation"],"falsifier":"Resolve every reference cited in Sections 4.1-4.4 and compare each attributed claim against the actual source; in particular, check whether [49], [63], [64], and [65] correspond to real published work and whether [64] is distinct from [30]. If the claims tied to those entries cannot be traced, the survey's evidential grounding for the post-interaction feedback category fails.","tokens_in":19785,"feed_emoji":"🗺️","tokens_out":8313,"duration_ms":75225,"temperature":0.7,"pith_summary":"This paper argues that personalization in AI copilots has a common core: preference optimization, the ability to detect, interpret, and align with what a user wants. It proposes a unified definition of an AI copilot as an interactive, task-oriented system that supports knowledge workers through real-time assistance, domain adaptation, and human-guided decision-making. The central contribution is a taxonomy that organizes preference handling into sources (explicit, implicit, hybrid, interactive, and LLM-generated), detection techniques across pre-, mid-, and post-interaction phases, personalized response generation at prompt, fine-tuning, and architecture levels, and feedback-driven optimization through Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO). If the taxonomy is sound, it gives researchers and builders a shared map of a fragmented design space, turning \"make it personalized\" into a set of concrete, comparable design decisions.","feed_headline":"Survey maps how AI copilots learn, model, and adapt user preferences","feed_subtitle":"The survey organizes preference sources, detection techniques, and feedback loops into one design space for builders.","key_machinery":"The load-bearing object is the paper's conceptual architecture for a preference-aware AI copilot: user input flows into preference sources, then into detection techniques, then into personalized response generation, and finally into feedback-driven optimization, with the loop closing back into the model. The taxonomy splits detection into before (predefined profiles and persona development), during (real-time persona extraction and in-dialogue learning), and after (post-interaction feedback and preference refinement) the interaction. A second named device is the generalized DPO objective, $L(\\theta) = -\\log\\sigma(S(y_w, y_l, x)) + R(y_w, y_l, x)$, which the paper uses to unify DPO and its variants in one table; the scoring function $S$ measures preference between chosen and rejected responses and the regularizer $R$ encodes corrections such as overconfidence penalties or length normalization.","core_discovery":"The paper's discovery is that preference optimization can be treated as a system-level lifecycle rather than a single algorithmic step. On its own terms, the paper establishes a conceptual architecture in which user preferences flow from input channels into a preference representation, guide response generation, and then get refined through feedback; within that flow it identifies five sources of preference signals, three temporal phases of detection, three levels of response personalization, and two dominant post-training alignment families (RLHF and DPO). It also distills DPO variants into one generalized objective, $L(\\theta) = -\\log\\sigma(S(y_w, y_l, x)) + R(y_w, y_l, x)$, showing that the many published variants differ only in their scoring function $S$ and a regularization term $R$. The paper's claim is that this unified view is a literature-grounded definition and organizing framework for building user-aligned, persona-aware copilots.","pith_inferences":["A testable next step the authors do not run: use the taxonomy to build an evaluation protocol that measures whether copilots combining pre-, mid-, and post-interaction signals outperform single-phase systems on personalization benchmarks.","The taxonomy could be extended to multi-user and team settings, where preferences conflict and must be aggregated or negotiated, a setting the paper only touches via collective behavior modeling.","Because several load-bearing references in Section 4.2.3 are unverifiable as printed, the post-interaction feedback category should be regarded as provisional until those sources are confirmed; the taxonomy's other phases rest on standard, traceable literature.","One can map each preference source to a measurable cost/benefit trade-off (annotation cost, user fatigue, privacy risk, grounding) and use the paper's tables as a checklist for choosing signals under deployment constraints."],"forward_implications":["Designers can treat preference handling as a lifecycle, choosing sources and techniques per phase rather than bolting a single personalization method onto a chatbot.","The taxonomy gives a common vocabulary across recommender systems, human-AI interaction, and LLM alignment, so results from one subfield can be transferred to copilot design.","The unified DPO loss form makes variants (CPO, ORPO, SimPO, IRPO, beta-DPO, DPOC, MODPO) directly comparable, helping practitioners select an alignment method by scoring function and regularization.","The survey implies that robust personalization needs signals from all three phases, since pre-defined profiles alone are rigid and real-time extraction alone is content-bound.","Post-interaction feedback is cast as essential for long-term alignment, pointing to continual learning loops as a core copilot capability."],"supporting_citations":[{"why":"Supplies the DPO variant landscape and Table 6's structure, which Section 4.4.2 generalizes into the unified loss form.","marker":"[2]"},{"why":"Anchors the explicit feedback source category: pairwise human comparisons of model outputs for benchmarking.","marker":"[33]"},{"why":"Anchors the implicit feedback category: web usage mining to infer preferences from behavior.","marker":"[36]"},{"why":"Anchors LLM-generated signals: generative user simulation for scalable evaluation of personalized response generation.","marker":"[45]"},{"why":"Anchors conversational elicitation: coached preference elicitation showing richer signals from structured dialogue.","marker":"[50]"},{"why":"Anchors in-dialogue learning and the DPOC criterion-based post-interaction refinement example.","marker":"[61]"},{"why":"Supplies the AFSPP continual preference-shaping example in post-interaction refinement, though the printed entry is not verifiable.","marker":"[65]"},{"why":"Defines the RLHF post-training pipeline that Section 4.4.1 reviews.","marker":"[88]"},{"why":"Defines DPO, the base algorithm whose scoring/regularization form the survey generalizes.","marker":"[89]"}],"fun_headline_variants":["Taxonomy unifies AI copilot preference optimization","Preference optimization as a lifecycle in AI copilots","New framework for user-aligned AI copilots","Unifying DPO variants for copilot preferences","From signal to feedback: a copilot preference cycle"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole taxonomy stands on the cited literature existing and supporting the specific claims attached to it, and that grounding is visibly fragile in the post-interaction section, where several references carry incomplete identifiers or no verifiable authorship.","fun_headline_variants_meta":{"raw":{"variants":["Taxonomy unifies AI copilot preference optimization","Preference optimization as a lifecycle in AI copilots","New framework for user-aligned AI copilots","Unifying DPO variants for copilot preferences","From signal to feedback: a copilot preference cycle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000339,"raw_usage":{"total_tokens":1878,"prompt_tokens":956,"completion_tokens":922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":857}},"tokens_in":572,"tokens_out":922,"duration_ms":7685,"temperature":1.0,"reasoning_tokens":857,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:19:27.531609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Resolve every reference cited in Sections 4.1-4.4 and compare each attributed claim against the actual source; in particular, check whether [49], [63], [64], and [65] correspond to real published work and whether [64] is distinct from [30]. If the claims tied to those entries cannot be traced, the survey's evidential grounding for the post-interaction feedback category fails.","supporting_citations":[{"cited_title":"Evaluating large language models as generative user simulators for conversational recommendation","cited_arxiv_id":null,"evidence_quote":"Anchors LLM-generated signals: generative user simulation for scalable evaluation of personalized response generation."},{"cited_title":"Coached conversational preference elicitation: A case study in understanding movie preferences","cited_arxiv_id":null,"evidence_quote":"Anchors conversational elicitation: coached preference elicitation showing richer signals from structured dialogue."},{"cited_title":"Afspp: An agent framework for shaping preference and personality with llms","cited_arxiv_id":null,"evidence_quote":"Supplies the AFSPP continual preference-shaping example in post-interaction refinement, though the printed entry is not verifiable."}],"review_version":1}