{"id":"7c5fa6b4-4eed-47b5-ab11-02aa6f86b5e4","arxiv_id":"2505.09065","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A PRISMA-based survey proposing a unified lifecycle framework for HCI in explainable recommender systems, with video-based explanations flagged as an open research opportunity.","lead":"This paper is a survey that organizes how explainable recommender systems explain their recommendations to users, grouping current work into display content, display methods, and evaluation. A generalist should read it because it provides a common vocabulary for a fragmented field and argues that video-based explanations are an untapped direction.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey framework and first-time video claim rest on a PRISMA sample that cannot be reconstructed; missing search date, study list, and coding protocol leave the taxonomy's completeness unverifiable.","rationale":"I agree with the reader's weakest assumption. The paper has real strengths: a clear lifecycle framing, a useful mapping of examples to categories, and an honest discussion of evaluation gaps such as responsiveness and cognitive load. But the central contribution is the framework, and a framework's completeness is only as strong as the corpus from which it is induced. The reported PRISMA flow is not enough: PRISMA itself requires a date, a list of included and excluded studies, and reproducible screening. The missing pieces matter concretely for the novelty claim. If a pre-existing video-based or multimedia explanation paper is found in a re-run of the same search, the claim in Section 5.4 that no one has explicitly proposed such methods would be unsupported. I would keep the reader's conditional verdict: the survey is useful and should be published after the authors provide the full screening log, the study list, the search date, and a clarification of EC4 versus IC4, and after they either substantiate or qualify the 'first time' claim. No change in the verdict is needed.","tokens_in":26697,"tokens_out":7060,"duration_ms":73144,"concrete_test":"Request from the authors the full PRISMA screening log: exact search date(s), per-database query strings, hit counts after each stage, and the complete list of 102 included papers with their EC/IC codes. Independently re-execute the same query on the four databases up to the manuscript submission date and screen results using the stated criteria. If the reproduced set contains relevant papers not in the 102, especially any video- or multimedia-based explanation study that would change Section 5.4 or the display-content taxonomy, the 'first time' claim and the completeness of the unified framework are unsupported. As an internal check, verify whether IC4 admitted algorithm-only papers that EC4 should have excluded; if so, the inclusion process is not internally consistent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that the four-stage framework and the video-based-explanation direction are a faithful synthesis of the XRS HCI literature, rests on the completeness of the 102-paper sample in Section 2.2. That sample cannot be independently checked: no search date range is reported, the 102 included papers are never listed, and no coding protocol or inter-rater reliability is given. The 'first time' statement in Section 5.4 and the Conclusion ('no scholars have explicitly proposed video-based explainable recommendation methods') is therefore unfalsifiable from the paper as published, in a field where relevant multimedia work may have appeared before or after the unstated search date. There is also an internal tension in the screening rules: EC4 excludes papers 'focused solely on algorithms without addressing explanation presentation,' but IC4 admits 'classic papers on explainable recommendation algorithms and models'; without the screening log one cannot tell whether algorithm-only papers entered the corpus despite EC4. If the corpus is biased or incomplete, the display-content taxonomy (user, item, feature, logical, hybrid), the four display methods, and the evaluation categories are a plausible but unvalidated organizing scheme rather than a verified synthesis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a survey of the human-computer interaction (HCI) layer of explainable recommender systems (XRS). Following a PRISMA-style literature selection that reportedly yields 102 primary studies, the authors propose a four-stage lifecycle framework (data input, recommendation algorithms, explanation display, and evaluation) and organize the surveyed work along three axes: display content (user-, item-, feature-, logical-, and hybrid-based explanations), display methods (textual, visualization, hybrid-element, and multimedia explanations), and evaluation methods (quantitative and qualitative). The paper claims to be the first to highlight video-based explanations as a distinct research direction for XRS, and it offers a structured overview of evaluation criteria, including a table of user-study items based on Tintarev's criteria. It concludes with future-work directions covering multimedia explanations, evaluation standardization, and cognitive load.","tokens_in":26969,"tokens_out":3719,"duration_ms":36163,"significance":"If the synthesis is reliable, this survey is a useful contribution because it consolidates several previously fragmented taxonomies (Vig et al., Papadimitriou et al., Nunes and Jannach, Mohamed et al.) into one lifecycle-oriented framework, and it draws attention to multimedia and video-based explanations as an underexplored HCI direction. The evaluation overview, especially Table 2, provides a practical checklist for designing user studies of explanation interfaces. The paper is also candid about open evaluation gaps such as responsiveness and cognitive load. However, the scientific value of the survey depends heavily on the completeness and reproducibility of the 102-paper corpus, and the manuscript does not currently make that corpus verifiable. The descriptive claims about individual systems are generally consistent with the cited references, but the central framework's status as a faithful synthesis rests on methodological transparency that is not yet provided.","major_comments":[{"comment":"The PRISMA procedure as reported is not reproducible. The manuscript states that 489 articles were retrieved from four databases, 9 from other sources, 272 remained after duplicate removal, and 102 were ultimately included, but it reports no search date range, no per-database query adaptations, no list of the 102 included studies, and no coding protocol or inter-rater reliability for applying EC1-EC4 and IC1-IC5. Because the four-stage framework and the claimed novelty of video-based explanations are generalizations over this corpus, the completeness of the sample is load-bearing. Please provide the search date range, the exact queries used per database, an appendix listing the included studies, and a description of how the eligibility criteria were applied and by how many reviewers.","section":"Section 2.2"},{"comment":"The screening rules are internally in tension. EC4 excludes papers 'focused solely on algorithms without addressing explanation presentation,' while IC4 admits 'classic papers on explainable recommendation algorithms and models.' Without a screening log, it is impossible to determine whether algorithm-only papers entered the corpus through IC4, which could bias the reported display-content and display-method categories. Please clarify how the two criteria were reconciled in practice and report the number of papers admitted through each inclusion criterion.","section":"Section 2.2, EC4 vs. IC4"},{"comment":"The novelty claim about video-based explanations is unsupported as stated. The paragraph beginning 'Therefore, although no scholars have explicitly proposed video-based explainable recommendation methods...' appears twice verbatim, once after the discussion of reference [20] and again after Fig. 28, and the surrounding text reviews [19], [20], and [40] as video-related recommendation techniques. Either these works are not explanation methods and should be explicitly distinguished from explanation, or the claim that 'no scholars' have proposed such methods is contradicted by the cited literature. The claim also cannot be verified without the search date and study list requested above. Please replace the assertion with a precise, evidence-backed statement about what existing video-based work does and does not cover.","section":"Section 5.4 and Conclusion"}],"minor_comments":[{"comment":"The LinkedVis system is attributed to 'Stetlin et al.' but reference [85] is Bostandjiev et al.; the author name is incorrect and should be fixed.","section":"Section 4.1"},{"comment":"The text refers to 'the Fidelity formula, as Eq. (1)' but no equation is shown in the manuscript; the formula is missing and must be inserted.","section":"Section 6.1"},{"comment":"The paragraph beginning 'Therefore, although no scholars have explicitly proposed video-based explainable recommendation methods' is duplicated verbatim; one copy should be removed.","section":"Section 5.4"},{"comment":"The abbreviation EMF is introduced for 'Explicit Factor Model,' which is nonstandard and may be confused with Expectation-Maximization-based factorization; consider using a less overloaded abbreviation.","section":"Section 3.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's central contribution is a survey taxonomy, so the missing search date, missing list of included studies, and missing coding protocol are not cosmetic: they are what would allow a reader to verify that the framework is a faithful synthesis rather than an illustrative selection. The taxonomy itself is plausible and the evaluation overview is useful, so I would not reject the manuscript, but the authors should be required to supply the methodology appendix before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the four-stage lifecycle framework and the five-part display-content taxonomy are genuinely useful, and the paper does a fair job of comparing earlier taxonomies. The weak link is the 102-paper PRISMA sample: it cannot be reconstructed or audited, and the \"for the first time\" video claim is not falsifiable from the paper as published. But the central organizing scheme survives the methodological criticism—it is a plausible synthesis, not a verified one.\n\nWhat's good: the comparison of taxonomies in Section 3.3 is clear and the proposed framework—data input, recommendation algorithms, explanation display, evaluation—is a reasonable way to organize a messy field. The display-content categories (user-based, item-based, feature-based, logical-based, hybrid) map well onto the cited systems, and the evaluation section gives a genuinely useful structured list of qualitative criteria. The discussion of cognitive load and the future directions are concrete and honest. They also do not pretend to present new empirical results; it is framed as a review.\n\nWhere it wobbles:\n\n1. Section 2.2 reports PRISMA steps but no search date range, no list of included studies, and no coding protocol or inter-rater reliability. That makes the sampling claims unverifiable and the completeness of the taxonomy uncheckable. This is the real soft spot.\n\n2. Section 5.4 contains a duplicated paragraph verbatim—the \"Therefore, although no scholars...\" passage appears twice. Editorial sloppiness that needs fixing.\n\n3. The \"first time\" claim about video-based explanations is too strong. The paper itself cites video-related recommendation work (Guo et al., SEMI, Poet), and while no direct prior work explicitly targeting explanation in video-based recommendation is named, the claim cannot be tested without a date-bounded search. It should be softened to \"we did not find prior work explicitly targeting explanation in video-based recommendation.\"\n\n4. The inclusion/exclusion criteria tension (EC4 vs IC4) is real but minor; with the screening log it would be solvable.\n\nIs it a serious thinker? Yes. The paper is coherent on its own terms; the framework is not derived from a numerical fit, so there is no circularity. The self-citation to [76] is not a problem here.\n\nWho benefits: researchers and PhD students entering XRS HCI who want a map of display content, methods, and evaluation metrics. It will be cited more as a starting point than as a canonical systematic review.\n\nRecommendation: send it to peer review, but require the authors to add a supplementary study list, search date range, and coding protocol, and to fix the duplicated paragraph and soften the \"first time\" claim.","headline":"Useful taxonomy-driven survey with a wobbly literature sample; the 'first video explanations' claim is not established but the framework itself is worth engaging.","tokens_in":27455,"tokens_out":2245,"would_cite":true,"duration_ms":22865,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of explainable recommender systems puts the HCI layer into one four-stage lifecycle framework—data input, recommendation algorithms, explanation display, and evaluation—and argues that video-based explanations are a newly…","keywords":["explainable recommender systems","human-computer interaction","explanation display","display content taxonomy","display methods","video-based explanation","explanation evaluation","lifecycle framework"],"falsifier":"A broader search that includes non-English publications, industry interfaces, and demo systems could settle the taxonomy's completeness: if it finds an explanation display that fits none of the five content types or none of the four method types, the proposed framework is incomplete. Likewise, locating a published video-based explainable recommendation method would directly contradict the paper's novelty claim about video explanations.","tokens_in":26525,"feed_emoji":"🎬","tokens_out":9101,"duration_ms":82728,"temperature":0.7,"pith_summary":"Explainable recommender systems are usually surveyed from the algorithm side, leaving the ways explanations are actually shown to users scattered across inconsistent taxonomies. This paper tries to fix that by proposing a unified, lifecycle-scale framework with four stages: data input, recommendation algorithms, explanation display, and evaluation. Within that framework it organizes the HCI layer into display content (user-based, item-based, feature-based, logical-based, hybrid), display methods (text, visualization, hybrid-element, multimedia), and both qualitative and quantitative evaluation. It also argues that video-based explanations are a distinct and promising display direction that prior reviews have neglected. A reader should care because the framework promises a common vocabulary for designing, comparing, and testing explanations across different recommender systems.","feed_headline":"Survey unifies explainable recommender displays under one framework","feed_subtitle":"It classifies what to show, how to show it, and how to measure it, and names video explanations the next open direction.","key_machinery":"The central object is the four-stage lifecycle framework: data input, recommendation algorithms, explanation display, and evaluation. It does not replace algorithm taxonomies; it relocates them inside a process, so an explanation is described by what data feeds it, which algorithm generates it, how it is displayed, and how it is measured. The display-content taxonomy (user-, item-, feature-, logical-, hybrid-based) is carried by the intermediary-entity idea—the user, item, or feature that links the current user to the recommended item—while the display-method taxonomy is carried by presentation format: text, visualization, hybrid-element, and multimedia. The evaluation stage is carried by quantitative metrics such as explainability precision and recall, feature matching, coverage, diversity, fidelity, and confidence, together with qualitative criteria anchored in the seven classic explanation goals plus dimensions such as stability and cognitive load.","core_discovery":"On the paper's own terms, the central discovery is that the HCI layer of explainable recommender systems is not a jumble of interface tricks but a structured space, and that this space is best described as a process rather than a static taxonomy. The proposed four-stage framework traces explanation content from user, item, feature, review, and process data through the recommendation or explanation algorithm into a display stage and then an evaluation stage. Display content is classified into user-, item-, feature-, logical-, and hybrid-based explanations; display methods into textual, visualization, hybrid-element, and multimedia explanations; and evaluation into quantitative methods (online and offline, with item-based or feature-based metrics) and qualitative methods (user-perception criteria such as transparency, trust, effectiveness, persuasiveness, efficiency, satisfaction, and newer dimensions like stability). The paper further claims that multimedia, especially video, is an underused but viable explanation method supported by existing video summarization, highlight detection, and captioning techniques, and that responsiveness and cognitive load are the clearest gaps in current evaluation practice.","pith_inferences":["If display content and display method are truly independent axes, then video explanations could be attached to any content type; testing all cells of that content-by-method matrix would be a natural next design exercise.","The paper's offline metrics could be extended to video by scoring whether the temporal segment a video highlights corresponds to the item feature named in the explanation; the paper does not propose such temporal metrics.","Because the underlying sample is English-language and full-text only, a broader corpus including non-English work, industry interfaces, or demo tracks might reveal display formats that do not fit the four method classes; that is a testable completeness check rather than a settled conclusion.","Video explanations may have a double effect: richer information for novices but higher cognitive load; the paper flags cognitive load generally but does not tie it to user expertise, so an A/B study split by user expertise would fill that gap."],"forward_implications":["A designer can use the framework to choose a data input, an algorithm family, an explanation content type, a display format, and an evaluation dimension from one coherent map instead of from competing taxonomies.","Video-based explanations become a fourth display method, with video summarization, highlight detection, and captioning as concrete technical pathways and content availability, relevance, and scalability as the open problems.","Using the evaluation stage, a researcher can choose between offline metrics (explainability precision and recall, feature matching and coverage, fidelity) and user-perception criteria (transparency, trust, effectiveness, efficiency, satisfaction), which is the paper's recommended way to test an XRS.","Because overly detailed explanations can reintroduce the cognitive burden that recommender systems were meant to remove, future hybrid and multimedia explanation designs should prioritize relevant information and clear presentation, and should measure cognitive load directly.","Qualitative evaluation remains organized around the seven classic explanation criteria, but the paper identifies responsiveness and cognitive load as dimensions that still lack standardized measurement methods."],"supporting_citations":[{"why":"Supplies the algorithm-side overview and prior taxonomy that the proposed four-stage framework unifies.","marker":"[8]"},{"why":"Provides the explanation objectives, content, and presentation dimensions that the display and evaluation sections refine.","marker":"[15]"},{"why":"Gives the visual explanation format taxonomy and the comparison baseline for the display-method categories.","marker":"[16]"},{"why":"Introduces the human, item, feature, and hybrid explanation styles that the display-content taxonomy extends.","marker":"[18]"},{"why":"Defines the seven explanation evaluation criteria that anchor the qualitative evaluation survey.","marker":"[22]"},{"why":"Establishes the intermediary-entity view linking users, items, and features, which is the conceptual basis for the content categories.","marker":"[24]"},{"why":"Supplies a working micro-video recommendation system as evidence for the multimedia display direction.","marker":"[19]"},{"why":"Supplies product-oriented video captioning as a technical pathway for video-based explanations.","marker":"[20]"},{"why":"Supplies video highlight detection in e-commerce as evidence that video can explain product recommendations.","marker":"[40]"},{"why":"Provides the quantitative-versus-qualitative framing used to organize the evaluation methods section.","marker":"[98]"}],"fun_headline_variants":["Unified framework for explainable recommender UI displays","Lifecycle survey: how XRS explain, show, and test","Video explanations called the next step for recommender transparency","Mapping display content, methods, and evaluation in XRS","Survey exposes gaps in recommender explanation evaluation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's completeness rests on the assumption that the 102-paper sample covers the range of HCI display and evaluation work in explainable recommender systems; if relevant displays or evaluation studies were missed, the unified categories and the video-explanation novelty claim would be unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Unified framework for explainable recommender UI displays","Lifecycle survey: how XRS explain, show, and test","Video explanations called the next step for recommender transparency","Mapping display content, methods, and evaluation in XRS","Survey exposes gaps in recommender explanation evaluation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1421,"prompt_tokens":983,"completion_tokens":438,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":599,"tokens_out":438,"duration_ms":4806,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:40:29.188552+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A broader search that includes non-English publications, industry interfaces, and demo systems could settle the taxonomy's completeness: if it finds an explanation display that fits none of the five content types or none of the four method types, the proposed framework is incomplete. Likewise, locating a published video-based explainable recommendation method would directly contradict the paper's novelty claim about video explanations.","supporting_citations":[{"cited_title":"Tintarev and J","cited_arxiv_id":null,"evidence_quote":"Defines the seven explanation evaluation criteria that anchor the qualitative evaluation survey."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the intermediary-entity view linking users, items, and features, which is the conceptual basis for the content categories."},{"cited_title":"Q.Tan, J.Yu, Z","cited_arxiv_id":null,"evidence_quote":"Supplies product-oriented video captioning as a technical pathway for video-based explanations."},{"cited_title":"Yu, TaoHighlight: Commodity-Aware Multi-Modal Video Highlight Detection in E-Commerce, IEEE Trans","cited_arxiv_id":null,"evidence_quote":"Supplies video highlight detection in e-commerce as evidence that video can explain product recommendations."},{"cited_title":"Balog and F","cited_arxiv_id":null,"evidence_quote":"Provides the quantitative-versus-qualitative framing used to organize the evaluation methods section."}],"review_version":1}