{"id":"b8ce6d27-b6e1-4127-b30c-9d52628793df","arxiv_id":"2501.08087","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"App reviews expressing explanation needs can be routed to internal teams by a taxonomy-based filter and hierarchy, reaching 79.2% reported agreement in this one-company case study.","lead":"App-store reviews that need an explanation can be pre-filtered with word lists, sorted into taxonomy categories, and routed to the right internal team by a hierarchy built from support staff input. The paper reports 79.2% routing accuracy at one navigation company, but that figure comes from an evaluation that uses the same interviews used to build the hierarchy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 79.2% team-assignment figure is a resubstitution result: the same four support employees' assignments define both the hierarchy and the ground truth, and the paper concedes resolution accuracy was never evaluated.","rationale":"The reader's weakest-assumption analysis correctly identifies the load-bearing premise: the support employees' assignments serve simultaneously as training data for the team hierarchy and as ground truth for its evaluation, with no validation against whether the assigned team actually resolves the user's request. My read of the full text confirms this. The paper's own Section V.C explicitly states that the accuracy of team assignments in resolving explanation needs has not been evaluated, which is a direct admission of the missing external criterion. The low kappa values for team assignment (Table VII) further undermine the stability of the 'correct team' label. The 79.2% figure is best interpreted as a top-3 hit rate on the training set; the 52.5% top-1 figure is a more honest but still resubstituted number. Given the paper's framing as a single-company case study, a conditional verdict is appropriate: the descriptive findings and the taxonomy extension have value, but the headline accuracy claim should not be read as a validated performance measure. The proposed leave-one-out test is the minimal check that would determine whether the hierarchy generalizes to unseen assignments; if it passes, the remaining concern about resolution-based ground truth would still need to be addressed before calling 79.2% 'correct team assignment' in a meaningful sense. Therefore, I agree with the reader's conditional verdict and do not recommend changing it.","tokens_in":14244,"tokens_out":3544,"duration_ms":37777,"concrete_test":"Recompute the 79.2% average validity using leave-one-out cross-validation: for each of the 158 explanation needs, rebuild the taxonomy-category-to-team hierarchy from the other 157 needs (using the same 25% inclusion threshold) and test whether the held-out support-team assignment appears among the ranked teams; report top-1 and top-3 hit rates separately. If the leave-one-out top-3 figure is materially below 79.2%, the headline number is an artifact of evaluating on the training data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim, that hierarchical team assignment achieves 79.2% accuracy versus 52.5% for single-team assignment (Sections V.A and IV.E), rests on an evaluation that is circular with respect to the data. The team hierarchy in Section III.E is built from the same interview and survey responses provided by four Graphmasters support employees, and the 79.2% figure is then computed by checking whether those same responses fall within the ranked team list, counting second- and third-ranked teams as correct. This is a resubstitution metric, not an estimate of routing accuracy on unseen reviews. The problem is compounded by the low interrater reliability of the support-team assignments themselves (Cohen's kappa 0.146-0.558 in Table VII), which means the 'correct team' is not a stable target even among the four raters. The paper acknowledges in Section V.C that 'the accuracy of these team assignments in resolving the explanation needs has not been evaluated.' Without a held-out evaluation or an independent ground truth based on actual resolution outcomes, the 79.2% number only demonstrates that the hierarchy reproduces the raters' majority tendencies on the training data. A leave-one-out cross-validation would directly test whether the hierarchy generalizes beyond the very assignments used to create it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on a semi-automated system for managing explanation needs in app reviews at Graphmasters GmbH. The pipeline scrapes 2,366 Google Play and Apple App Store reviews, uses a word/phrase filter to detect explicit and implicit explanation needs, assigns reviews to taxonomy categories, and then maps categories to internal teams via a ranked hierarchy built from interviews and surveys with four support employees. External sources such as support articles and prior review responses are retrieved to answer the needs. The central reported results are that hierarchical team assignment achieves 79.2% validity (versus 52.5% for single-team assignment) and that 139 of 158 explanation needs were addressed by the company. The paper also reports low interrater agreement in taxonomy and team assignments, and several validity threats are discussed.","tokens_in":14478,"tokens_out":2990,"duration_ms":32972,"significance":"If the central figure were a true out-of-sample accuracy estimate, the paper would provide a useful industrial case study: it demonstrates a concrete taxonomy-to-team routing workflow, integrates real support articles and past review responses, and is honest about the challenges of interrater agreement. The authors also make their data collection process transparent, including interview guidelines and the use of a publicly documented taxonomy. However, the headline 79.2% team-assignment validity is computed on the same interview and survey data used to construct the team hierarchy, with a 25% inclusion threshold fit to those assignments, and the paper itself concedes that resolution accuracy was not evaluated. As presented, the quantitative claim does not measure routing quality for unseen reviews; it measures reproduction of the raters' majority tendencies on training data. The qualitative findings about taxonomy ambiguity and the practical difficulties of team assignment remain valuable, but the main accuracy claim needs a proper validation procedure before it can be accepted.","major_comments":[{"comment":"The 79.2% team-assignment validity is computed against the same interview and survey responses that were used to build the team hierarchy in Section III.E.1. This is a resubstitution metric: the hierarchy is fit to the raters' assignments and then evaluated on those same assignments, while counting second- and third-ranked teams as correct. It therefore does not estimate how well the hierarchy would route unseen reviews. A leave-one-out cross-validation over the four raters, or an independently held-out set of reviews, is required before the claim that 'the correct team could be assigned in 79.2% of cases' can be supported.","section":"IV.E and V.A"},{"comment":"The ground truth for 'correct team' is unstable across the four support employees: Table VII reports Fleiss' κ = 0.307 for reviews 1-25, Cohen's κ = 0.558 for reviews 26-50, and Cohen's κ = 0.146 for reviews 51-75. Because the 25% inclusion threshold in Section III.E.1 is itself fit to these same assignments, the resulting hierarchy and the validity estimate are jointly determined by the same noisy data. The paper should either justify the threshold with independent data or report a sensitivity analysis over the threshold and over subsets of raters.","section":"III.E.1 and Table VII"},{"comment":"The manuscript explicitly concedes that 'the accuracy of these team assignments in resolving the explanation needs has not been evaluated.' This directly limits the central claim: the 79.2% figure measures agreement with the raters' majority assignments, not whether the assigned team actually resolves the user's request. The abstract and RQ1 currently present the figure as routing accuracy, which overstates what the evaluation supports. Either the evaluation must be extended to resolution outcomes, or the claims should be reworded to describe rater-agreement, not correctness.","section":"V.C"},{"comment":"The dataset size is reported inconsistently: the text states 2,366 reviews were scraped, while Section III.C.4 later mentions 2,376 reviews in the manual labeling, and Table I sums to 2,365 reviews. The paper also reports 158 explanation needs from a total that changes by ten because of multi-need reviews. These numbers need to be reconciled, since the recall and precision figures in Section IV.A depend on a clearly defined denominator.","section":"III.C.3 and III.C.4"}],"minor_comments":[{"comment":"There is a typo in Section III.I: 'Cohne's kappa' should be 'Cohen's kappa'.","section":"III.I"},{"comment":"The row for 'nunav truck' appears to have missing or misaligned entries; the table should be reformatted so the counts for explicit, implicit, potential, and none are clear.","section":"Table I"},{"comment":"The phrase 'requriement engineer' in Section IV.C contains a typo; it should be 'requirements engineer'.","section":"IV.C"},{"comment":"The taxonomy description in Section II.A lists five main categories plus Meta Information and Timing, but Section V.B.1 adds Business and Feature Questions, and Table VIII contains additional categories such as Operation, Tutorial, and Consequences. The relationship between these category sets should be clarified, since the reader cannot tell which taxonomy was used for the final team assignments.","section":"II.A and Table VIII"},{"comment":"References [23] and [24] both attribute the same Landis and Koch work to different venues and years; one of them appears to be a duplicate with incorrect bibliographic details.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core issue is not the industrial case study itself but the validity of the headline accuracy figure. The resubstitution problem and the unstable ground truth are clearly visible in the paper's own tables and limitations section. I would ask the authors to reanalyze the data with a proper cross-validation scheme or to reframe the contribution as a qualitative case study without the 79.2% routing-accuracy claim. The paper is not beyond repair, but the central quantitative claim needs substantial work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nTwo things to know. First, this is the first real exercise of the Droste et al. explanation-need taxonomy inside a live company, and it adds the practical machinery of routing reviews to internal teams and to reusable sources. Second, the headline 79.2% team-assignment accuracy is a same-data resubstitution figure: the same four support employees' assignments were used both to build the team hierarchy and to serve as ground truth for the evaluation. The paper itself concedes in Section V.C that the accuracy of these team assignments in resolving the explanation needs has not been evaluated. So treat 79.2% as internal consistency, not predicted routing quality.\n\nWhat the paper does well: it is a transparent single-company case study. The taxonomy extension (Business, Meta Information, Feature Questions) is a genuine practical addition. The reporting of word-filter precision/recall trade-offs is honest, and the observation that 13 of 15 newly written responses were for Apple App Store reviews gives practitioners a concrete platform asymmetry to work with. The limitations section covers the small sample, the single-company scope, and the lack of longitudinal data.\n\nWhere the soft spots land: the core accuracy claim is circular. A leave-one-out evaluation would have tested whether the hierarchy generalizes; they did not run one. Interrater agreement for team assignment is low (Cohen's kappa 0.146–0.558), so the 'correct team' is not a stable target even among the four raters. No code or data is shipped, which limits independent verification, though that is a moderate issue for a case study. If the claims were reframed as 'our process reproduces the support team's majority assignments on the training data,' the paper would be solid.\n\nWho should read it: researchers working on app review analysis, requirements engineers interested in explainability in practice, and people designing support tooling. I would send it to a serious referee, not desk reject. The referee should demand that the evaluation claims be scaled back to internal consistency or supported by held-out validation.\n\nMy bottom line: worth reading and worth citing for the taxonomy extension and the industrial process, but do not quote the 79.2% as an accuracy result.","headline":"Industrial deployment of the explanation-need taxonomy with a useful routing process, but the headline 79.2% team-assignment figure is resubstitution, so read the paper for the process and not for the accuracy claim.","tokens_in":15031,"tokens_out":3794,"would_cite":true,"duration_ms":34581,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Team ranking lifts app-review routing to 79.2% correct","keywords":["explainability","app reviews","explanation needs","taxonomy","team assignment","requirements engineering","survey study","navigation app"],"falsifier":"Take a fresh batch of explanation-needy reviews from the same company, have the support teams actually work the tickets, and record which team resolves each request; if the hierarchy's top-ranked team matches the resolving team no more than 52.5% of the time, the 79.2% figure overstates routing quality.","tokens_in":14032,"feed_emoji":"📱","tokens_out":7828,"duration_ms":66278,"temperature":0.7,"pith_summary":"The paper proposes a semi-automated pipeline that turns app-store reviews containing explanation needs into a taxonomy category, a ranked list of internal teams that could answer, and a source that already contains the answer. Working with a navigation-app company and 2,366 scraped reviews, the authors find that ranking several candidate teams instead of picking one raises the chance of a correct assignment from 52.5% to 79.2%. The company was able to address 139 of 158 explanation needs, and only 15 required drafting a new response. If the results hold, companies can make explainability a managed requirement rather than a repetitive manual chore.","feed_headline":"Team ranking lifts app-review routing to 79.2% correct","feed_subtitle":"A taxonomy-plus-ranking pipeline cuts repetitive support work and meets 88% of user explanation needs from existing content.","key_machinery":"The central machinery is an extended taxonomy of explanation needs plus a hierarchical team-assignment rule. The taxonomy, taken from Droste et al., sorts reviews into categories such as System Behavior, Interaction, Privacy & Security, Domain Knowledge, and User Interface, and is extended here with Business, Meta Information, and Feature Questions to capture needs that otherwise fall between categories. A word-and-phrase filter built from Obaidi's 245 phrases and Droste's trigger words flags candidate reviews, and two requirements engineers verify the labels as ground truth. For each taxonomy category, the system lists internal teams that received at least 25% of the support staff's assignments, ranked by frequency; this ranked list is the 'reference point.' A separate script compares the review text against support articles using Python's difflib SequenceMatcher and against past Google Play responses to identify a 'source.'","core_discovery":"The central claim is that a taxonomy-to-team mapping, built from the judgments of four support-staff members, lets a company route explanation-needy reviews to a ranked list of internal teams with 79.2% cumulative accuracy. The pipeline first flags candidate reviews with a word-and-phrase filter, then human requirements engineers verify the labels to produce ground truth. For each taxonomy category, any team that received at least 25% of the support staff's assignments is placed in a hierarchy, ranked by frequency, and the top three options count as a correct assignment. Using only the first-ranked team yields 52.5% accuracy, so the ranking adds 26.7 percentage points. A separate script finds a source—a support article, a past Google Play response, or a newly drafted answer—and the company reports it can answer 88% of the identified explanation needs.","pith_inferences":["The 79.2% accuracy is measured against the support staff's own assignments, not against whether the assigned team actually resolves the request; the paper itself flags this, so a pilot deployment with ticket outcomes is the natural next test.","The 25% threshold likely needs recalibration as team sizes and responsibilities change, but the ranked-list idea should transfer to any organization that has a ticket taxonomy and a support team.","The pipeline's practical value depends on the cost of the human verification step, which the paper does not quantify; a cost-benefit estimate would show whether the automation pays for itself.","Feeding the ranked team list into a lightweight classifier or language model could test whether the manual verification step can be removed once more labeled data accumulates."],"forward_implications":["If the 79.2% accuracy reproduces, companies can auto-route reviews to a short ranked list and have a human only confirm the top choice, cutting repetitive work.","The 88% addressability figure suggests most user explanation needs are recurring and can be served from existing support articles or past responses, not fresh writing.","The 25% inclusion threshold is a simple data-driven rule for turning a small set of expert labels into a team-assignment hierarchy, no machine learning required.","The taxonomy extensions (Business, Meta Information, Feature Questions) provide a template for other companies whose reviews ask about the provider or span multiple categories.","Apple App Store reviews are harder to source because past review responses are not available, so the pipeline works best where a reply history exists."],"supporting_citations":[{"why":"supplies the taxonomy categories that the whole team-assignment pipeline is built on","marker":"[3]"},{"why":"provides the 245 phrases powering the explanation-need filter","marker":"[20]"},{"why":"supplies the trigger words and labeling guidelines for implicit and explicit needs","marker":"[19]"},{"why":"documents prior NLP detection of explanation needs, the baseline this rule-based approach builds on","marker":"[14]"},{"why":"defines explainability in terms of support articles and support team answers, motivating the reference-point/source split","marker":"[11]"}],"fun_headline_variants":["Ranked team picks lift app-review routing to 79.2% accuracy","Taxonomy-based routing reaches 79.2% team accuracy","Ranking teams boosts app-review routing from 52.5% to 79.2%","Semi-automated review routing hits 79.2% team-match rate","Existing content meets 88% of app-review explanation needs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the assignments made by four support employees in interviews and surveys define the 'correct team' and can serve as both the training data and the ground truth for the evaluation; the paper never checks whether the assigned team actually resolves the user's request.","fun_headline_variants_meta":{"raw":{"variants":["Ranked team picks lift app-review routing to 79.2% accuracy","Taxonomy-based routing reaches 79.2% team accuracy","Ranking teams boosts app-review routing from 52.5% to 79.2%","Semi-automated review routing hits 79.2% team-match rate","Existing content meets 88% of app-review explanation needs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000704,"raw_usage":{"total_tokens":3175,"prompt_tokens":947,"completion_tokens":2228,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2129}},"tokens_in":563,"tokens_out":2228,"duration_ms":15584,"temperature":1.0,"reasoning_tokens":2129,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:29:48.586403+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh batch of explanation-needy reviews from the same company, have the support teams actually work the tickets, and record which team resolves each request; if the hierarchy's top-ranked team matches the resolving team no more than 52.5% of the time, the 79.2% figure overstates routing quality.","supporting_citations":[{"cited_title":"While the primary focus was on the Nunav Navigation App , three additional apps were included to expand the dataset due to their structural similarity","cited_arxiv_id":null,"evidence_quote":"supplies the taxonomy categories that the whole team-assignment pipeline is built on"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the 245 phrases powering the explanation-need filter"},{"cited_title":"[3], as was also discussed in their original publication","cited_arxiv_id":null,"evidence_quote":"supplies the trigger words and labeling guidelines for implicit and explicit needs"},{"cited_title":"UI/UX 55%","cited_arxiv_id":null,"evidence_quote":"documents prior NLP detection of explanation needs, the baseline this rule-based approach builds on"},{"cited_title":"Support 35%","cited_arxiv_id":null,"evidence_quote":"defines explainability in terms of support articles and support team answers, motivating the reference-point/source split"}],"review_version":1}