{"id":"4c081337-d030-4869-b462-58a2de415128","arxiv_id":"2507.13725","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A reflection paper argues that POI recommendation research is held back by 20 systemic pitfalls in datasets, algorithms, and evaluation, and outlines six solutions.","lead":"This paper identifies 20 recurring pitfalls in how point-of-interest recommender systems are designed, tested, and evaluated, and proposes a six-part research agenda to fix them. It matters because these systems guide tourists' high-stakes travel decisions, yet the research behind them is often built on outdated, biased data and unrealistic evaluations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'majority have failed' claim in Section 8 is an unquantified empirical generalization; the narrative selection of 20 pitfalls and subjective Table 1 mapping do not establish that a majority of proposed solutions fail to address stakeholder needs.","rationale":"The paper is a useful and transparent reflection piece, and the reader's ACCEPT is defensible within the position-paper genre. However, the single most load-bearing element is the conclusion's empirical claim that a majority of proposed POI recommendation solutions have failed to address stakeholder needs. Everything else—the 20 pitfalls, Table 1, the six-part agenda—is diagnostic scaffolding for that claim. If the claim overstates the failure rate, the agenda may be aimed at a distorted picture of the field. The reader's weakest_assumption focused on the selection of pitfalls; my concern is adjacent but sharper: even if the 20 pitfalls were representative, the paper provides no operational definition of 'failed to address important needs' and no systematic measurement of the proportion of solutions that fail. The only quantitative support (13/310 releasing code) targets one narrow pitfall. Thus the central claim is an assertion, not a finding. Because this is a position paper, the honest fix is inexpensive: either soften 'majority' to 'many' or 'a substantial portion,' or add a small systematic review coding a sample of papers against the pitfalls. I recommend conditional acceptance rather than rejection, as the synthesis and agenda remain valuable and the concern is about calibration of the headline claim, not about internal inconsistency. This concern does not change the reader's assessment of the paper's overall utility, but it should change the verdict to CONDITIONAL pending revision of the quantifier or addition of evidence.","tokens_in":24695,"tokens_out":5887,"duration_ms":64249,"concrete_test":"Select a stratified random sample of 100 POI recommendation papers from the 310 proposals surveyed in [95] (or from RecSys/SIGIR/UMAP 2015-2024); define an operational rubric mapping the 20 pitfalls to binary indicators (e.g., uses post-2015 data, evaluates non-accuracy metrics, considers non-user stakeholders, releases code); have two independent coders score each paper; compute the proportion failing on at least one critical stakeholder-need dimension. If the proportion is not >50%, the 'majority' claim in Section 8 is unsupported; if it is >50%, the claim gains direct evidence.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Section 8) that 'the majority of the proposed solutions have failed to address the important needs and demands of the involved stakeholders' is a quantitative statement about the field, but the paper never measures it. The support consists of 20 pitfalls, which Section 1 states are 'not comprehensive' and selected 'based on our review of the literature and discussions.' Each pitfall is illustrated with selected references, and Table 1 maps pitfalls to solutions using the authors' primary/secondary judgments, with no data on how many published systems actually exhibit each pitfall. The only concrete statistic is Pitfall 20: 13 of 310 surveyed proposals released code [95], which addresses reproducibility only and does not imply failure on stakeholder needs. Consequently, the 'majority' quantifier exceeds what the narrative evidence can support; the paper demonstrates that many pitfalls are common, not that most solutions fail.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This reflection paper argues that POI recommendation research, despite its volume, remains poorly suited for real-world deployment. The authors identify 20 pitfalls across three dimensions—datasets (Pitfalls 1–10), algorithms (Pitfalls 11–15), and evaluation (Pitfalls 16–20)—and propose a six-part research agenda: multistakeholder design, context-awareness, data collection, trustworthiness design, novel interactions, and real-world evaluations. The paper reviews the state of the art in POI data (LBSNs, Flickr, mobility traces, user studies, synthetic data), algorithmic approaches (classical, neural, LLM-based, reinforcement learning, operations research), and evaluation practices. It closes with Table 1 mapping pitfalls to primary and secondary solutions, and a conclusion asserting that the majority of proposed POI solutions have failed to address stakeholder needs.","tokens_in":24801,"tokens_out":3747,"duration_ms":43977,"significance":"If treated as a position statement, the paper is useful and timely. It consolidates scattered concerns about LBSN data biases, MNAR issues, popularity bias, reproducibility, multistakeholder trade-offs, and the limitations of offline evaluation into a single structured catalog, and it translates these into a concrete research agenda with named solution directions. The paper also contains at least one concrete quantitative anchor: the 13-out-of-310 code-availability figure for POI recommendation papers, taken from [95]. Its main value is as a community reference for what has gone wrong and what could be done next, rather than as a new technical result. However, because the central claim is a quantitative generalization about the field, the evidentiary basis for that generalization must be examined carefully; this is where the manuscript currently needs strengthening.","major_comments":[{"comment":"The concluding claim that “the majority of the proposed solutions have failed to address the important needs and demands of the involved stakeholders” is not supported by the evidence presented in the paper. The 20 pitfalls are introduced in Section 1 as selected “based on our review of the literature and discussions,” not from a systematic survey. Table 1 encodes the authors’ qualitative primary/secondary judgments, and the only concrete prevalence statistic in the paper (Pitfall 20: 13 of 310 proposals released code, citing [95]) concerns reproducibility rather than stakeholder need satisfaction. The text therefore demonstrates that many pitfalls are common, but it does not establish that a majority of proposed solutions fail. Either soften the quantifier (e.g., “many” or “a substantial share”) or provide a systematic evidence base, such as a coded sample of recent publications indicating the frequency of each pitfall.","section":"8 (Lessons Learned and Conclusions)"},{"comment":"The research agenda’s plausibility rests in part on Table 1’s assignment of each pitfall to primary and secondary solutions, but no criteria for these assignments are given. For example, Pitfall 3 (Incomplete Data) is assigned Data-Collection as primary and Context-Awareness as secondary, while one could reasonably argue that richer evaluation protocols or data augmentation are equally primary. Section 8’s statement that Real-World Evaluations “emerges as the most important area of improvement” is a counting result over these subjective assignments. The authors should either state explicit coding rules and ideally report inter-coder agreement, or explicitly present Table 1 as an informed opinion rather than a derived, reproducible result.","section":"7 and Table 1"}],"minor_comments":[{"comment":"The first sentence contains a typo: “he incompleteness of LBSN datasets” should read “The incompleteness of LBSN datasets.”","section":"5, Pitfall 11"},{"comment":"The phrase “undermining the system’s real-world applicability of the system” is redundant; consider “undermining the system’s real-world applicability.”","section":"4.1, Pitfall 1"},{"comment":"The column heading “Real-World Evalua- tions” is hyphenated across a line break; use “Real-World Evaluations” consistently.","section":"Table 1"},{"comment":"The paper carefully disclaims that the pitfall list is not comprehensive, which is good; the same caution should be applied to the framing that the six solutions are “viable” or sufficient, since the manuscript does not demonstrate that the proposed agenda, if followed, would resolve the identified pitfalls.","section":"1 and 8"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reflection/synthesis piece rather than a novel technical contribution, and its fit will depend on whether the venue welcomes such works. I would also note that several load-bearing examples are drawn from the authors’ own prior work (e.g., [94], [95], [97]); this is acceptable in a synthesis, but the editors may want to ensure that independent evidence is proportionately represented. The main revision needed is the calibration of the “majority” claim and the transparency of the Table 1 mapping methodology; neither requires new experiments if the claims are softened appropriately."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid reflection paper that compiles 20 known pitfalls across datasets, algorithms, and evaluation, maps them to six solution directions, and packages it in a way that is genuinely usable. The novelty is organizational rather than empirical, which is fine for a position paper. If you work in POI recommendation, you will want a copy on your desk.\n\nThe 20-pitfall structure is the real asset. Each pitfall is concrete, cited, and illustrated with examples, and Table 1 links pitfalls to primary and secondary solutions. That table is a good starting point for a research agenda and for teaching. The paper is also honest about its own scope: it says the list is not comprehensive and relies on literature review and discussions. The six solutions are not new individually, but the synthesis and the explicit mapping are useful.\n\nThe soft spot is the Section 8 claim that 'the majority of the proposed solutions have failed to address the important needs and demands of the involved stakeholders.' That is an empirical generalization, and the paper never measures it. The support is the 20 pitfalls, each illustrated by selected references, but there is no systematic count of how many published systems exhibit each pitfall. The only concrete statistic is the 13/310 code-release figure from the authors' own earlier survey, and that is about reproducibility, not stakeholder needs. So the 'majority' quantifier exceeds the evidence. It is a strong opinion in a position paper, which is defensible, but it should be phrased as an opinion rather than a finding, and the phrase 'we have shown' elsewhere is similarly a bit strong.\n\nSelf-citation is frequent, but it is concentrated where it is earned: their 2022 survey [95] is the natural source for the code-availability statistic and the density figures. I do not see a circularity problem.\n\nOverall: for a reader who wants a structured critique of where POI recommendation research has gone wrong, and a roadmap for better evaluation and design, this paper is valuable. It deserves peer review and publication as a reflection paper. My main advice to the authors would be to soften or rephrase the 'majority' claim.","headline":"A genuinely useful, well-organized reflection paper that will serve as a practical checklist for POI researchers, but the central claim that 'the majority' of solutions have failed is an opinion, not a measured finding.","tokens_in":25358,"tokens_out":2071,"would_cite":true,"duration_ms":23896,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that point-of-interest recommendation research has failed its stakeholders because of 20 recurring pitfalls in datasets, algorithms, and evaluation, and that a six-part agenda—multistakeholder design, context awareness…","keywords":["point-of-interest recommendation","tourist recommender systems","evaluation pitfalls","multistakeholder recommender systems","location-based social networks","context-aware recommendation","reproducibility","tourism data"],"falsifier":"A systematic coding of the POI recommendation literature published since 2011, using a transparent protocol on the same 310 papers the authors rely on, would settle the prevalence claim: if a majority of those papers already avoid most of the 20 pitfalls—for example, by using post-2020 data, user studies, or multistakeholder metrics—the paper's central diagnosis would be false.","tokens_in":24466,"feed_emoji":"🗺️","tokens_out":7798,"duration_ms":76433,"temperature":0.7,"pith_summary":"The paper is a critical review of point-of-interest (POI) recommendation in tourism. It argues that after years of research, most proposed systems still fail the people and organizations they are meant to serve, because three parts of the research pipeline—datasets, algorithms, and evaluation—carry 20 systematic pitfalls. These include outdated and biased location-based social network data, accuracy-obsessed models that ignore context and other stakeholders, and offline evaluations that measure the wrong things. The authors' positive claim is that a structured agenda built around six directions—multistakeholder design, context-awareness, data collection, trustworthiness, novel interactions, and real-world evaluation—can turn the field toward systems that work outside the laboratory. A sympathetic reader should care because tourism recommendation is high-stakes: users spend time, money, and effort on the suggestions, and destinations face overcrowding and sustainability problems that current systems can worsen.","feed_headline":"20 pitfalls block point-of-interest recommenders from real use","feed_subtitle":"A six-part agenda—multistakeholder design, context, trust, data, interaction, real-world tests—is the proposed way out.","key_machinery":"The load-bearing object is the catalogue of 20 pitfalls, organized along three dimensions (datasets, algorithms, evaluation) and mapped to six solutions in a pitfall–solution table, with primary and secondary matches. The catalogue does the work of structuring the diagnosis, and the table does the work of showing that the solutions are neither arbitrary nor independent: for example, real-world evaluation is the primary remedy for six pitfalls and a secondary remedy for seven, while context-awareness is the only solution that does not touch any evaluation pitfall. The multistakeholder framing is the conceptual motor: it defines POI recommendation as a balancing act among tourists, destination managers, local businesses, and communities, which makes accuracy metrics an obviously insufficient target and turns the six solutions into a coherent agenda.","core_discovery":"On the paper's own terms, the central discovery is a diagnosis and a route forward: POI recommender research has largely optimized accuracy on LBSN check-in histories, and in doing so has built systems that are outdated, biased, sparsely populated, unmindful of tourists' real contexts and group behavior, ignorant of the interests of destination managers and local communities, and impossible to compare or reproduce. The authors identify 20 pitfalls across the data, algorithmic, and evaluation dimensions and claim these are the most prevalent and critical blockers of real-world applicability. They then claim that a coordinated research agenda, with real-world evaluation as the keystone, is a viable solution: multistakeholder design addresses whose goals count, context-awareness handles in-trip re-planning, data collection replaces stale biased logs, trustworthiness designs make recommendations credible, novel interactions integrate trip components, and real-world evaluation tests systems in living labs and calibrated simulations. The conclusion they draw is that the majority of proposed solutions have failed to address the important needs and demands of the involved stakeholders.","pith_inferences":["A likely consequence the authors do not spell out is that the next shared POI benchmark should incorporate post-COVID mobility data and stakeholder outcomes such as crowding reduction and small-business visibility, not just ranking accuracy.","The pitfall–solution mapping suggests a testable extension: because context-awareness is the only solution that does not address any evaluation pitfall, context-rich models tested with traditional offline protocols may show limited gains unless the evaluation is reformed at the same time.","Regulatory pressure—GDPR's consent requirements and the Digital Services Act's transparency duties—may push the data-collection and trustworthiness agenda faster than internal research incentives alone, making the proposed directions increasingly hard to ignore.","A concrete way to test the central claim would be a living-lab comparison of a multistakeholder, context-aware, explainable recommender against an accuracy-optimized baseline, measuring user satisfaction, crowding, and coverage of lesser-known venues; the authors' diagnosis predicts the former wins on those metrics."],"forward_implications":["If the diagnosis is correct, accuracy-focused offline benchmarks on Foursquare and Gowalla cannot serve as evidence that a POI recommender will work in practice; they need to be supplemented or replaced by newer, complete, stakeholder-aware data.","Algorithms that optimize only precision and recall will keep producing popular-biased, context-blind suggestions, so future models must adopt multi-objective targets that include diversity, fairness, and stakeholder interests.","Real-world evaluation—through living labs and calibrated simulations—becomes the primary filter for selecting candidate systems, with offline experiments used mainly to shortlist them.","The pitfall-solution mapping implies that no single fix is enough: data collection reforms change what algorithms can learn, and trustworthiness constraints shape novel interactions, so the agenda must be pursued as a coordinated program.","Reproducibility will require changes in publication norms—sharing code, data splits, and evaluation protocols—alongside new methods, since the field's own survey found that only 13 of 310 proposals published their code."],"supporting_citations":[{"why":"The experimental survey of POI recommender systems that supplies the density figures for Foursquare and Gowalla and documents the 310-proposal corpus used to substantiate the reproducibility pitfall.","marker":"[95]"},{"why":"The handbook chapter that defines multistakeholder recommender systems and grounds the framing that POI recommendation must balance conflicting stakeholder objectives.","marker":"[1]"},{"why":"The LBSN survey that establishes the age, sparsity, and incompleteness of the common check-in datasets behind the dataset pitfalls.","marker":"[14]"},{"why":"The simulation study showing that a multistakeholder recommender can improve both user utility and the dispersal of tourists, used as evidence that the first solution is viable.","marker":"[71]"},{"why":"The empirical analysis of Foursquare check-ins that documents weekday/weekend and resident-like temporal patterns behind the training-data mismatch pitfall.","marker":"[75]"},{"why":"The study of aggregation strategies showing that most Foursquare users check in within a single city, supporting the claim that LBSN data is biased toward locals.","marker":"[94]"},{"why":"The travelers-versus-locals cluster analysis used to justify the segmentation pitfall and the call for user and item segmentation in evaluation.","marker":"[97]"},{"why":"The reproducibility analysis of recommender systems research that grounds the reproducibility pitfall and the argument for sharing code and protocols.","marker":"[25]"},{"why":"The description of the YFCC100M photo dataset used to illustrate platform-specific temporal and geospatial biases in photo-sharing data.","marker":"[107]"}],"fun_headline_variants":["POI recommenders hit 20 pitfalls; six-step agenda offers fix","POI recs flawed: 20 pitfalls, six research directions to fix","20 pitfalls block POI recommenders; agenda paves way out","POI rec research: 20 known pitfalls, six-part fix agenda","Tourism recs fail: 20 pitfalls, six research priorities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that these 20 pitfalls are the most prevalent and critical rests on the authors' narrative reading of the literature, not on a systematic, transparent selection methodology; if that selection is skewed, the research agenda could be aimed at the wrong problems.","fun_headline_variants_meta":{"raw":{"variants":["POI recommenders hit 20 pitfalls; six-step agenda offers fix","POI recs flawed: 20 pitfalls, six research directions to fix","20 pitfalls block POI recommenders; agenda paves way out","POI rec research: 20 known pitfalls, six-part fix agenda","Tourism recs fail: 20 pitfalls, six research priorities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3106,"prompt_tokens":985,"completion_tokens":2121,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":2026}},"tokens_in":601,"tokens_out":2121,"duration_ms":18484,"temperature":1.0,"reasoning_tokens":2026,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:17:00.988170+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic coding of the POI recommendation literature published since 2011, using a transparent protocol on the same 310 papers the authors rely on, would settle the prevalence claim: if a majority of those papers already avoid most of the 20 pitfalls—for example, by using post-2020 data, user studies, or multistakeholder metrics—the paper's central diagnosis would be false.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The handbook chapter that defines multistakeholder recommender systems and grounds the framing that POI recommendation must balance conflicting stakeholder objectives."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The LBSN survey that establishes the age, sparsity, and incompleteness of the common check-in datasets behind the dataset pitfalls."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The empirical analysis of Foursquare check-ins that documents weekday/weekend and resident-like temporal patterns behind the training-data mismatch pitfall."}],"review_version":1}