{"id":"3a0e2a8d-70d1-49b7-9939-073b6341b619","arxiv_id":"2508.05634","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A crowd navigation method augmenting reinforcement learning with conformal uncertainty estimates is claimed to cut collisions under distribution shift, but the manuscript body is an unrelated live streaming dataset paper.","lead":"This preprint's abstract describes a robot crowd navigation method that uses conformal prediction uncertainty estimates to guide a reinforcement learning policy. The paper's body contains a completely different paper about live streaming recommendation, so the navigation claims could not be checked.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manuscript body is an unrelated live-streaming dataset paper; the crowd-navigation claims are entirely unsupported, so the paper cannot be verified.","rationale":"The reader's verdict is UNVERDICTED, and my stress-test supports that verdict, so the verdict should not change. I identify the structural mismatch between abstract and full text as the single most load-bearing concern: the quantitative claims in the abstract have no corresponding method or experimental section in the submitted document. This is not a scientific disagreement with consensus; it is an internal inconsistency that prevents any assessment of correctness. The reader's primary weakest_assumption concerned whether conformal uncertainty estimates remain calibrated when used as constraints inside a reinforcement learning loop. That is an important substantive concern, but it presupposes the actual crowd-navigation paper exists in the submission. Since the body is an unrelated live-streaming dataset paper, that substantive concern cannot even be analyzed. I therefore partially agree with the reader: they flagged the structural mismatch as a secondary issue, but my load-bearing concern is that mismatch itself. If the correct PDF were supplied, the conformal-in-the-loop question would be the natural next target. The concrete test is deliberately minimal: verifying the document's identity settles the current concern decisively. No ad hominem is intended; an upload error is a plausible and charitable explanation, but it does not change the epistemic status of this submitted text.","tokens_in":11358,"tokens_out":3813,"duration_ms":39174,"concrete_test":"Retrieve the PDF for arXiv:2508.05634 and programmatically compare the title/abstract with the full-text heading structure, specifically checking for a crowd-navigation method section, an experiments/results section reporting the claimed metrics, and the gen-safe-nav URL. If the body is the KuaiLive dataset paper and contains none of these, the mismatch is confirmed; the paper should remain UNVERDICTED (or be returned for resubmission of the correct file).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The full text of arXiv:2508.05634 is a SIGIR '26 paper titled 'KuaiLive: A Real-time Interactive Dataset for Live Streaming Recommendation' by Qu, Dai, Guo, et al. The abstract supplied for the submission describes a crowd navigation method using adaptive conformal inference and constrained RL, with quantitative claims: 96.93% success rate, >8.80% over SOTA, 3.72x fewer collisions, 2.43x fewer intrusions, and real-robot deployment. None of this appears in the body: Sections 1-7 describe live streaming data collection, dataset statistics (Table 2), and recommendation tasks (top-K, CTR, watch time, gift price). There is no method section for the navigation policy, no conformal uncertainty derivation, no simulation setup, no baseline table, no OOD-shift experiment, and no gen-safe-nav website. The body even identifies itself as a dataset paper. Thus the central claim is supported only by the abstract; there is no experiment, derivation, or algorithm to check. This is a structural mismatch between the abstract and the document, making every load-bearing premise (calibrated uncertainty under policy feedback, OOD robustness, robot safety) unverifiable. I am not alleging intent; an upload or metadata mix-up could explain it. But as a submitted manuscript, the claim cannot be assessed on the evidence in hand.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission, arXiv:2508.05634 (cs.RO), is presented with an abstract and title claiming a crowd navigation method that uses adaptive conformal inference to produce uncertainty estimates, which are then fed into a constrained reinforcement learning policy to improve safety and out-of-distribution robustness. The abstract reports specific quantitative results (96.93% success, >8.80% over state-of-the-art, 3.72x fewer collisions, 2.43x fewer intrusions) and mentions real-robot deployment. However, the manuscript body is an entirely different paper: 'KuaiLive: A Real-time Interactive Dataset for Live Streaming Recommendation,' a SIGIR '26 dataset paper about a live-streaming recommendation dataset. The body contains no crowd navigation, no conformal inference, no reinforcement learning, no simulation experiments, and no robot deployment. None of the abstract's claims are supported by the submitted text.","tokens_in":11700,"tokens_out":1940,"duration_ms":20647,"significance":"If the claimed method and results were actually present and correct, they would represent a meaningful contribution to safe crowd navigation, particularly the use of conformal uncertainty estimates within constrained RL to handle distribution shifts. The quantitative improvements over state-of-the-art baselines, if reproducible, would be practically relevant. However, the manuscript as submitted does not contain the claimed work at all. There is no method, no experiment, no derivation, and no result to evaluate. The significance of the submission is therefore entirely unrealized; the paper cannot be assessed on its merits because the content is missing.","major_comments":[{"comment":"The manuscript body is the KuaiLive dataset paper on live-streaming recommendation. Sections 1, 3, and 7 describe a dataset from Kuaishou, with user/streamer/room statistics, data analysis, and recommendation tasks. There is no mention of crowd navigation, pedestrians, conformal prediction, constrained reinforcement learning, or robots. The abstract's claimed contribution is completely absent from the body, making the central claim unsupported by the submitted document.","section":"Abstract vs. Sections 1-7"},{"comment":"The quantitative claims (96.93% success rate, >8.80% over baselines, 3.72x fewer collisions, 2.43x fewer intrusions) appear only in the abstract. There is no experimental protocol, no simulator description, no baseline table, no error bars, and no statistical analysis anywhere in the paper. These numbers cannot be verified or meaningfully reviewed. This is a load-bearing issue: the paper's stated findings have no supporting evidence in the manuscript.","section":"Abstract (quantitative claims)"},{"comment":"The page header identifies the document as 'arXiv:2508.05633v2 [cs.IR]', which is the KuaiLive paper, while the submission is numbered arXiv:2508.05634 (cs.RO). This internal metadata confirms that the uploaded text is a different manuscript from the one described in the abstract. This is not a minor formatting issue; it indicates that the paper under review does not contain the work it claims to present.","section":"Page header (arXiv identifier)"},{"comment":"Even setting aside the topic mismatch, there is no method section for the proposed approach. There is no description of the conformal predictor, its calibration set, the uncertainty estimates, the constrained RL objective, the constraint formulation, or the out-of-distribution scenarios (velocity variations, policy changes, group dynamics). The sole method-related statement is the abstract's one-sentence mention of 'augments agent observations with prediction uncertainty estimates... through constrained reinforcement learning.' This is insufficient to assess correctness, novelty, or soundness.","section":"Sections 1-7 (methodological content)"}],"minor_comments":[{"comment":"The abstract states that code and videos are available at https://gen-safe-nav.github.io/, but the body references the KuaiLive project page (https://imgkkk574.github.io/KuaiLive). No gen-safe-nav resources appear anywhere in the manuscript.","section":"Abstract"},{"comment":"The reference list consists entirely of recommender-systems and live-streaming papers, with no citations to crowd navigation, conformal prediction, or reinforcement learning literature. This is consistent with the body being a different paper, but it further underscores the absence of the claimed research context.","section":"References"},{"comment":"The paper's title and subject classification (cs.RO) do not match the content (cs.IR). While this could be an upload or metadata mix-up, as a submitted manuscript it is internally inconsistent.","section":"Title and metadata"}],"recommendation":"reject","confidential_remarks":"This appears to be a manuscript mix-up or error in the submission pipeline rather than a deliberate misrepresentation: the body is a complete, coherent dataset paper by different authors. However, as submitted, the paper is un-reviewable. The abstract's claims are entirely unsupported by the text, and there is no way to evaluate the proposed method or results. Rejection is appropriate; the authors or editors may wish to resubmit the correct manuscript (or correct the metadata) if this was an error."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, this submission is two different papers stapled together. The abstract describes a crowd-navigation method using adaptive conformal inference and constrained RL, with specific numbers (96.93% success, 8.80% over SOTA, 3.72x fewer collisions). The body is a SIGIR '26 dataset paper, 'KuaiLive: A Real-time Interactive Dataset for Live Streaming Recommendation.' None of the crowd-navigation content appears in the body: no method, no conformal derivation, no simulation setup, no baseline table, no OOD experiments, no robot deployment. This is a structural mismatch, not a missing appendix. I'm not alleging intent; an upload or metadata mix-up could explain it. But as a submitted manuscript, the central claims are unverifiable.\n\nTo give credit: the KuaiLive body looks like a competent dataset paper. It presents a large-scale real-world dataset (23,772 users, 452,621 streamers, 11.6M live rooms) with fine-grained interaction types, room lifecycle timestamps, and benchmarks over several recommendation methods. If this were the submission, it would be a plausible dataset contribution. But it is not the paper the abstract advertises.\n\nThe mismatch is load-bearing. The abstract's quantitative claims come with no protocol, no derivations, no error bars, and no baselines in the submitted text. I can't check the claimed novelty against prior crowd-navigation work because the body's related work is about recommendation systems. The reader's theoretical concern about conformal calibration under policy feedback is legitimate, but it's untestable here because there is no algorithm to audit. Self-citation isn't the issue; the cited crowd-navigation work simply isn't in the document.\n\nWho benefits from this as submitted? Nobody. The crowd-navigation paper, if it exists elsewhere, deserves a serious referee; this upload does not. A desk reject is the right call, with the practical next step being a request to correct the submission or resubmit the right manuscript. Do not send this to peer review.","headline":"The uploaded manuscript body is a live-streaming dataset paper, not the crowd-navigation paper promised in the abstract, so the reported safety gains are unsupported by anything in the submission.","tokens_in":12123,"tokens_out":3876,"would_cite":false,"duration_ms":33695,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Crowd navigation gets safer by feeding robot policies calibrated uncertainty bounds from conformal inference, the abstract claims.","keywords":["crowd navigation","conformal inference","uncertainty handling","constrained reinforcement learning","distribution shift","pedestrian trajectory prediction","robot safety","adaptive conformal prediction"],"falsifier":"Run the same benchmark with the conformal uncertainty feature replaced by the pedestrian predictor's own variance estimate or by a fixed uncertainty constant: if success rate and collision counts stay roughly unchanged, the conformal bounds carry no causal safety benefit. A second decisive check is to measure empirical coverage of the claimed uncertainty intervals under the closed-loop policy and confirm it remains near the nominal level when pedestrian behavior shifts.","tokens_in":11325,"feed_emoji":"🤖","tokens_out":1680,"duration_ms":20610,"temperature":0.7,"pith_summary":"The paper's abstract claims that a crowd-navigation robot can stay safe under distribution shifts if its policy is fed uncertainty estimates for each pedestrian rather than raw predictions alone. The proposed method augments the robot's observations with adaptive conformal inference uncertainty bounds and uses constrained reinforcement learning to keep the robot out of regions where predictions are unreliable. In the in-distribution benchmark the abstract reports a 96.93% success rate, 8.80 percentage points above prior baselines, with 3.72 times fewer collisions and 2.43 times fewer intrusions into pedestrians' future paths. Under shifted velocities, changed pedestrian policies, and transitions from individual to group motion, the abstract claims the same approach degrades far less than baselines, and a real-robot deployment is reported. The uploaded manuscript body, however, is a different paper on a live-streaming recommendation dataset, so the claims rest entirely on the abstract.","feed_headline":"Conformal uncertainty bounds make crowd navigation safer, abstract claims","feed_subtitle":"Robots that read pedestrian prediction uncertainty may cut collisions by 3.7x and stay safe under shifts.","key_machinery":"The key mechanism is the coupling of adaptive conformal inference with constrained reinforcement learning. Adaptive conformal inference produces a per-pedestrian uncertainty bound on trajectory predictions, and the constrained RL policy treats that bound as a safety constraint rather than as mere side information. The mechanism is meant to work by making the robot slow down, stop, or reroute when the predicted uncertainty in a nearby pedestrian's motion is high, and the adaptive component is meant to keep the bounds calibrated as the environment changes.","core_discovery":"The central claim, as stated in the abstract, is that accounting for prediction uncertainty is what makes a crowd-navigation policy robust to distribution shifts. The proposed system generates pedestrian trajectory prediction uncertainty with adaptive conformal inference, appends those uncertainty estimates to the robot's observations, and then trains the robot's policy with constrained reinforcement learning so that the uncertainty signals actively regulate the agent's actions. The abstract reports that this yields a 96.93% success rate in the in-distribution environment, over 8.80% higher than the previous state-of-the-art baselines, with more than 3.72 times fewer collisions and 2.43 time","pith_inferences":["If the claims are accurate, a testable extension is to replace adaptive conformal inference with other calibration schemes and compare whether the safety gains persist; the abstract's framing suggests the calibration feedback is the active ingredient, not the specific algorithm.","A subtle consequence the abstract does not address: because the RL policy actively reacts to the uncertainty signal, the conformal coverage guarantee is violated in principle, so the reported safety margins may depend on how mild the distribution shifts actually are.","The reported 2.43x reduction in intrusions into ground-truth future trajectories implies the method is not just avoiding instantaneous collisions but is also avoiding paths where pedestrians are later predicted to be, which could be reformulated as a safety metric for evaluating any crowd navigation policy.","The real-robot result, if reproducible, suggests that uncertainty-constrained RL could generalize beyond pedestrians to other dynamic obstacles, such as cyclists or drones, where prediction uncertainty is large and non-stationary."],"forward_implications":["If the abstract's claims hold, uncertainty-aware observation augmentation becomes a direct recipe for making learned crowd navigation robust to distribution shifts without retraining on new pedestrian behaviors.","The 3.72x reduction in collisions and 2.43x reduction in future-trajectory intrusions suggest that most of the safety gain comes from avoiding high-uncertainty regions, not from better point predictions.","The claimed robustness across velocity, policy, and group-dynamics shifts points toward a general principle: safe navigation policies should be trained on uncertainty features, not just on predicted trajectories.","Deployment on a real robot implies the uncertainty bounds can be computed online at control rates, making conformal inference a practically usable safety layer for mobile robots.","The method, as described, would apply to any pedestrian-prediction backbone, since the uncertainty estimate is an augmentation of the observation rather than a replacement of the predictor."],"supporting_citations":[],"fun_headline_variants":["Conformal uncertainty cuts robot collisions 3.7x in crowds","Robots navigate crowds safer by reading prediction uncertainty","Uncertainty-aware policy boosts crowd nav success to 96.93%","How conformal inference makes crowd navigation robust to shifts","Predicting uncertainty helps robots avoid collisions in crowds"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the adaptive conformal uncertainty estimates stay calibrated and informative even while the robot's own policy is actively reacting to them, so that the safety constraints they impose are meaningful rather than empty; the manuscript body provides no argument or experiment showing this.","fun_headline_variants_meta":{"raw":{"variants":["Conformal uncertainty cuts robot collisions 3.7x in crowds","Robots navigate crowds safer by reading prediction uncertainty","Uncertainty-aware policy boosts crowd nav success to 96.93%","How conformal inference makes crowd navigation robust to shifts","Predicting uncertainty helps robots avoid collisions in crowds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1036,"prompt_tokens":728,"completion_tokens":308,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":226}},"tokens_in":472,"tokens_out":308,"duration_ms":3224,"temperature":1.0,"reasoning_tokens":226,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:11:06.883392+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same benchmark with the conformal uncertainty feature replaced by the pedestrian predictor's own variance estimate or by a fixed uncertainty constant: if success rate and collision counts stay roughly unchanged, the conformal bounds carry no causal safety benefit. A second decisive check is to measure empirical coverage of the claimed uncertainty intervals under the closed-loop policy and confirm it remains near the nominal level when pedestrian behavior shifts.","supporting_citations":[],"review_version":1}