{"id":"68d8d5a2-7252-40f9-bfbe-8558862429b6","arxiv_id":"2608.10784","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In an O-RAN simulation testbed, three UAVs are detected far more often than they are tracked because two targets often share one detection, and emulation wall-clock delays distort the tracking metrics.","lead":"This paper tests a simulated 5G network that uses uplink reference signals as passive radar to track three drones at the same time. It finds that once every drone is detected, the bottleneck is keeping detections and identities straight, and that a vertical antenna array mainly helps decide which detection belongs to which drone.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The contention-vs-sensitivity conclusion depends on a detector capacity that is sized from the true target count and on a single-scenario contention estimate; both should be stress-tested before the claim is treated as general.","rationale":"The reader's weakest_assumption correctly identifies the declared N_T as load-bearing. My stress-test confirms this is the central soft spot: the paper's own caveats (Section VIII-C) admit that N_T is not estimated and that the fix does not generalize, and the strongest claim about contention binding before sensitivity is constructed on a detector that is pre-informed of the target count. The claim is internally consistent under the stated configuration, but it is not robust to one of the paper's own declared limitations. The reader's verdict of CONDITIONAL is appropriate; the paper should be accepted only under the condition that the N_T dependence is disclosed and ideally stress-tested. The alternative-heading experiment (Section IX-G) provides some evidence that the contention findings transfer to a more stressing geometry, but the paper itself notes the control did not hold, so the transferability of the magnitude is still uncertain. I do not find a basis for REJECT: the paper is honest about its limitations, the metrics are well-defined, and the mechanisms are argued from causes. The missing code URL is a separate, addressable issue; the paper claims reproducibility but the link is absent. Overall, the concern is real but not fatal, hence CONDITIONAL.","tokens_in":36662,"tokens_out":1605,"duration_ms":14966,"concrete_test":"Re-run the detector with an over-declared and under-declared target count (e.g., N_T=2 and N_T=4) on the same recorded detections, and recompute per-target availability and the contested-cell fraction. If the weakest target's availability drops below the contention-limited level, or if the gap between detection and tracking narrows, the \"contention, not sensitivity\" claim is an artifact of the declared N_T. Additionally, re-score the diagonal-geometry campaign with the detector capacity derived from N_T=3 as in the baseline; if the 97.4% contested fraction persists and availability falls for all targets, the contention claim holds under higher contention, but if the detector drops a target outright, the sensitivity limit resurfacing must be acknowledged.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim, that contention rather than sensitivity is the binding constraint, rests on two coupled conditions. First, the detector's capacity budgets (raw CFAR crosser list and detection cap) are derived from a declared target count N_T=3 (Section VIII-C), which the paper itself flags as indefensible in the field. An under-declared N_T silently drops the weakest target, so the measured per-target availability (83.7/72.7/53.7%) and the \"detected far more often than tracked\" gap are conditional on the detector being pre-told how many targets to expect. The paper's contention metric is also scenario-specific: the 73.7±4.8% contested-cell fraction is measured on one favorable geometry, and the single alternative-heading experiment shows contention rising to 97.4±1.4% (Section IX-G), with a caveat that the control did not hold because all targets' availability fell. This means the \"contention, not sensitivity\" ordering could invert or attenuate under a wrong or unknown N_T, or under geometries with higher contention, because availability drops when contention rises (as the diagonal experiment shows). The authors themselves state N_T is \"declared, not estimated\" and that the fix \"does not generalize\" (Section VIII-C), so the weakness is acknowledged but not resolved; the headline conclusion is therefore not robust to the paper's own acknowledged limitation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates multi-target UAV detection, association, and tracking on an O-RAN simulation testbed that reuses the 5G NR uplink sounding reference signal as a passive radar waveform. Three UAVs with different altitudes, velocities, and radar cross sections fly a single bistatic geometry over a 2x4 planar receive array. The evaluation is scored against what a counter-UAS C2 consumer needs, not detection accuracy alone. The main claims are: (i) once every target is detected, contention—two targets sharing one nearest detection—is the binding limit rather than sensitivity; (ii) targets are detected far more often than they are tracked; (iii) the vertical aperture buys association rather than localization; and (iv) concurrent tracks can be exported to a C2 fusion node over a SAPIENT adapter carrying a calibrated detection confidence. The statistical design is 10 noise seeds on one deterministic trajectory and one gNB siting, plus a single alternative-heading experiment, all in emulation without radios.","tokens_in":36947,"tokens_out":3632,"duration_ms":38828,"significance":"If the contention-not-sensitivity claim holds, the paper makes a useful and falsifiable point: for this class of cellular ISAC sensors, the bottleneck in multi-target tracking is association and identity preservation, not raw detection. The paper is unusually disciplined in its evaluation reporting. It discloses that the N=10 campaign is actually two batches of 5 from different repository revisions, runs permutation tests at the revision boundary, withdraws the small-sample identity claim when it does not survive doubling, conditions the gated RMSE and quantifies its selection effect, and repeatedly states that magnitudes are scenario-specific while mechanisms are what transfer. It also ships committed artifacts and a single script that regenerates all figures, which is exemplary for reproducibility. These strengths make the paper's limitations especially important: the central measurement, per-target availability under contention, is obtained with a detector whose capacity is set from the true target count, a quantity unavailable to a real counter-UAS sensor.","major_comments":[{"comment":"The detector's raw CFAR crosser list and per-CPI detection cap are derived from a declared target count N_T=3 (560 and 18 entries), and the paper states that N_T is 'declared, not estimated' and that an under-declared count 'silently drops the weakest target.' Because the per-target availability figures (83.7/72.7/53.7%) and the contested-cell fraction (73.7±4.8%) are measured under a detector pre-told the true number of targets, the headline 'contention, not sensitivity' conclusion is conditional on exactly the field condition that fails. Please add a robustness experiment with over- and under-declared N_T (e.g., 1, 2, 4, 5) showing whether the weakest target is dropped and whether the contention-versus-sensitivity ordering survives, or explicitly reframe the conclusion as applying only to a detector with a priori target-count knowledge.","section":"Section VIII-C and Table I"},{"comment":"The diagonal-heading experiment is the only geometric variation, and its control did not hold: availability fell for all three targets, including the two whose trajectories were unchanged (UA V-A 83.71→77.93, UA V-C 53.73→46.94). The paper reports this honestly, but it means the experiment cannot attribute the contention rise to the heading change rather than to other scene differences, and the 73.7±4.8% contention estimate therefore rests on a single favorable geometry. Please provide at least one additional geometry in which the unchanged targets hold still, or restrict the contention claim to the rendered scenarios rather than presenting it as a transferable mechanism.","section":"Section IX-G and Table IX"},{"comment":"The 7.5 m bistatic range-bias constant was configured but never reached the estimator, so every bistatic range figure carries the uncorrected under-read. This directly affects the per-target accuracy rows in Table VII and the statement that UA V-A meets the 1 m to 10 m use-case requirement (8.13 m). While the disclosure is scrupulous, the one quantitative comparison to an external KPI is made against a knowingly biased measurement. Please report the corrected values, or state quantitatively why the bias does not change the comparison, before the KPI table is used.","section":"Section X, Table X, and Table VII"}],"minor_comments":[{"comment":"The string 'UA V' appears throughout the title and abstract where 'UAV' is intended; this appears to be a rendering artifact and should be normalized.","section":"Title and Abstract"},{"comment":"The caption says '4 object_id' and the text says the tracker carried the 3 targets on five slots of which four reached confirmation; this is consistent but the relationship between 'five slots' and 'four identities' would be clearer if explained in the caption or a table.","section":"Section XII-D and Figure 3"},{"comment":"The statement that the bridged wall-clock continuity figure is below the unbridged sensor-timeline one is central to the timeline argument; a small table reporting single-target continuity on both timelines with and without bridging would make this easier to read than the current prose.","section":"Section XII-B"},{"comment":"The companion manuscript [8] carries the coordinate frame, measurement model, and platform description; please state its review status or provide a stable citation so readers can verify that the deferred derivations are publicly available.","section":"Section I and Reference [8]"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is defensible only under the declared-target-count condition that the authors themselves flag as indefensible in the field. The robustness experiment I request in major comment 1 is cost-effective and decisive for the contribution's scope. The paper's reporting discipline is unusually high; if the N_T sensitivity can be resolved, I would support publication. The journal fit is appropriate for cs.NI given the O-RAN and C2 focus, though the companion paper owns much of the platform description."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on arXiv:2608.10784. It's a genuinely useful evaluation paper, one of the most disciplined simulation studies I've seen in this area. The new content is the end-to-end integration: UL-SRS passive sensing in OAI, detector to EKF xApp on FlexRIC, and multi-track export over SAPIENT to a mock C2. That chain, with the timeline replay and the failure-mode taxonomy (sensitivity, contention, availability, admission), is real new material. The paper earns credit for separating wall-clock dilation from sensor behavior—the 38% coverage ceiling argument is a nice catch—and for its statistical hygiene: it discloses the 5+5 batch split, runs permutation tests at the revision boundary, and withdraws the N=5 identity claims once N=10 didn't support them. That kind of honesty is rare.\n\nThe soft spots are real but mostly acknowledged. The load-bearing one is N_T. The detector's capacity budgets come from the declared target count, and the paper says plainly this is 'declared, not estimated' and 'does not generalize.' That means the availability numbers (83.7/72.7/53.7) and the contention-vs-sensitivity conclusion hold only for a detector pre-told how many targets to expect. Under-declare and the weakest target is silently dropped; the whole ordering could change. The single alternative-heading experiment also has a broken control—all targets' availability fell—which the paper reports honestly but which limits what that experiment can conclude. Add emulation-only, one geometry, ten seeds, and the missing code link (the appendix says artifacts exist but gives no repository URL) and you have a paper whose magnitudes are definitely not portable.\n\nStill, the mechanisms are argued from their causes and the analysis is careful. The per-target geometry screens, the elevation-observability inequality, the gated-RMSE conditioning—these are the work of someone thinking clearly. The conclusions are scoped properly for the most part. I'd send this to a serious referee. It deserves engagement, and the referees can push on the N_T dependence and ask for a code link. It's not a field-changing paper, but it's a solid, reproducible-in-principle testbed evaluation that the ISAC community will want to read.","headline":"A disciplined, honestly-scoped O-RAN ISAC testbed evaluation whose headline claim (contention, not sensitivity) is real but conditional on a declared target count the authors themselves flag as indefensible in the field.","tokens_in":37501,"tokens_out":2300,"would_cite":true,"duration_ms":24271,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that multi-target UAV tracking on a 5G integrated-sensing link is limited by detection contention and identity management rather than by sensitivity, and that wall-clock evaluation artifacts can masquerade as tracking…","keywords":["integrated sensing and communication","O-RAN","5G NR uplink sounding reference signal","multi-target tracking","data association","counter-UAS","bistatic radar","uniform planar array"],"falsifier":"Run the same three-target scene with the detector's capacity supplied as two targets instead of three: if the weakest target's detections vanish silently from the range-Doppler map, then the contention-not-sensitivity conclusion is an artifact of the declared target count rather than a property of the sensor.","tokens_in":36459,"feed_emoji":"📡","tokens_out":6798,"duration_ms":65840,"temperature":0.7,"pith_summary":"The paper tries to establish that, for cellular integrated sensing and communication, the binding constraint on multi-target tracking is contention for detections rather than sensitivity, and that a counter-UAS consumer needs identity and track availability, not just detection accuracy. It builds an emulation-only testbed that repurposes the 5G uplink sounding reference signal as a passive radar waveform, flies three UAVs of differing size and altitude through one bistatic pair, and measures where the pipeline loses targets. The claim, if true, redirects effort: a sensor that detects every target can still fail to track them because two targets share one detection in about 74% of coherent processing intervals, and the weakest targets are detected far more often than they are tracked. The paper also argues that evaluation artifacts, such as the testbed running far slower than real time, can change the diagnostic; re-scoring the same detections on the sensor's own timeline lifts coverage and inverts the error decomposition.","feed_headline":"Contention, not sensitivity, is the multi-drone tracking bottleneck","feed_subtitle":"In 74% of scans two targets share one detection; identity, not sensitivity, is what breaks.","key_machinery":"The carrier is the uplink-sounding-reference-signal sensing chain: clutter-subspace deflation, range-Doppler order-statistic CFAR with non-maximum suppression, single-snapshot interferometric angle estimation on a 2×4 planar array, and an extended Kalman filter with probabilistic data association and birth inhibition running as a RIC xApp. The decisive quantity is the contested-cell fraction, the share of coherent processing intervals in which two targets' nearest detection is the same detection; it is what separates the contention claim from a sensitivity claim. A second load-bearing mechanism is the sensor-timeline replay, which re-spaces recorded detections at the 0.64 s coherent-processing-interval period so that tracker constants carrying wall-clock units, such as publication lifetime and coverage polling, no longer dominate the metric.","core_discovery":"This paper argues that for three simultaneous UAVs observed by one bistatic pair with an 8-element planar receive array, the sensor detects all three targets but cannot feed a usable multi-target picture to a counter-UAS command-and-control system. The binding limit is contention, not sensitivity: in 73.7±4.8% of coherent processing intervals two targets' nearest detection is the same detection, and targets are detected far more often than they are tracked (one target is detected in 72.7±4.0% of intervals yet holds a track in 12.9% of its life). Elevation from the planar array does not separate targets that share a range-Doppler cell, but it is the deciding association discriminant in 58% of intervals. The paper also claims that the testbed's wall-clock slowness imposes a 38.0% coverage ceiling that makes 'mostly tracked' unreachable by construction, and that re-scoring the same detections on the sensor's own timeline raises coverage from 14.6% to 41.5% and changes the GOSPA error decomposition.","pith_inferences":["If the contention result transfers beyond this geometry, then the highest-value work is identity management, including birth admission, multi-hypothesis tracking, and track-level confidence, rather than waveform or sensitivity engineering.","The declared-target-count caveat suggests a deployable detector needs an online count estimator or a capacity law robust to under-declaration; the paper shows exactly why silent drops are the failure mode to guard against.","The sensor-timeline replay discipline likely generalizes to any hardware-in-the-loop or emulation campaign: metrics scored on host wall-clock time can be dominated by execution-rate artifacts, so publications should state the units of every tracker constant.","A downstream fusion node should not integrate the exported detection confidence over a track's life, because the calibration is per cell and ignores association ambiguity; the paper explicitly leaves track-level confidence uncalibrated, implying that is the next necessary step."],"forward_implications":["For cellular-ISAC multi-target tracking, once every target is detectable, further sensitivity investment will not improve C2-useful tracking; the levers are association and birth-admission rules.","A counter-UAS consumer that needs stable identifiers will judge this sensor unusable even though detection is largely successful, because two of the three targets hold tracks in well under a fifth of their detection-bearing intervals.","Any evaluation of a real-time-constrained digital twin must either run at or above real time or state which timeline its tracker constants live in; wall-clock scoring can create a coverage ceiling that looks like tracker failure.","A vertical aperture on a single array should be budgeted for association rather than promised as 3-D localization, since it cannot resolve coincident range-Doppler cells but can settle association in 58% of intervals in this geometry."],"supporting_citations":[{"why":"Provides the single-target sensing chain, coordinate frame, and measurement model that this multi-target evaluation inherits.","marker":"[8]"},{"why":"Supplies the 1 m to 10 m horizontal accuracy use-case range used to judge per-target tracking results.","marker":"[1]"},{"why":"Defines the order-statistic CFAR detector whose guard bands and thresholds shape multi-target detection behavior.","marker":"[23]"},{"why":"Supplies the disturbance-removal clutter-deflation approach adapted to the SRS channel estimates.","marker":"[22]"},{"why":"Establishes the ESPRIT relation with which the closed-form interferometric angle estimator coincides on a uniform linear array.","marker":"[24]"},{"why":"Provides the ray-tracing channel emulator that produces per-target echo paths and Doppler.","marker":"[21]"},{"why":"Provides the near-real-time RIC that hosts the EKF association and tracking xApp.","marker":"[19]"},{"why":"Provides the 5G NR protocol stack whose uplink sounding reference signal is repurposed as the sensing waveform.","marker":"[20]"},{"why":"Demonstrates multi-target tracking on a 5G-compliant ISAC proof-of-concept, serving as the comparison point for this testbed's approach.","marker":"[9]"}],"fun_headline_variants":["74% of scans: two drones share one detection, breaking tracking","Shared detections, not sensitivity, limit multi-drone tracking","Detection is easy, tracking is hard: contention is the bottleneck","Multi-drone tracking's real limit: shared detections, not sensitivity","When two drones share one detection, tracking fails"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The detector's internal buffers are sized from the true target count, which a real counter-UAS sensor would not know, so every availability and coverage figure in the paper assumes the sensor has been told how many targets to expect.","fun_headline_variants_meta":{"raw":{"variants":["74% of scans: two drones share one detection, breaking tracking","Shared detections, not sensitivity, limit multi-drone tracking","Detection is easy, tracking is hard: contention is the bottleneck","Multi-drone tracking's real limit: shared detections, not sensitivity","When two drones share one detection, tracking fails"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000961,"raw_usage":{"total_tokens":4123,"prompt_tokens":1003,"completion_tokens":3120,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":3034}},"tokens_in":619,"tokens_out":3120,"duration_ms":23507,"temperature":1.0,"reasoning_tokens":3034,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:25:00.082032+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three-target scene with the detector's capacity supplied as two targets instead of three: if the weakest target's detections vanish silently from the range-Doppler map, then the contention-not-sensitivity conclusion is an artifact of the declared target count rather than a property of the sensor.","supporting_citations":[{"cited_title":"5G ISAC-Based UAV Detection and 3-D Tracking Using Uplink Sounding Reference Signals on an End-to-End O-RAN Simulation Testbed","cited_arxiv_id":"2608.05826","evidence_quote":"Provides the single-target sensing chain, coordinate frame, and measurement model that this multi-target evaluation inherits."},{"cited_title":"Experimental Demonstration of Multi-Target Tracking in Integrated Sensing and Communication","cited_arxiv_id":"2510.22180","evidence_quote":"Demonstrates multi-target tracking on a 5G-compliant ISAC proof-of-concept, serving as the comparison point for this testbed's approach."}],"review_version":1}