{"id":"c246af1b-f252-4196-9a02-25e44327fb2c","arxiv_id":"2508.20892","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A taxonomy that classifies unified perception methods in autonomous driving into Early, Late, and Full Unified Perception based on task integration, tracking formulation, and representation flow.","lead":"This survey proposes a new way to organize research on self-driving car perception: it groups methods that combine detection, tracking, and prediction into three paradigms. It gives the community a shared vocabulary and a map of open problems, which may speed up work on more robust and interpretable perception systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Taxonomy's own classification rule contradicts its LUP table: Sec. V declares PNA infeasible for LUP/FUP, yet MTP is reviewed and included as 'a form of PNA.'","rationale":"I read the paper as a survey whose primary contribution is a taxonomy that is both complete and internally consistent. The three paradigms are plausible, the tables are detailed, and the authors acknowledge prior FUP work by Dal'Col et al. The strongest challenge I found is not the broad worry that tracking may not be the central design element for every occupancy- or flow-based method, although that is a legitimate concern. The more immediately falsifiable problem is the internal contradiction in Sec. V: PNA is declared infeasible for LUP/FUP, yet MTP is simultaneously included and explicitly called 'a form of PNA.' This is a concrete inconsistency in the taxonomy's central axis, and it means the 'internally consistent mapping' claim is not currently true as written. The reader's weakest assumption was about tracking centrality; my concern is adjacent but distinct—it is about the consistency of the tracking-formulation axis at the LUP boundary. The issue is fixable without abandoning the taxonomy, so a conditional acceptance remains appropriate. I would not reject the paper on this basis, and I am not accusing the authors of dishonesty; the contradiction appears to be an oversight in wording rather than a fundamental flaw. The proposed test—reclassifying MTP by the paper's own definitions—would settle whether the contradiction is real or whether I have misread the scope of the infeasibility claim.","tokens_in":35046,"tokens_out":4722,"duration_ms":54010,"concrete_test":"Reclassify MTP using the paper's own Sec. III definitions. Specifically, check whether Murty's H-best assignment is (a) performed outside the network, (b) non-differentiable, and (c) used to form track hypotheses that are then passed to the prediction module. If all three hold, MTP satisfies the paper's definition of PNA. Then test the conjunction 'PNA is not feasible in LUP' against this classification; if both are true, the contradiction is confirmed. This analytical check requires reading MTP (Weng et al., 2022) and comparing it with Sec. V and Table III; no experiments are needed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's central claim is that the taxonomy provides a complete and internally consistent mapping of the unified perception design space. That claim is undercut by a direct internal contradiction in the tracking-formulation axis. In Sec. V the authors write: 'As tracking is addressed in union with prediction, PNA is not feasible, as post hoc association would break the joint reasoning process and prevent shared optimization between tasks.' But in the same section, MTP is described as generating tracking hypotheses via 'a deterministic, non-differentiable process external to the network, a form of PNA, meaning MTP does not fully realize end-to-end implicit tracking.' MTP is also included in Table III as an LUP method. If MTP is correctly characterized as PNA, then PNA is in fact feasible in LUP (at least as a partially unified method), and the stated infeasibility rule is false. If the rule is correct, then MTP should be excluded or reclassified, contradicting its inclusion. Either way, the taxonomy's organizing axis—explicit/implicit tracking and PNA/INA—does not consistently classify the very methods it surveys. This is not a stylistic quibble: the contradiction sits at the boundary between two paradigms and affects the completeness claim, since LUP is claimed to be a distinct paradigm with a well-defined design space. The fix is likely small (e.g., 'PNA cannot support full end-to-end LUP' or 'MTP is a boundary case'), but as written the text cannot simultaneously assert both statements.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey proposes a taxonomy for unified perception in autonomous driving, organizing methods into Early, Late, and Full Unified Perception (EUP, LUP, FUP) along three axes: task integration, tracking formulation (explicit/implicit, PNA/INA), and representation flow. It reviews a large set of existing joint detection/tracking/prediction methods, tabulates their attributes (modality, input, paradigm, output, architecture, learning strategy, datasets, code availability), and discusses intermediate representations, training strategies, and future research directions.","tokens_in":35374,"tokens_out":5655,"duration_ms":57498,"significance":"If the taxonomy is internally consistent, it fills a clear gap in the survey literature by providing the first systematic framework covering EUP, LUP, and FUP, and by positioning them relative to modular and end-to-end stacks. The paper's strengths are its broad method coverage, explicit scope decisions (e.g., separating FUP from planning), structured tables, and concrete discussion of representational and training trade-offs. The main risk to the central claim is the consistent application of the tracking-formulation axis, and the manuscript currently contains several internal contradictions on exactly this point.","major_comments":[{"comment":"The taxonomy's PNA/INA axis is applied inconsistently at the LUP boundary. Sec. V states: 'As tracking is addressed in union with prediction, PNA is not feasible, as post hoc association would break the joint reasoning process and prevent shared optimization between tasks.' Yet the same section describes MTP [118] as generating tracking hypotheses via 'a deterministic, non-differentiable process external to the network, a form of PNA,' and Table III includes MTP as an LUP method under 'Imp. T.' If MTP is PNA, then PNA is feasible in LUP; if PNA is infeasible, then MTP is misclassified. Please revise the claim (e.g., 'full end-to-end PNA is not feasible') and align the MTP table entry with the text.","section":"Sec. V and Table III"},{"comment":"Sec. IV.C states 'As the association is performed internally, JDTQ methods are inherently INA.' Two paragraphs later, TransTrack [75] is classified as JDTQ but described as relying on 'post-decoding IoU matching, which makes it PNA,' and Table II lists it as 'JDTQ - PNA.' This directly contradicts the 'inherently INA' claim and the definition of JDTQ. Either TransTrack should be treated as a boundary/hybrid case, or the general statement should be qualified to exclude methods with external association.","section":"Sec. IV.C and Table II"},{"comment":"Sec. VI.C introduces 'an alternative field' that 'generates occupancy outputs without relying on object detection, implicit tracking, and prediction as separate goals,' then groups Khurana et al. [179] and Occ4Cast [180] under this description. However, the subsection heading is 'Occupancy Output and Implicit Tracking' and Table IV labels both methods 'Imp. T.' This is an internal contradiction in the tracking axis that affects the completeness claim for FUP. Please reclassify these methods, or revise the narrative to explain why they are considered implicit tracking despite the text saying they do not rely on it.","section":"Sec. VI.C and Table IV"}],"minor_comments":[{"comment":"The caption uses 'JQDT' while the text and Table II consistently use 'JDTQ.' Please correct the typo.","section":"Fig. 4"},{"comment":"There are numerous LaTeX spacing artifacts, e.g., 'V oxelNeXt,' 'V oxel,' and similar in tables and body text. Please fix the source to render 'Voxel...' correctly.","section":"General formatting"},{"comment":"The legends for paradigm abbreviations ('Imp. T', 'Exp. T', 'A+M', 'M', etc.) are split across table footnotes and are not fully self-contained. A reader should be able to parse the tables without going back to Sec. III. Please add complete legend entries directly under each table.","section":"Tables III and IV"},{"comment":"The open-source statistics (31.1% EUP, 85.7% LUP, 75.0% FUP) are stated without derivation. Since 'Code' availability is marked only as ✓/✗ in Tables II-IV, please clarify how these percentages are computed and which entries count as open-source.","section":"Sec. VII.C"}],"recommendation":"major_revision","confidential_remarks":"The paper is a survey and the central contribution is the taxonomy; no derivations or empirical results are at stake. The three major comments all concern the same root cause: the tracking-formulation axis is applied inconsistently within and across sections. These are fixable with wording changes and table updates, but they directly affect the paper's claim of a complete and internally consistent mapping, so major revision is appropriate. The paper is within the scope of the journal and the authors' positioning against prior surveys is fair."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this survey earns a serious referee. The three-paradigm split — Early, Late, and Full Unified Perception — is a real organizational contribution, and the LUP category is new in this form. I learned something from the tables, and the authors are honest about where they build on Dal'Col et al. and where they depart.\n\nWhat's good: the taxonomy is simple enough to stick but fine-grained enough to discriminate methods. Tracking formulation (PNA/INA, explicit/implicit) and representation flow are the right axes for most of this literature, and applying them across EUP, LUP, and FUP gives the community a working vocabulary. The tables are dense but useful, and the open-source availability statistics in the discussion are a nice touch. The writing is clear; I had no trouble following the categorization logic, which is not always true for surveys.\n\nThe soft spots are real but narrow. The reader's and stress-test both caught the same thing: Sec. V says PNA is not feasible for LUP/FUP because post hoc association breaks joint reasoning, yet the same section describes MTP's external hypothesis generation as \"a form of PNA\" and MTP is in the LUP table. As written, the text can't have both. The fix is small — soften to \"PNA cannot support full end-to-end LUP\" or mark MTP as a boundary case — but the contradiction sits on the taxonomy's main organizing axis, so it should be corrected before publication.\n\nSecond, the \"first comprehensive framework\" claim is mildly overstated. LUP is genuinely under-surveyed and the EUP/LUP/FUP framing is new, but FUP overlaps with Dal'Col et al. and the literature selection process is not documented. A sentence on search strategy or inclusion criteria would remove the worry.\n\nThird, the tracking-centric assumption is defensible but not universal. A method built around shared occupancy flow without any identity-like structure would not fit the organizing axis cleanly. The paper says tracking is \"the central design element\" — that's a scope choice, not an error, but it deserves a caveat.\n\nNothing here smells circular or self-serving. The few self-citations point to prior reviews, not to the taxonomy's own support. The math is a taxonomy, so there's no derivation to check — soundness is about internal consistency and coverage, and aside from the PNA/MTP issue, it holds up.\n\nWho's this for: anyone working on unified detection-tracking-prediction in AVs, particularly early-career researchers who need a map of the space. It deserves peer review and a conditional accept after the contradiction and framing issues are addressed. I'd cite it and bring it to a reading group.","headline":"A genuinely useful AV perception taxonomy with a small internal contradiction the authors can fix before publication.","tokens_in":35845,"tokens_out":1786,"would_cite":true,"duration_ms":22657,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Unified perception gets one taxonomy, three named levels","keywords":["unified perception","autonomous driving","scene understanding","multi-object tracking","motion prediction","taxonomy","detection-tracking-prediction","representation flow"],"falsifier":"Search for a published or constructable unified perception system that performs detection and prediction in one module with no identity assignment and no temporal-continuity mechanism in its latent features — e.g., a pure occupancy-flow predictor from raw sensors. If such a system performs competitively and cannot be placed in the taxonomy except by relabeling 'implicit tracking' as any temporal correlation, the tracking-centered claim is weakened. A second test: a controlled benchmark holding the backbone fixed and toggling only explicit vs implicit tracking would show whether the tracking ax","tokens_in":34965,"feed_emoji":"🚗","tokens_out":6021,"duration_ms":60871,"temperature":0.7,"pith_summary":"This survey sets out to give unified perception — the autonomous-driving idea of merging object detection, tracking, and motion prediction into shared modules instead of separate pipeline stages — a systematic map. Its claim is that every method in this scattered literature lands in one of three named levels: Early Unified Perception (detection + tracking), Late Unified Perception (tracking + prediction), or Full Unified Perception (all three), and that the dividing axis is tracking. The paper sorts methods by how tracking is done (explicit identity assignment vs. implicit temporal continuity in latent features) and by what representation flows between stages (bounding boxes, latent features, or occupancy). A reader gets a shared vocabulary for comparing systems, a way to spot empty design cells, and a statement of why the field's fragmentation has been blocking progress.","feed_headline":"Unified perception gets one taxonomy, three named levels","feed_subtitle":"Detection, tracking, and prediction merge in Early, Late, or Full Unified Perception; here is the design space.","key_machinery":"The organizing device is a three-level taxonomy centered on tracking. The levels are defined by which adjacent tasks share a module; the tracking formulation axis distinguishes explicit tracking (persistent object identities, assigned either after the network or inside it) from implicit tracking (soft temporal continuity in latent space), and the representation-flow axis distinguishes bounding-box, latent-feature, and occupancy intermediates. This machinery does the classification work: it decides which level a method belongs to, which training constraints apply (e.g., post-network association is ruled out in Late and Full Unified Perception because it would break joint reasoning), and which","core_discovery":"The paper's central claim is that unified perception is a design space, not a single architecture genre, and that the space is organized by the tracking stage. The taxonomy names three levels of task integration: Early Unified Perception integrates detection with tracking; Late Unified Perception integrates tracking with prediction; Full Unified Perception integrates detection, tracking, and prediction, optionally absorbing localization. Within that, tracking formulation is the load-bearing axis: explicit tracking assigns persistent identities via post-network or in-network association, while implicit tracking keeps temporal continuity in latent representations without committing to identiti","pith_inferences":["Beyond the paper: if tracking is truly the central axis, then identity-free occupancy-flow models are a stress test — a system that predicts future occupancy without any notion of object permanence may fit the taxonomy only by stretching 'implicit tracking' to mean any temporal consistency.","Beyond the paper: the taxonomy suggests a concrete controlled experiment — hold one detection backbone fixed and vary only the tracking formulation (implicit vs explicit); divergence in performance would show the tracking axis is causally meaningful, while no divergence would suggest the taxonomy is descriptive rather than functional.","Beyond the paper: the open 'occupancy intermediate representation' cell hints at a next generation of unified perception that treats the scene as a field rather than as a set of tracked instances, shifting evaluation from identity metrics toward occupancy and flow metrics.","Beyond the paper: the survey's own closed-source statistics (31.1% in EUP, 85.7% in LUP, 75.0% in FUP) imply that the least-studied levels are the least reproducible, so progress may depend as much on releasing artifacts as on proposing new algorithms."],"forward_implications":["Early Unified Perception methods, because they output bounding boxes with identities, are always explicitly tracked; no implicit-tracking or occupancy-based EUP method exists yet.","Late Unified Perception cannot use post-network association: a hard association step after the network would sever the joint tracking-prediction optimization, so LUP is forced toward either implicit affinity-based tracking or in-network differentiable association.","Full Unified Perception absorbs planning-adjacent methods only up to prediction; methods that extend to planning (e.g., end-to-end planners) are deliberately excluded from the level.","Intermediate representation choice is predictive of design constraints: bounding-box interfaces create differentiability bottlenecks and information compression, while latent representations avoid both and allow modular submodule replacement.","Open research directions follow directly: unified-vs-modular benchmarking, closed-loop evaluation, modality balance (EUP is image-heavy, FUP is LiDAR-heavy), and semi- or self-supervised training to reduce joint-task annotation cost."],"supporting_citations":[{"why":"Introduces the post-network vs in-network association distinction that the taxonomy's explicit-tracking axis is built on.","marker":"[34]"},{"why":"The only prior survey of full detection-plus-prediction methods, which this survey refines and separates from planning.","marker":"[33]"},{"why":"Recent deep-learning MOT survey that acknowledges joint detection-tracking but leaves it unsystematized; one piece of evidence for the gap.","marker":"[16]"},{"why":"Historical MOT survey with the same gap: tracking-by-detection is reviewed but early unified detection-tracking is not surveyed as a paradigm.","marker":"[17]"},{"why":"Supplies the widely accepted modular 3D detection taxonomy that the unified design space extends from.","marker":"[8]"},{"why":"Survey of motion prediction used to establish modular perception's limitations and the need for inter-task synergy.","marker":"[7]"},{"why":"Empirical study of perception inputs for motion forecasting; motivates LUP and FUP and the missing unified-vs-modular evaluation.","marker":"[115]"},{"why":"MOTR is the foundational in-network query-propagation method that later Full Unified Perception systems build on; anchors the query-based detection-tracking line.","marker":"[80]"}],"fun_headline_variants":["Unified perception: one taxonomy, three levels, tracking decides","Early, Late, or Full: how unified perception merges tasks","Tracking stage defines the shape of unified perception","Implicit vs explicit tracking: key to unified perception"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The taxonomy assumes tracking — persistent identity or at least temporal continuity — is the necessary hinge between detection and prediction in every unified perception system; if a unified method built on shared scene occupancy or flow needs no tracking-like structure, the organizing axis would misclassify or omit it.","fun_headline_variants_meta":{"raw":{"variants":["Unified perception: one taxonomy, three levels, tracking decides","Early, Late, or Full: how unified perception merges tasks","Tracking stage defines the shape of unified perception","Implicit vs explicit tracking: key to unified perception"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1309,"prompt_tokens":647,"completion_tokens":662,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":391,"completion_tokens_details":{"reasoning_tokens":596}},"tokens_in":391,"tokens_out":662,"duration_ms":7014,"temperature":1.0,"reasoning_tokens":596,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:43:16.095132+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search for a published or constructable unified perception system that performs detection and prediction in one module with no identity assignment and no temporal-continuity mechanism in its latent features — e.g., a pure occupancy-flow predictor from raw sensors. If such a system performs competitively and cannot be placed in the taxonomy except by relabeling 'implicit tracking' as any temporal correlation, the tracking-centered claim is weakened. A second test: a controlled benchmark holding the backbone fixed and toggling only explicit vs implicit tracking would show whether the tracking ax","supporting_citations":[{"cited_title":"Joint Perception and Prediction for Autonomous Driving: A Survey","cited_arxiv_id":"2412.14088","evidence_quote":"The only prior survey of full detection-plus-prediction methods, which this survey refines and separates from planning."}],"review_version":1}