{"id":"fce8b212-69aa-49b2-bd60-aa0a124b71cb","arxiv_id":"2411.10291","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A review of autonomous driving systems concludes that future in-car computers should combine general-purpose, specialized, and processing-in-memory accelerators, supported by a small CPU/GPU benchmark of three end-to-end models.","lead":"This paper surveys autonomous driving software and hardware, covering sensors, datasets, simulators, modular and end-to-end AI stacks, and CPU, GPU, FPGA, and SoC platforms. It adds a small benchmark suggesting that off-the-shelf CPUs cannot reach a 10 FPS target for three end-to-end driving models while GPUs can, and it argues for heterogeneous future hardware.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The V.A CPU/GPU benchmark cannot carry the V.D heterogeneity thesis: it never measures a heterogeneous configuration, and its uncited 10 FPS threshold plus missing methodology leave the quantitative premise fragile.","rationale":"The reader correctly identified that the Fig. 8 benchmark is under-specified and that the 10 FPS threshold lacks a source; I agree with that concern. My stress-test extends it: even if the three models were representative and the measurements were repeated carefully, the presented data would still not establish the central heterogeneity thesis, because the experiment never tests a heterogeneous configuration or a homogeneous accelerator other than the CPU. The paper's evidence can show 'these CPUs are weak; GPUs are stronger,' but not that 'a single type of hardware is suboptimal,' which is the exact premise Section V.B uses to motivate V.D. The survey has genuine strengths: a broad review of sensors, datasets, software stacks, and hardware platforms, plus a useful Fig. 9 layer-level arithmetic intensity analysis and a reasonable synthesis of PIM work. Those strengths support the plausibility of the prediction but not its quantitative grounding. Since the central claim is forward-looking and consensus-aligned, a conditional acceptance requiring the benchmark methodology, error bars, a sourced threshold, and ideally a direct homogeneous-versus-heterogeneous comparison remains appropriate; no verdict change is needed beyond the reader's CONDITIONAL.","tokens_in":971,"tokens_out":1053,"duration_ms":65093,"concrete_test":"Run the authors' benchmark, or the official TransFuser/InterFuser/MILE repositories, on CPU_4 (Ryzen 9 7950X, 32 threads) and GPU_1 (Jetson AGX Orin), fixing batch size 1, native input resolution, FP32, pinned threads, and measuring wall power over 5 independent runs; report mean and 95% CI per model. If the borderline CPU_4 FPS or the GPU_1 numbers move across the 10 FPS threshold, the 'CPUs cannot meet 10 FPS' and efficiency conclusions fail. Then add a third arm to settle the heterogeneity claim: the same workloads on a homogeneous GPU-only configuration versus a heterogeneous CPU+GPU+FPGA/CGRA or PIM-simulated configuration under equal power/area budgets; if the homogeneous arm already meets 10 FPS at comparable FPS/W, Fig. 8 cannot support Section V.D.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section V.D's central prediction is that future AD accelerators will be increasingly heterogeneous, combining CPUs, GPUs, PIM-based cores, FPGAs, and CGRAs. Section V.B grounds this in the assertion that 'managing these diverse models with a single type of hardware leads to suboptimal performance, as briefly demonstrated in the previous section.' That demonstration is Fig. 8: three end-to-end models (TransFuser, InterFuser, MILE) on four CPUs and four GPUs. This supports the narrow claim that these CPUs are slower than these GPUs on these models; it does not compare a homogeneous GPU-only platform with a heterogeneous CPU+GPU+FPGA/CGRA/PIM platform on the same workloads, nor does it show any single homogeneous accelerator failing to meet requirements. A GPU that is faster than a CPU is not evidence that heterogeneity is needed. The quantitative premise also lacks methodological detail: no batch size, input resolution, precision, framework version, measurement protocol, or error bars are given; the '10 FPS' minimum is asserted without a citation; and Table VII describes systems containing both CPU and GPU, so the isolation of 'CPU-only' versus 'GPU-only' is unclear. The 'except one with borderline results' qualifier is especially fragile: without confidence intervals, that single data point cannot support the strong statement that CPUs cannot meet 10 FPS. The paper's claim may be plausible and consensus-aligned, but its load-bearing empirical support is under-specified and logically insufficient for the heterogeneity conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of autonomous driving (AD) software and hardware systems. It reviews sensor inputs, datasets, simulators, the modular and end-to-end software architectures, and commercial hardware platforms including GPUs, FPGAs, and SoCs such as Tesla FSD, NVIDIA DRIVE Orin, and Mobileye EyeQ6. It contributes a small benchmark study (Fig. 8) measuring FPS and FPS/W of three end-to-end models—TransFuser, InterFuser, and MILE—on four CPU/GPU systems, plus a layer-level arithmetic-intensity analysis of InterFuser (Fig. 9). On this basis, Sections V.B–V.D argue that single-type homogeneous hardware is suboptimal and that future AD accelerators will be increasingly heterogeneous, combining CPUs, GPUs, task-specific accelerators including PIM, and programmable logic such as FPGAs or CGRAs.","tokens_in":22161,"tokens_out":5439,"duration_ms":49832,"significance":"If its central thesis is accepted, the paper provides a useful organizing perspective for AD hardware design, aligning with industry trends in heterogeneous SoCs. Its main strength is the broad synthesis of current software and hardware stacks with specific commercial examples, and the attempt to ground a hardware argument in a small set of measurements rather than speculation. The benchmark data, however, are not currently sufficient to carry the heterogeneity conclusion: the measurements lack methodological detail and error bars, the '10 FPS' threshold is unsourced, and the test models do not span the diversity of the software stack described earlier. The qualitative conclusion is defensible and consistent with the literature, but the empirical support needs substantial strengthening. The paper is likely to be useful to practitioners and researchers entering the area, but in its present form it does not meet the standard for a fully supported experimental claim.","major_comments":[{"comment":"The quantitative claim in Section V.A that 'none of the CPUs solely—except for one with borderline results—can meet the minimum required performance of 10 FPS' is not supported by the information provided. There is no description of the measurement methodology: batch size, input resolution, precision, inference framework and version, CPU/GPU power states, number of runs, or the method used to compute FPS/W are all absent. The '10 FPS' threshold is asserted without any citation or derivation, and the single 'borderline' CPU data point cannot be interpreted without confidence intervals or run-to-run variance. Because this result is later cited in Section V.B as the 'brief demonstration' that single-type hardware is suboptimal, the missing methodology is load-bearing for the paper's hardware argument.","section":"§V.A, Fig. 8, Table VII"},{"comment":"The comparison labeled CPU-only versus GPU-only is confounded by the system configurations. System 1 is a Jetson AGX Orin, which is a heterogeneous SoC containing both CPU and GPU on the same die, and System 2 is a laptop-class system with both a Ryzen 9 CPU and an RTX 3060 GPU. The paper does not state how 'CPU-only' execution was isolated (e.g., whether the GPU was disabled), how the GPU measurements were taken on the same systems, or how power and thermal sharing between the CPU and GPU affected the measurements. Without this information, the direct CPU-versus-GPU comparison in Fig. 8 is not reproducible and its validity is unclear.","section":"§V.A, Table VII"},{"comment":"The logical link from Fig. 8 to the heterogeneity thesis is incomplete. The benchmark compares different homogeneous CPU systems against different homogeneous GPU systems; it does not compare a homogeneous GPU-only configuration with a heterogeneous CPU+GPU+FPGA/CGRA/PIM configuration on the same workloads. Showing that these CPUs are slower than these GPUs on three end-to-end models does not demonstrate that 'managing these diverse models with a single type of hardware leads to suboptimal performance.' The paper should either add a direct homogeneous-versus-heterogeneous comparison or substantially soften the claim, explicitly stating that Fig. 8 supports only the narrower observation that the evaluated CPUs are less performant than the evaluated GPUs on these models.","section":"§V.B, §V.D"},{"comment":"The selection of TransFuser, InterFuser, and MILE is not justified as representative of the autonomous driving software stack surveyed in Section III. These three models are all transformer-based, camera+LiDAR, end-to-end driving models; they do not cover the diverse workloads described earlier, such as 2D/3D object detection CNNs (YOLO, VoxelNet, PointPillars), point-cloud networks, tracking, trajectory-prediction GNNs, or planning algorithms. The paper's conclusion that future accelerators must handle 'diverse computational and memory requirements' relies on this representativeness, but no argument or evidence is given that the three chosen models span that diversity. Without such justification, the empirical results in Fig. 8 and the layer analysis in Fig. 9 (which is only for InterFuser) cannot be generalized to the full AD stack.","section":"§V.A, Table VI"}],"minor_comments":[{"comment":"The KITTI dataset is cited as reference [17], which is actually the Contraction Hierarchies routing paper (Geisberger et al.); a proper citation for the KITTI vision benchmark suite is missing or mis-numbered.","section":"§II.B"},{"comment":"The text states that TransFuser uses 'ResNets [84] and RegNets [85]', but reference [85] is a model-predictive motion planner for the IARA car; the RegNet backbone should be cited to Radosavovic et al., which is already reference [90].","section":"§III.B.1"},{"comment":"The device labels CPU_1 through GPU_4 in Fig. 8 are not mapped directly to the System 1–4 rows of Table VII, making the figure hard to interpret; adding a legend or using the system names directly would improve clarity.","section":"Fig. 8, Table VII"},{"comment":"There is a typo: 'To set he stage for this' should read 'To set the stage for this.'","section":"§I (page 2)"},{"comment":"The sentence 'Worth noting is that these results indicate even a single GPU can exhibit varying performance and efficiency across different models' is grammatically awkward and should be rephrased for clarity.","section":"§V.A"},{"comment":"The statement that projections indicate autonomous vehicles 'will dominate 95% of the market by 2050' is an imprecise reading of reference [103], which is primarily about emissions from onboard computing; the claim should be reworded to match the source's actual projection (e.g., vehicle-miles traveled share) or removed.","section":"§V.A"}],"recommendation":"major_revision","confidential_remarks":"The survey is broad and timely, and the qualitative picture it paints is consistent with current industry directions. The empirical component in Section V.A is the part that would differentiate this work from other surveys, but in its present form it is not reproducible and its logical connection to the heterogeneity thesis is weak. The citation errors (KITTI, RegNet) suggest the reference list needs a careful pass. I believe the central claim is defensible and the paper can be brought to an acceptable standard with a major revision that adds measurement methodology, error bars, a sourced threshold, and a more careful framing of what Fig. 8 does and does not show."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a worthwhile survey for hardware designers entering autonomous driving, and the authors did real work by benchmarking three end-to-end models on four CPU/GPU systems and profiling InterFuser's layers. But the benchmark is the weakest part of the paper, and the way it is used to support the heterogeneity conclusion is logically under-powered.\n\nWhat's new: the FPS and FPS/W numbers for TransFuser, InterFuser, and MILE on those systems, plus the layer-wise arithmetic intensity plots, are original. The survey itself covers a lot of ground—sensors, datasets, simulators, modular and end-to-end software, commercial SoCs, and PIM research—and the descriptions match the cited literature. The tables summarizing hardware platforms are handy. The paper is clearly written and honest about being speculative in the final section.\n\nWhere it's soft: the experimental methodology is almost absent. No batch size, input resolution, precision, framework version, measurement protocol, or power measurement detail. There are no error bars, and the '10 FPS minimum' is asserted without a citation. Given that, the claim that 'none of the CPUs solely—except one with borderline results—can meet 10 FPS' is fragile; one borderline data point without confidence intervals can't carry that statement. More importantly, the CPU-vs-GPU comparison doesn't demonstrate that a single homogeneous accelerator is inadequate, nor that heterogeneity is needed. It shows these GPUs beat these CPUs on these three models. The paper's Section V.B uses this as 'briefly demonstrated' evidence for suboptimality of single-type hardware, which overstates what the data can support. The future-architecture prediction is plausible and consensus-aligned, but it should be framed as a qualitative argument, not a conclusion forced by Fig. 8.\n\nThe survey portion deserves credit. This is not a case of flawed science across the board; it's a case where the original empirical contribution needs to be either strengthened or explicitly de-emphasized. If the authors add methodology, confidence intervals, a source for the FPS threshold, and soften the logical claim, the paper is a solid contribution.\n\nFor peer review: yes, I'd send it out. It is a competent survey with an original data point that a referee can push on. It should not be desk-rejected. My recommendation: accept with major revision, with the benchmark reframed as a limited case study.","headline":"A competent, useful survey whose original CPU/GPU benchmark is too under-specified to carry the heterogeneity thesis; worth refereeing with requests for methodology and reframing.","tokens_in":22646,"tokens_out":2658,"would_cite":true,"duration_ms":24999,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that no single hardware type can efficiently run the full autonomous-driving stack, so future accelerators will be heterogeneous: CPUs and GPUs plus task-specific cores and programmable FPGAs or CGRAs.","keywords":["autonomous driving","hardware acceleration","heterogeneous architecture","processing-in-memory","end-to-end driving models","GPU vs CPU performance","deep learning systems","self-driving SoCs"],"falsifier":"A head-to-head study that runs the full modular stack (perception, prediction, planning) alongside the three end-to-end models on the same four CPUs and four GPUs, with a sourced real-time threshold, and finds a single device class meeting all latency and energy budgets would directly weaken the claim that heterogeneous hardware is required.","tokens_in":21666,"feed_emoji":"🚗","tokens_out":7136,"duration_ms":63929,"temperature":0.7,"pith_summary":"This paper is a survey of autonomous-driving software and hardware with a forward-looking argument: the compute platform that will support high-level autonomy cannot be a single device type. Using three end-to-end driving models (TransFuser, InterFuser, MILE) measured on four CPUs and four GPUs, it reports that nearly all CPUs fall below the 10 FPS real-time bar while GPUs deliver higher throughput and better FPS/W. From the wide variation in arithmetic intensity across tasks and layers, it concludes that one kind of hardware is suboptimal and that future self-driving accelerators will be heterogeneous—general-purpose CPUs and GPUs combined with task-specific PIM-based cores and programmable FPGAs or CGRAs. A sympathetic reader would care because the choice of accelerator architecture directly determines whether level-4/5 autonomy can meet latency, energy, and adaptability requirements in a vehicle that must last 10–15 years.","feed_headline":"Future self-driving computers should combine GPU, CPU, and memory-based cores","feed_subtitle":"One chip class can't handle full autonomy; the answer is specialized cores plus programmable hardware.","key_machinery":"The argument is carried by workload-diversity characterization. The paper defines arithmetic intensity (FLOPs per byte, equivalently FLOPs per memory operation) for the three end-to-end benchmarks and plots layer-level intensity for InterFuser (Fig. 9), showing some layers are compute-bound and others memory-bound. This intensity spread is the mechanism that motivates heterogeneous accelerators and processing-in-memory, since memory-bound layers benefit from computation placed near or inside DRAM while compute-bound layers benefit from parallel SIMD-style engines.","core_discovery":"The central claim is that the future of self-driving accelerators lies in heterogeneous architectures rather than a single dominant device class. The paper reaches this through workload diversity: the evaluated end-to-end models have markedly different parameter counts, FLOPs, and arithmetic intensities, and within a single model such as InterFuser, individual layers range over several orders of magnitude in FLOPs-per-memory-operation. It reports that on the four CPU systems almost none of the end-to-end models can sustain the required 10 FPS, while GPUs both meet and exceed it with better energy efficiency, and it argues that no single hardware type can be optimal across such a spread. The conclusion is therefore a multi-core SoC that mixes general-purpose CPUs/GPUs, task-specific accelerators including processing-in-memory cores, and programmable components such as FPGAs or CGRAs to preserve long-term adaptability.","pith_inferences":["Editorial inference: the paper's proof-of-concept measurements would be much stronger if repeated on a standard benchmark suite covering full modular stacks (perception, prediction, planning) and end-to-end policies, since the chosen three models sample only part of the software space.","Editorial inference: the arithmetic-intensity evidence suggests the first PIM deployments in an autonomous vehicle would target LiDAR point-cloud preprocessing, sensor-fusion attention layers, and other low-intensity layers, but the paper does not commit to a specific placement.","Editorial inference: the 10 FPS threshold is treated as given, yet it is load-bearing; an independently sourced real-time requirement could shift the CPU-vs-GPU conclusion and deserves explicit validation."],"forward_implications":["If the heterogeneity claim is right, next-generation automotive SoCs will need to co-design general-purpose CPU/GPU cores with task-specific accelerators and programmable fabric rather than relying on a single accelerator type.","Memory-bound layers identified by low arithmetic intensity become the natural targets for processing-in-memory cores, while compute-bound layers can stay on GPU or neural-network accelerators.","Because autonomous vehicles have 10- to 15-year lifespans, the winning hardware platforms will include programmable elements (FPGAs or CGRAs) so they can absorb software updates and new models.","The CPU-only path to level-4/5 autonomy becomes untenable if the 10 FPS threshold and the measured CPU results hold, reinforcing GPUs as the baseline and specialized accelerators as the next step."],"supporting_citations":[{"why":"Supplies TransFuser, one of the three end-to-end models whose FPS and FPS/W measurements anchor the CPU-vs-GPU comparison.","marker":"[76]"},{"why":"Supplies InterFuser, the second benchmark model and the source of the layer-level arithmetic-intensity data in Fig. 9.","marker":"[77]"},{"why":"Supplies MILE, the third benchmark model used in the scaling evaluation.","marker":"[102]"},{"why":"Documents Tesla's FSD SoC with CPUs, GPU, and NNAs as an industry example of heterogeneous autonomous-driving hardware.","marker":"[93]"},{"why":"Documents Mobileye's EyeQ heterogeneous SoC combining CPUs, GPU, accelerators, and CGRA-like arrays.","marker":"[99]"},{"why":"Provides the FPGA-vs-GPU comparison on Apollo perception tasks that shows specialized programmable hardware can beat GPUs in some tasks.","marker":"[100]"},{"why":"Gives the Pony.ai example of an FPGA handling sensor data to offload the CPU, supporting the heterogeneous-system argument.","marker":"[92]"},{"why":"Describes Neurocube, a near-memory PIM accelerator that demonstrates high-bandwidth neural-network inference.","marker":"[108]"},{"why":"Reports TETRIS's 4.1x performance and 1.5x energy gains over a 2D NN accelerator, providing the quantitative PIM motivation.","marker":"[110]"},{"why":"Shows Ambit's in-DRAM bitwise operations, the base technique that later PIM accelerators build on.","marker":"[112]"}],"fun_headline_variants":["Self-driving chips need CPU, GPU, and memory cores","Autonomy demands a mix of processors, not one chip","Heterogeneous hardware is the key to self-driving","Future self-driving hardware: mix CPU, GPU, and memory","One chip can't handle autonomy; mix cores instead"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on assuming the three end-to-end models and the 10 FPS threshold fairly represent what autonomous-driving hardware must run, so the Figure 8 measurements can stand in for the full software stack.","fun_headline_variants_meta":{"raw":{"variants":["Self-driving chips need CPU, GPU, and memory cores","Autonomy demands a mix of processors, not one chip","Heterogeneous hardware is the key to self-driving","Future self-driving hardware: mix CPU, GPU, and memory","One chip can't handle autonomy; mix cores instead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000681,"raw_usage":{"total_tokens":3111,"prompt_tokens":979,"completion_tokens":2132,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":2052}},"tokens_in":595,"tokens_out":2132,"duration_ms":13308,"temperature":1.0,"reasoning_tokens":2052,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:46:43.506999+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A head-to-head study that runs the full modular stack (perception, prediction, planning) alongside the three end-to-end models on the same four CPUs and four GPUs, with a sourced real-time threshold, and finds a single device class meeting all latency and energy budgets would directly weaken the claim that heterogeneous hardware is required.","supporting_citations":[{"cited_title":"Trans- fuser: Imitation with transformer-based sensor fusion for autonomous driving,","cited_arxiv_id":null,"evidence_quote":"Supplies TransFuser, one of the three end-to-end models whose FPS and FPS/W measurements anchor the CPU-vs-GPU comparison."},{"cited_title":"Safety-enhanced autonomous driving using interpretable sensor fusion transformer,","cited_arxiv_id":null,"evidence_quote":"Supplies InterFuser, the second benchmark model and the source of the layer-level arithmetic-intensity data in Fig. 9."},{"cited_title":"Compute solution for tesla’s full self-driving computer,","cited_arxiv_id":null,"evidence_quote":"Documents Tesla's FSD SoC with CPUs, GPU, and NNAs as an industry example of heterogeneous autonomous-driving hardware."},{"cited_title":"Eyeq chip technology","cited_arxiv_id":null,"evidence_quote":"Documents Mobileye's EyeQ heterogeneous SoC combining CPUs, GPU, accelerators, and CGRA-like arrays."},{"cited_title":"Hardware acceleration of deep neural networks for autonomous driving on fpga- based soc,","cited_arxiv_id":null,"evidence_quote":"Provides the FPGA-vs-GPU comparison on Apollo perception tasks that shows specialized programmable hardware can beat GPUs in some tasks."},{"cited_title":"Accelerating the pony.ai av sensor data pro- cessing pipeline","cited_arxiv_id":null,"evidence_quote":"Gives the Pony.ai example of an FPGA handling sensor data to offload the CPU, supporting the heterogeneous-system argument."},{"cited_title":"Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory,","cited_arxiv_id":null,"evidence_quote":"Describes Neurocube, a near-memory PIM accelerator that demonstrates high-bandwidth neural-network inference."},{"cited_title":"Tetris: Scalable and efficient neural network acceleration with 3d memory,","cited_arxiv_id":null,"evidence_quote":"Reports TETRIS's 4.1x performance and 1.5x energy gains over a 2D NN accelerator, providing the quantitative PIM motivation."},{"cited_title":"Ambit: In- memory accelerator for bulk bitwise operations using commodity dram technology,","cited_arxiv_id":null,"evidence_quote":"Shows Ambit's in-DRAM bitwise operations, the base technique that later PIM accelerators build on."}],"review_version":1}