Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Moving Forward: A Review of Autonomous Driving Software and Hardware Systems

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper argues that no single hardware type can efficiently run the full autonomous-driving stack, so future accelerators will be heterogeneous: CPUs and GPUs plus task-specific cores and programmable FPGAs or CGRAs.

desk verdict A competent, useful survey whose original CPU/GPU benchmark is too under-specified to carry the heterogeneity thesis; worth refereeing with requests for methodology and reframing. read the letter →

arxiv 2411.10291 v1 pith:SAAVA7D5 submitted 2024-11-15 cs.RO

classification cs.RO
keywords autonomousdrivinghardwareaccelerationheterogeneousarchitectureprocessing-in-memoryend-to-endmodelsGPUvsCPUperformancedeeplearningsystemsself-drivingSoCs
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a survey of autonomous-driving software and hardware with a forward-looking argument: the compute platform that will support high-level autonomy cannot be a single device type. Using three end-to-end driving models (TransFuser, InterFuser, MILE) measured on four CPUs and four GPUs, it reports that nearly all CPUs fall below the 10 FPS real-time bar while GPUs deliver higher throughput and better FPS/W. From the wide variation in arithmetic intensity across tasks and layers, it concludes that one kind of hardware is suboptimal and that future self-driving accelerators will be heterogeneous—general-purpose CPUs and GPUs combined with task-specific PIM-based cores and programmable FPGAs or CGRAs. A sympathetic reader would care because the choice of accelerator architecture directly determines whether level-4/5 autonomy can meet latency, energy, and adaptability requirements in a vehicle that must last 10–15 years.

What carries the argument

The argument is carried by workload-diversity characterization. The paper defines arithmetic intensity (FLOPs per byte, equivalently FLOPs per memory operation) for the three end-to-end benchmarks and plots layer-level intensity for InterFuser (Fig. 9), showing some layers are compute-bound and others memory-bound. This intensity spread is the mechanism that motivates heterogeneous accelerators and processing-in-memory, since memory-bound layers benefit from computation placed near or inside DRAM while compute-bound layers benefit from parallel SIMD-style engines.

What would settle it

A head-to-head study that runs the full modular stack (perception, prediction, planning) alongside the three end-to-end models on the same four CPUs and four GPUs, with a sourced real-time threshold, and finds a single device class meeting all latency and energy budgets would directly weaken the claim that heterogeneous hardware is required.

Watch

Extended reading notes

Core claim

The central claim is that the future of self-driving accelerators lies in heterogeneous architectures rather than a single dominant device class. The paper reaches this through workload diversity: the evaluated end-to-end models have markedly different parameter counts, FLOPs, and arithmetic intensities, and within a single model such as InterFuser, individual layers range over several orders of magnitude in FLOPs-per-memory-operation. It reports that on the four CPU systems almost none of the end-to-end models can sustain the required 10 FPS, while GPUs both meet and exceed it with better energy efficiency, and it argues that no single hardware type can be optimal across such a spread. The conclusion is therefore a multi-core SoC that mixes general-purpose CPUs/GPUs, task-specific accelerators including processing-in-memory cores, and programmable components such as FPGAs or CGRAs to preserve long-term adaptability.

Load-bearing premise

The argument rests on assuming the three end-to-end models and the 10 FPS threshold fairly represent what autonomous-driving hardware must run, so the Figure 8 measurements can stand in for the full software stack.

Editorial extensions

If this is right

  • If the heterogeneity claim is right, next-generation automotive SoCs will need to co-design general-purpose CPU/GPU cores with task-specific accelerators and programmable fabric rather than relying on a single accelerator type.
  • Memory-bound layers identified by low arithmetic intensity become the natural targets for processing-in-memory cores, while compute-bound layers can stay on GPU or neural-network accelerators.
  • Because autonomous vehicles have 10- to 15-year lifespans, the winning hardware platforms will include programmable elements (FPGAs or CGRAs) so they can absorb software updates and new models.
  • The CPU-only path to level-4/5 autonomy becomes untenable if the 10 FPS threshold and the measured CPU results hold, reinforcing GPUs as the baseline and specialized accelerators as the next step.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the paper's proof-of-concept measurements would be much stronger if repeated on a standard benchmark suite covering full modular stacks (perception, prediction, planning) and end-to-end policies, since the chosen three models sample only part of the software space.
  • Editorial inference: the arithmetic-intensity evidence suggests the first PIM deployments in an autonomous vehicle would target LiDAR point-cloud preprocessing, sensor-fusion attention layers, and other low-intensity layers, but the paper does not commit to a specific placement.
  • Editorial inference: the 10 FPS threshold is treated as given, yet it is load-bearing; an independently sourced real-time requirement could shift the CPU-vs-GPU conclusion and deserves explicit validation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript is a survey of autonomous driving (AD) software and hardware systems. It reviews sensor inputs, datasets, simulators, the modular and end-to-end software architectures, and commercial hardware platforms including GPUs, FPGAs, and SoCs such as Tesla FSD, NVIDIA DRIVE Orin, and Mobileye EyeQ6. It contributes a small benchmark study (Fig. 8) measuring FPS and FPS/W of three end-to-end models—TransFuser, InterFuser, and MILE—on four CPU/GPU systems, plus a layer-level arithmetic-intensity analysis of InterFuser (Fig. 9). On this basis, Sections V.B–V.D argue that single-type homogeneous hardware is suboptimal and that future AD accelerators will be increasingly heterogeneous, combining CPUs, GPUs, task-specific accelerators including PIM, and programmable logic such as FPGAs or CGRAs.

Significance. If its central thesis is accepted, the paper provides a useful organizing perspective for AD hardware design, aligning with industry trends in heterogeneous SoCs. Its main strength is the broad synthesis of current software and hardware stacks with specific commercial examples, and the attempt to ground a hardware argument in a small set of measurements rather than speculation. The benchmark data, however, are not currently sufficient to carry the heterogeneity conclusion: the measurements lack methodological detail and error bars, the '10 FPS' threshold is unsourced, and the test models do not span the diversity of the software stack described earlier. The qualitative conclusion is defensible and consistent with the literature, but the empirical support needs substantial strengthening. The paper is likely to be useful to practitioners and researchers entering the area, but in its present form it does not meet the standard for a fully supported experimental claim.

major comments (4)
  1. [§V.A, Fig. 8, Table VII] The quantitative claim in Section V.A that 'none of the CPUs solely—except for one with borderline results—can meet the minimum required performance of 10 FPS' is not supported by the information provided. There is no description of the measurement methodology: batch size, input resolution, precision, inference framework and version, CPU/GPU power states, number of runs, or the method used to compute FPS/W are all absent. The '10 FPS' threshold is asserted without any citation or derivation, and the single 'borderline' CPU data point cannot be interpreted without confidence intervals or run-to-run variance. Because this result is later cited in Section V.B as the 'brief demonstration' that single-type hardware is suboptimal, the missing methodology is load-bearing for the paper's hardware argument.
  2. [§V.A, Table VII] The comparison labeled CPU-only versus GPU-only is confounded by the system configurations. System 1 is a Jetson AGX Orin, which is a heterogeneous SoC containing both CPU and GPU on the same die, and System 2 is a laptop-class system with both a Ryzen 9 CPU and an RTX 3060 GPU. The paper does not state how 'CPU-only' execution was isolated (e.g., whether the GPU was disabled), how the GPU measurements were taken on the same systems, or how power and thermal sharing between the CPU and GPU affected the measurements. Without this information, the direct CPU-versus-GPU comparison in Fig. 8 is not reproducible and its validity is unclear.
  3. [§V.B, §V.D] The logical link from Fig. 8 to the heterogeneity thesis is incomplete. The benchmark compares different homogeneous CPU systems against different homogeneous GPU systems; it does not compare a homogeneous GPU-only configuration with a heterogeneous CPU+GPU+FPGA/CGRA/PIM configuration on the same workloads. Showing that these CPUs are slower than these GPUs on three end-to-end models does not demonstrate that 'managing these diverse models with a single type of hardware leads to suboptimal performance.' The paper should either add a direct homogeneous-versus-heterogeneous comparison or substantially soften the claim, explicitly stating that Fig. 8 supports only the narrower observation that the evaluated CPUs are less performant than the evaluated GPUs on these models.
  4. [§V.A, Table VI] The selection of TransFuser, InterFuser, and MILE is not justified as representative of the autonomous driving software stack surveyed in Section III. These three models are all transformer-based, camera+LiDAR, end-to-end driving models; they do not cover the diverse workloads described earlier, such as 2D/3D object detection CNNs (YOLO, VoxelNet, PointPillars), point-cloud networks, tracking, trajectory-prediction GNNs, or planning algorithms. The paper's conclusion that future accelerators must handle 'diverse computational and memory requirements' relies on this representativeness, but no argument or evidence is given that the three chosen models span that diversity. Without such justification, the empirical results in Fig. 8 and the layer analysis in Fig. 9 (which is only for InterFuser) cannot be generalized to the full AD stack.
minor comments (6)
  1. [§II.B] The KITTI dataset is cited as reference [17], which is actually the Contraction Hierarchies routing paper (Geisberger et al.); a proper citation for the KITTI vision benchmark suite is missing or mis-numbered.
  2. [§III.B.1] The text states that TransFuser uses 'ResNets [84] and RegNets [85]', but reference [85] is a model-predictive motion planner for the IARA car; the RegNet backbone should be cited to Radosavovic et al., which is already reference [90].
  3. [Fig. 8, Table VII] The device labels CPU_1 through GPU_4 in Fig. 8 are not mapped directly to the System 1–4 rows of Table VII, making the figure hard to interpret; adding a legend or using the system names directly would improve clarity.
  4. [§I (page 2)] There is a typo: 'To set he stage for this' should read 'To set the stage for this.'
  5. [§V.A] The sentence 'Worth noting is that these results indicate even a single GPU can exhibit varying performance and efficiency across different models' is grammatically awkward and should be rephrased for clarity.
  6. [§V.A] The statement that projections indicate autonomous vehicles 'will dominate 95% of the market by 2050' is an imprecise reading of reference [103], which is primarily about emissions from onboard computing; the claim should be reworded to match the source's actual projection (e.g., vehicle-miles traveled share) or removed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's benchmark and survey reasoning are externally grounded, and no prediction reduces to its inputs by construction.

full rationale

This paper is a survey with a speculative forward-looking conclusion, not a derivation chain in which outputs are fitted from inputs. Section V.A reports empirical FPS and FPS/W measurements of three published end-to-end models (TransFuser, InterFuser, MILE) on four CPU and four GPU systems. These measurements are not used to fit any parameter that is then renamed as a prediction; the numbers are external benchmark observations, and the conclusion that GPUs outperform CPUs on these workloads follows directly from the reported measurements. The central Section V.D claim, that future self-driving accelerators will be heterogeneous, is justified by the diverse computational and memory requirements of the surveyed software stack, by the architectural diversity of existing systems (Tesla FSD, DRIVE Orin, EyeQ6), and by qualitative arguments about sparsity and arithmetic intensity. Even though the CPU-vs-GPU comparison does not logically prove that heterogeneous architectures are necessary, that is an evidentiary gap or a weakness in inductive support, not circularity. There are no load-bearing self-citations: the references are to external prior work, and no uniqueness theorem or prior result by the same authors is invoked to force a conclusion. The unsourced 10 FPS threshold and missing measurement methodology are rigor concerns, but they do not make the argument circular. The paper is therefore self-contained against external benchmarks, and the appropriate circularity score is 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central argument rests on three unstated assumptions about the benchmark: the selected models are representative, TDP approximates power, and the FPS measurements are reliable. The hand-chosen 10 FPS threshold is the only free parameter. No invented entities are introduced.

free parameters (1)
  • Minimum 10 FPS performance threshold = 10 FPS
    Section V.A sets 10 FPS as the minimum required performance for self-driving systems and uses it to judge CPUs; no citation or derivation is given, so it is a hand-chosen threshold that drives the conclusion that CPUs are insufficient.
assumptions (3)
  • ad hoc to paper The three end-to-end models TransFuser, InterFuser, and MILE are representative of the computational and memory demands of autonomous driving software stacks.
    Section V.A selects these three models without argument that they span perception, prediction, planning, and control workloads in real AD systems.
  • domain assumption TDP values from hardware specifications are an adequate proxy for power consumption when computing FPS/W efficiency.
    Fig. 8(b) labels efficiency as FPS/W using TDP values from Table VII; no measured power is reported.
  • domain assumption The reported FPS measurements are accurate despite the absence of a measurement methodology.
    The paper does not describe input resolution, batch size, precision, software versions, or number of runs behind Fig. 8.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Moving Forward: A Review of Autonomous Driving Software and Hardware Systems." pith.science (2026). https://pith.science/paper/SAAVA7D5

@misc{pith2026241110291,
  author       = {Pith},
  title        = {Pith review of: Moving Forward: A Review of Autonomous Driving Software and Hardware Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SAAVA7D5}},
  note         = {Machine review of arXiv:2411.10291}
}
read the original abstract

With their potential to significantly reduce traffic accidents, enhance road safety, optimize traffic flow, and decrease congestion, autonomous driving systems are a major focus of research and development in recent years. Beyond these immediate benefits, they offer long-term advantages in promoting sustainable transportation by reducing emissions and fuel consumption. Achieving a high level of autonomy across diverse conditions requires a comprehensive understanding of the environment. This is accomplished by processing data from sensors such as cameras, radars, and LiDARs through a software stack that relies heavily on machine learning algorithms. These ML models demand significant computational resources and involve large-scale data movement, presenting challenges for hardware to execute them efficiently and at high speed. In this survey, we first outline and highlight the key components of self-driving systems, covering input sensors, commonly used datasets, simulation platforms, and the software architecture. We then explore the underlying hardware platforms that support the execution of these software systems. By presenting a comprehensive view of autonomous driving systems and their increasing demands, particularly for higher levels of autonomy, we analyze the performance and efficiency of scaled-up off-the-shelf GPU/CPU-based systems, emphasizing the challenges within the computational components. Through examples showcasing the diverse computational and memory requirements in the software stack, we demonstrate how more specialized hardware and processing closer to memory can enable more efficient execution with lower latency. Finally, based on current trends and future demands, we conclude by speculating what a future hardware platform for autonomous driving might look like.

Figures

Figures reproduced from arXiv: 2411.10291 by the authors.

Figure 1
Figure 1. Autonomous driving system modules. A high-level simplified view of an autonomous driving system is depicted in [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. System diagram of: (a) generic modular system, (b) end-to-end system. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. The Apollo system-driving framework based on version 9.0 with only perception breakdown for clarity. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: An overview of transformer-based TransFuser archi [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: An overview of transformer-based InterFuser architec [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Simplified system diagram of Tesla’s autonomous system. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: (a) Conventional CPU+GPU-based system, (b) Integrating an FPGA as a data processor on CPU+GPU-based system [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: (a) Performance (in terms of Frame Per Seconds) of the three end-to-end systems on various devices (b) Efficiency (in [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: (a) The arithmetic intensity of sample layers in the InterFuser AD system versus memory operations (reading data from [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Architecture of TETRIS accelerator adopted from [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Open-Source Autonomous Driving Software Platforms: Comparison of Autoware and Apollo

    cs.RO 2025-01 conditional novelty 4.0 of 10

    Autoware and Apollo differ in module design, and Apollo's shared-memory middleware is faster but more memory-hungry than Autoware's serialized DDS middleware.

Reference graph

Works this paper leans on

114 extracted references · 67 canonical work pages · cited by 1 Pith paper

  1. [1]

    Evaluation of fuel consumption and emissions benefits of connected and automated vehicles in mixed traffic flow,

    H. Li, H. Li, Y . Hu, T. Xia, Q. Miao, and J. Chu, “Evaluation of fuel consumption and emissions benefits of connected and automated vehicles in mixed traffic flow,” Frontiers in Energy Research, vol. 11, p. 1207449, 2023

  2. [2]

    Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,

    S. International, “Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,” 2021

  3. [3]

    Autonomous vehicle decision-making and control in complex and unconventional scenar- ios—a review,

    F. Sana, N. L. Azad, and K. Raahemifar, “Autonomous vehicle decision-making and control in complex and unconventional scenar- ios—a review,” Machines, vol. 11, no. 7, p. 676, 2023

  4. [4]

    A survey of autonomous driving: Common practices and emerging technologies,

    E. Yurtsever, J. Lambert, A. Carballo, and K. Takeda, “A survey of autonomous driving: Common practices and emerging technologies,” IEEE Access, vol. 8, pp. 58443–58469, 2019

  5. [5]

    Autonomous driving in urban environments: Boss and the urban challenge,

    C. Urmson, J. Anhalt, J. A. Bagnell, C. R. Baker, R. Bittner, M. N. Clark, J. M. Dolan, D. Duggins, T. Galatali, C. Geyer, M. Gittle- man, S. Harbaugh, M. Hebert, T. M. Howard, S. Kolski, A. Kelly, M. Likhachev, M. McNaughton, N. Miller, K. M. Peterson, B. Pilnick, R. R. Rajkumar, P. E. Rybski, B. Salesky, Y .-W. Seo, S. Singh, J. M. Snider, A. Stentz, W....

  6. [6]

    Self-driving cars: A survey,

    C. S. Badue, R. Guidolini, R. V . Carneiro, P. Azevedo, V . B. Cardoso, A. Forechi, L. F. R. Jesus, R. Berriel, T. M. Paix ˜ao, F. W. Mutz, T. Oliveira-Santos, and A. F. de Souza, “Self-driving cars: A survey,” ArXiv, vol. abs/1901.04407, 2019

  7. [7]

    Introducing the 5th-generation waymo driver: Informed by experience, designed for scale, engineered to tackle more environments

    “Introducing the 5th-generation waymo driver: Informed by experience, designed for scale, engineered to tackle more environments.” https: //waymo.com/blog/2020/03/introducing-5th-generation-waymo-driver. html. Online

  8. [8]

    Apollo: Open source autonomous driving

    “Apollo: Open source autonomous driving.” https://github.com/ ApolloAuto/apollo. Online

Show all 114 references
  1. [9]

    Autopilot

    I. Tesla, “Autopilot.” https://www.tesla.com/autopilot, 2023. Online

  2. [10]

    Environmental-driven approach towards level 5 self-driving,

    M. Hurair, J. Ju, and J. Han, “Environmental-driven approach towards level 5 self-driving,” Sensors, vol. 24, no. 2, p. 485, 2024

  3. [11]

    A novel approach for detecting road based on two-stream fusion fully convolutional network,

    X. Lv, Z. yi Liu, J. Xin, and N. Zheng, “A novel approach for detecting road based on two-stream fusion fully convolutional network,” 2018 IEEE Intelligent Vehicles Symposium (IV) , pp. 1464–1469, 2018

  4. [12]

    The pascal visual object classes (voc) challenge,

    M. Everingham, L. V . Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International Journal of Computer Vision , vol. 88, pp. 303–338, 2010

  5. [13]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” 2009 IEEE Conference on Computer Vision and Pattern Recognition , pp. 248–255, 2009

  6. [14]

    Microsoft coco: Common objects in context,

    T.-Y . Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision , 2014

  7. [15]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009

  8. [16]

    The cityscapes dataset for semantic urban scene understanding,

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 3213–3223, 2016

  9. [17]

    Exact routing in large road networks using contraction hierarchies,

    R. Geisberger, P. Sanders, D. Schultes, and C. Vetter, “Exact routing in large road networks using contraction hierarchies,”Transp. Sci., vol. 46, pp. 388–404, 2012

  10. [18]

    V oxelnet: End-to-end learning for point cloud based 3d object detection,

    Y . Zhou and O. Tuzel, “V oxelnet: End-to-end learning for point cloud based 3d object detection,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 4490–4499, 2017

  11. [19]

    Pointpillars: Fast encoders for object detection from point clouds,

    A. H. Lang, S. V ora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), pp. 12689–12697, 2018

  12. [20]

    Multi-view 3d object detection network for autonomous driving,

    X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object detection network for autonomous driving,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 6526–6534, 2016

  13. [21]

    Joint 3d proposal generation and object detection from view aggregation,

    J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. L. Waslander, “Joint 3d proposal generation and object detection from view aggregation,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 1–8, 2017

  14. [22]

    Scalability in percep- tion for autonomous driving: Waymo open dataset,

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine, V . Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. M. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y . Zhang, J. Shlens, Z. Chen, and D. Anguelov, “Scalability in p...

  15. [23]

    Large scale interactive motion forecasting for autonomous driving : The waymo open motion dataset,

    S. M. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y . Chai, B. Sapp, C. Qi, Y . Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V . Vasudevan, A. McCauley, J. Shlens, and D. Anguelov, “Large scale interactive motion forecasting for autonomous driving : The waymo op...

  16. [24]

    Womd-lidar: Raw sensor dataset benchmark for motion forecasting,

    K. Chen, R. Ge, H. Qiu, R. Ai-Rfou, C. Qi, X. Zhou, Z. Yang, S. M. Ettinger, P. Sun, Z. Leng, M. A. Mustafa, I. Bogun, W. Wang, M. Tan, and D. Anguelov, “Womd-lidar: Raw sensor dataset benchmark for motion forecasting,” ArXiv, vol. abs/2304.03834, 2023

  17. [25]

    Autoware

    A. Foundation, “Autoware.” https://github.com/autowarefoundation/ autoware, 2023. Online

  18. [26]

    The apolloscape open dataset for autonomous driving and its application,

    P. Wang, X. Huang, X. Cheng, D. Zhou, Q. Geng, and R. Yang, “The apolloscape open dataset for autonomous driving and its application,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 42, pp. 2702–2719, 2018

  19. [27]

    Lgsvl simulator: A high fidelity simulator for autonomous driving,

    G. Rong, B. H. Shin, H. Tabatabaee, Q. Lu, S. Lemke, M. Mo ˇzeiko, E. Boise, G. Uhm, M. Gerow, S. Mehta, et al. , “Lgsvl simulator: A high fidelity simulator for autonomous driving,” arXiv preprint arXiv:2005.03778, 2020

  20. [28]

    CARLA: An open urban driving simulator,

    A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun, “CARLA: An open urban driving simulator,” in Proceedings of the 1st Annual Conference on Robot Learning , pp. 1–16, 2017

  21. [29]

    Faster r-cnn: Towards real- time object detection with region proposal networks,

    S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster r-cnn: Towards real- time object detection with region proposal networks,” IEEE Transac- tions on Pattern Analysis and Machine Intelligence , vol. 39, pp. 1137– 1149, 2015

  22. [30]

    You only look once: Unified, real-time object detection,

    J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 779–788, 2015

  23. [31]

    Ssd: Single shot multibox detector,

    W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C.-Y . Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European Conference on Computer Vision , 2015

  24. [32]

    Second: Sparsely embedded convolutional detection,

    Y . Yan, Y . Mao, and B. Li, “Second: Sparsely embedded convolutional detection,” Sensors (Basel, Switzerland) , vol. 18, 2018

  25. [33]

    Pointnet: Deep learning on point sets for 3d classification and segmentation,

    C. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 77–85, 2016

  26. [34]

    Pointnet++: Deep hierarchical feature learning on point sets in a metric space,

    C. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in Neural Information Processing Systems, 2017

  27. [35]

    Frustum pointnets for 3d object detection from rgb-d data,

    C. Qi, W. Liu, C. Wu, H. Su, and L. J. Guibas, “Frustum pointnets for 3d object detection from rgb-d data,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 918–927, 2017

  28. [36]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” CoRR, vol. abs/1409.1556, 2014

  29. [37]

    Attention is all you need,

    A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems , 2017

  30. [38]

    Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,

    X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1080–1089, 2022

  31. [39]

    Unifying voxel- based representation with transformer for 3d object detection,

    Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel- based representation with transformer for 3d object detection,” ArXiv, vol. abs/2206.00630, 2022

  32. [40]

    3d object detection for au- tonomous driving: A comprehensive survey,

    J. Mao, S. Shi, X. Wang, and H. Li, “3d object detection for au- tonomous driving: A comprehensive survey,” International Journal of Computer Vision, vol. 131, pp. 1909 – 1963, 2022

  33. [41]

    3d object tracking using rgb and lidar data,

    A. Asvadi, P. Gir ˜ao, P. Peixoto, and U. J. C. Nunes, “3d object tracking using rgb and lidar data,” 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC) , pp. 1255–1260, 2016

  34. [42]

    Fast multiple objects detection and tracking fusing color camera and 3d lidar for intelligent vehicles,

    S. Hwang, N. Kim, Y . Choi, S. Lee, and I.-S. Kweon, “Fast multiple objects detection and tracking fusing color camera and 3d lidar for intelligent vehicles,” 2016 13th International Conference on Ubiquitous Robots and Ambient Intelligence (URAI) , pp. 234–239, 2016

  35. [43]

    Fast r-cnn,

    R. B. Girshick, “Fast r-cnn,” 2015

  36. [44]

    Spin-images: A representation for 3-d surface match- ing,

    A. E. Johnson, “Spin-images: A representation for 3-d surface match- ing,” 1997

  37. [45]

    Using spin images for efficient object recognition in cluttered 3d scenes,

    A. E. Johnson and M. Hebert, “Using spin images for efficient object recognition in cluttered 3d scenes,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 21, pp. 433–449, 1999

  38. [46]

    Yolov3: An incremental improvement,

    J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” ArXiv, vol. abs/1804.02767, 2018

  39. [47]

    Yolo9000: Better, faster, stronger,

    J. Redmon and A. Farhadi, “Yolo9000: Better, faster, stronger,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6517–6525, 2016

  40. [48]

    Yolov4: Optimal speed and accuracy of object detection,

    A. Bochkovskiy, C.-Y . Wang, and H.-Y . M. Liao, “Yolov4: Optimal speed and accuracy of object detection,” ArXiv, vol. abs/2004.10934, 2020

  41. [49]

    Traffic light recognition using convolutional neural networks: A survey,

    S. Pavlitska, N. Lambing, A. K. Bangaru, and J. M. Z ¨ollner, “Traffic light recognition using convolutional neural networks: A survey,” 2023 IEEE 26th International Conference on Intelligent Transportation Systems (ITSC), pp. 2790–2796, 2023

  42. [50]

    Traffic light detection using tensorflow object detection framework,

    T. V . Janahiraman and M. S. M. Subuhan, “Traffic light detection using tensorflow object detection framework,” 2019 IEEE 9th International Conference on System Engineering and Technology (ICSET) , pp. 108– 113, 2019

  43. [51]

    Traffic lights detection and recognition method based on the improved yolov4 algorithm,

    Q. Wang, Q. Zhang, X. Liang, Y . Wang, C. Zhou, and V . I. Mikulovich, “Traffic lights detection and recognition method based on the improved yolov4 algorithm,” Sensors (Basel, Switzerland) , vol. 22, 2021

  44. [52]

    A two-stage framework for diverse traffic light recognition based on individual signal detection,

    S.-Y . Lin and H.-Y . Lin, “A two-stage framework for diverse traffic light recognition based on individual signal detection,” in Mediterranean Conference on Pattern Recognition and Artificial Intelligence , 2021

  45. [53]

    A sensor fusion approach for localization with cumulative error elimination,

    F. Zhang, H. Stahle, G. Chen, C.-W. Chen, C. Simon, C. Buckl, and A. Knoll, “A sensor fusion approach for localization with cumulative error elimination,” 2012 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI) , pp. 1–6, 2012

  46. [54]

    Road marking detection using lidar reflective intensity data and its application to vehicle localization,

    A. Y . Hata and D. F. Wolf, “Road marking detection using lidar reflective intensity data and its application to vehicle localization,” 17th International IEEE Conference on Intelligent Transportation Systems (ITSC), pp. 584–589, 2014

  47. [55]

    Sensor fusion-based low-cost vehicle localization system for complex urban environments,

    J. K. Suhr, J. Jang, D. Min, and H. G. Jung, “Sensor fusion-based low-cost vehicle localization system for complex urban environments,” IEEE Transactions on Intelligent Transportation Systems , vol. 18, pp. 1078–1086, 2017

  48. [56]

    A survey of the state-of-the-art localization techniques and their potentials for autonomous vehicle applications,

    S. Kuutti, S. Fallah, K. V . Katsaros, M. Dianati, F. Mccullough, and A. Mouzakitis, “A survey of the state-of-the-art localization techniques and their potentials for autonomous vehicle applications,” IEEE Inter- net of Things Journal , vol. 5, pp. 829–846, 2018

  49. [57]

    Multimodal trajectory pre- dictions for autonomous driving using deep convolutional networks,

    H. Cui, V . Radosavljevic, F.-C. Chou, T.-H. Lin, T. Nguyen, T.-K. Huang, J. G. Schneider, and N. Djuric, “Multimodal trajectory pre- dictions for autonomous driving using deep convolutional networks,” 2019 International Conference on Robotics and Automation (ICRA) , pp. 2090–...

  50. [58]

    Vectornet: Encoding hd maps and agent dynamics from vectorized representation,

    J. Gao, C. Sun, H. Zhao, Y . Shen, D. Anguelov, C. Li, and C. Schmid, “Vectornet: Encoding hd maps and agent dynamics from vectorized representation,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11522–11530, 2020

  51. [59]

    Bert: Pre- training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre- training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics, 2019

  52. [60]

    Wayformer: Motion forecasting via simple & efficient attention networks,

    N. Nayakanti, R. Al-Rfou, A. Zhou, K. Goel, K. S. Refaat, and B. Sapp, “Wayformer: Motion forecasting via simple & efficient attention networks,” 2023 IEEE International Conference on Robotics and Automation (ICRA) , pp. 2980–2987, 2022

  53. [61]

    Risky action recognition in lane change video clips using deep spatiotemporal networks with segmentation mask transfer,

    E. Yurtsever, Y . Liu, J. Lambert, C. Miyajima, E. Takeuchi, K. Takeda, and J. L. Hansen, “Risky action recognition in lane change video clips using deep spatiotemporal networks with segmentation mask transfer,” 2019 IEEE Intelligent Transportation Systems Conference (ITSC), p...

  54. [62]

    Concrete problems for autonomous vehicle safety: Advantages of bayesian deep learning,

    R. T. McAllister, Y . Gal, A. Kendall, M. van der Wilk, A. Shah, R. Cipolla, and A. Weller, “Concrete problems for autonomous vehicle safety: Advantages of bayesian deep learning,” in International Joint Conference on Artificial Intelligence , 2017

  55. [63]

    Highway hierarchies hasten exact shortest path queries,

    P. Sanders and D. Schultes, “Highway hierarchies hasten exact shortest path queries,” in Embedded Systems and Applications , 2005

  56. [64]

    Reach for a*: Efficient point-to-point shortest path algorithms,

    A. V . Goldberg, H. Kaplan, and R. F. Werneck, “Reach for a*: Efficient point-to-point shortest path algorithms,” in Workshop on Algorithm Engineering and Experimentation , 2006

  57. [65]

    Efficient constrained path planning via search in state lattices,

    M. Pivtoraiko and A. Kelly, “Efficient constrained path planning via search in state lattices,” 2005

  58. [66]

    Time-bounded lattice for efficient planning in dynamic environments,

    A. Kushleyev and M. Likhachev, “Time-bounded lattice for efficient planning in dynamic environments,” 2009 IEEE International Confer- ence on Robotics and Automation , pp. 1662–1668, 2009

  59. [67]

    On the application of the d* search algorithm to time-based planning on lattice graphs,

    M. Rufli and R. Y . Siegwart, “On the application of the d* search algorithm to time-based planning on lattice graphs,” in European Conference on Mobile Robots , 2009

  60. [68]

    Center-based 3d object detection and tracking,

    T. Yin, X. Zhou, and P. Kr ¨ahenb¨uhl, “Center-based 3d object detection and tracking,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11779–11788, 2020

  61. [69]

    Multi-camera fusion in apollo software distribution,

    G. Kulathunga, A. Buyval, and A. S. Klimchik, “Multi-camera fusion in apollo software distribution,” IFAC-PapersOnLine, 2019

  62. [70]

    Spatial as deep: Spatial cnn for traffic scene understanding,

    X. Pan, J. Shi, P. Luo, X. Wang, and X. Tang, “Spatial as deep: Spatial cnn for traffic scene understanding,” in AAAI Conference on Artificial Intelligence, 2017

  63. [71]

    Mobilenetv2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 4510–4520, 2018

  64. [72]

    Data driven prediction archi- tecture for autonomous driving and its application on apollo platform,

    K. Xu, X. Xiao, J. Miao, and Q. Luo, “Data driven prediction archi- tecture for autonomous driving and its application on apollo platform,” 2020 IEEE Intelligent Vehicles Symposium (IV) , pp. 175–181, 2020

  65. [73]

    End-to-end autonomous driving: Challenges and frontiers,

    L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li, “End-to-end autonomous driving: Challenges and frontiers,” ArXiv, vol. abs/2306.16927, 2023

  66. [74]

    Learning to drive in a day,

    A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J. M. Allen, V .-D. Lam, A. Bewley, and A. Shah, “Learning to drive in a day,” 2019 International Conference on Robotics and Automation (ICRA) , pp. 8248–8254, 2018

  67. [75]

    Cirl: Controllable imitative reinforcement learning for vision-based self-driving,

    X. Liang, T. Wang, L. Yang, and E. P. Xing, “Cirl: Controllable imitative reinforcement learning for vision-based self-driving,” ArXiv, vol. abs/1807.03776, 2018

  68. [76]

    Trans- fuser: Imitation with transformer-based sensor fusion for autonomous driving,

    K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger, “Trans- fuser: Imitation with transformer-based sensor fusion for autonomous driving,” IEEE Transactions on Pattern Analysis and Machine Intelli- gence, vol. 45, pp. 12878–12895, 2022

  69. [77]

    Safety-enhanced autonomous driving using interpretable sensor fusion transformer,

    H. Shao, L. Wang, R. Chen, H. Li, and Y . T. Liu, “Safety-enhanced autonomous driving using interpretable sensor fusion transformer,” in Conference on Robot Learning , 2022

  70. [78]

    End to end learning for self-driving cars,

    M. Bojarski, D. W. del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Zieba, “End to end learning for self-driving cars,” ArXiv, vol. abs/1604.07316, 2016

  71. [79]

    Urban driving with conditional imitation learning,

    J. Hawke, R. Shen, C. Gurau, S. Sharma, D. Reda, N. Nikolov, P. Mazur, S. Micklethwaite, N. Griffiths, A. Shah, and A. Kendall, “Urban driving with conditional imitation learning,” 2020 IEEE Inter- national Conference on Robotics and Automation (ICRA), pp. 251–257, 2019

  72. [80]

    Exploring the limitations of behavior cloning for autonomous driving,

    F. Codevilla, E. Santana, A. M. L ´opez, and A. Gaidon, “Exploring the limitations of behavior cloning for autonomous driving,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 9328–9337, 2019

  73. [81]

    Alvinn, an autonomous land vehicle in a neural network,

    D. A. Pomerleau, “Alvinn, an autonomous land vehicle in a neural network,” 2015

  74. [82]

    Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,

    D. Sun, X. Yang, M.-Y . Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 8934– 8943, 2017

  75. [83]

    Self- attention generative adversarial networks,

    H. Zhang, I. J. Goodfellow, D. N. Metaxas, and A. Odena, “Self- attention generative adversarial networks,” ArXiv, vol. abs/1805.08318, 2018

  76. [84]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2015

  77. [85]

    A model-predictive motion planner for the iara autonomous car,

    V . B. Cardoso, J. Oliveira, T. Teixeira, C. S. Badue, F. W. Mutz, T. Oliveira-Santos, L. de Paula Veronese, and A. F. de Souza, “A model-predictive motion planner for the iara autonomous car,” 2017 IEEE International Conference on Robotics and Automation (ICRA) , pp. 225–230, 2016

  78. [86]

    Mastering the game of go with deep neural networks and tree search,

    D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V . Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis,...

  79. [87]

    Continuous control with deep reinforcement learning,

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. M. O. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” CoRR, vol. abs/1509.02971, 2015

  80. [88]

    Driving with llms: Fusing object- level vector modality for explainable autonomous driving,

    L. Chen, O. Sinavski, J. H ¨unermann, A. Karnsund, A. J. Willmott, D. Birch, D. Maund, and J. Shotton, “Driving with llms: Fusing object- level vector modality for explainable autonomous driving,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) , pp. 14...

  81. [89]

    Tesla ai day 2021

    “Tesla ai day 2021.” https://www.youtube.com/watch?v= j0z4FweCy4M. Online

  82. [90]

    Designing network design spaces,

    I. Radosavovic, R. P. Kosaraju, R. B. Girshick, K. He, and P. Doll ´ar, “Designing network design spaces,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 10425–10433, 2020

  83. [91]

    Efficientdet: Scalable and efficient object detection,

    M. Tan, R. Pang, and Q. V . Le, “Efficientdet: Scalable and efficient object detection,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10778–10787, 2019

  84. [92]

    Accelerating the pony.ai av sensor data pro- cessing pipeline

    “Accelerating the pony.ai av sensor data pro- cessing pipeline.” https://developer.nvidia.com/blog/ accelerating-the-pony-av-sensor-data-processing-pipeline/. Accessed: 2024-05-31

  85. [93]

    Compute solution for tesla’s full self-driving computer,

    E. Talpes, A. Gorti, G. S. Sachdev, D. D. Sarma, G. Venkataramanan, P. J. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, and S. Arora, “Compute solution for tesla’s full self-driving computer,” IEEE Micro, vol. 40, pp. 25–35, 2020

  86. [94]

    Technology - pony.ai

    “Technology - pony.ai.” https://pony.ai/tech?lang=en. Accessed: 2024- 05-31

  87. [95]

    The architectural implications of autonomous driv- ing: Constraints and acceleration,

    S.-C. Lin, Y . Zhang, C.-H. Hsu, M. Skach, M. E. Haque, L. Tang, and J. Mars, “The architectural implications of autonomous driv- ing: Constraints and acceleration,” Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and...

  88. [96]

    NUVO-6108GC Industrial-grade GPU Computing Platform

    “NUVO-6108GC Industrial-grade GPU Computing Platform.” Online

  89. [97]

    Jetson Orin Family — NVIDIA

    NVIDIA Corporation, “Jetson Orin Family — NVIDIA.” https: //www.nvidia.com/en-us/autonomous-machines/embedded-systems/ jetson-orin/, 2024. Accessed: 2024-05-21

  90. [98]

    Nvidia drive in-vehicle computing for autonomous vehicles

    “Nvidia drive in-vehicle computing for autonomous vehicles.” https: //www.nvidia.com/en-us/self-driving-cars/in-vehicle-computing/. On- line

  91. [99]

    Eyeq chip technology

    “Eyeq chip technology.” https://www.mobileye.com/technology/ eyeq-chip/. Online

  92. [100]

    Hardware acceleration of deep neural networks for autonomous driving on fpga- based soc,

    G. Sciangula, F. Restuccia, A. Biondi, and G. C. Buttazzo, “Hardware acceleration of deep neural networks for autonomous driving on fpga- based soc,” 2022 25th Euromicro Conference on Digital System Design (DSD), pp. 406–414, 2022

  93. [101]

    Tesla vision update: Replacing ultrasonic sensors with tesla vision

    “Tesla vision update: Replacing ultrasonic sensors with tesla vision.” https://www.tesla.com/en eu/support/transitioning-tesla-vision. On- line

  94. [102]

    Model-based imitation learning for urban driving,

    A. Hu, G. Corrado, N. Griffiths, Z. Murez, C. Gurau, H. Yeo, A. Kendall, R. Cipolla, and J. Shotton, “Model-based imitation learning for urban driving,” ArXiv, vol. abs/2210.07729, 2022

  95. [103]

    Data centers on wheels: Emis- sions from computing onboard autonomous vehicles,

    S. Sudhakar, V . Sze, and S. Karaman, “Data centers on wheels: Emis- sions from computing onboard autonomous vehicles,” IEEE Micro , vol. 43, no. 1, pp. 29–39, 2023

  96. [104]

    Lazypim: An efficient cache coherence mechanism for processing-in-memory,

    A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, K. Hsieh, K. T. Malladi, H. Zheng, and O. Mutlu, “Lazypim: An efficient cache coherence mechanism for processing-in-memory,” IEEE Computer Architecture Letters, vol. 16, pp. 46–50, 2017

  97. [105]

    Cellular logic-in-memory arrays,

    W. H. Kautz, “Cellular logic-in-memory arrays,” IEEE Transactions on Computers, vol. C-18, pp. 719–727, 1969

  98. [106]

    A logic-in-memory computer,

    H. S. Stone, “A logic-in-memory computer,” IEEE Transactions on Computers, vol. C-19, pp. 73–78, 1970

  99. [107]

    New potential breakthrough memory: Hbm2 aquabolt

    “New potential breakthrough memory: Hbm2 aquabolt.” https:// semiconductor.samsung.com/dram/hbm/hbm2-aquabolt/. Online

  100. [108]

    Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory,

    D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory,” in 2016 ACM/IEEE 43rd Annual Interna- tional Symposium on Computer Architecture (ISCA) , pp. 380–392, 2016

  101. [109]

    Decomposing a scene into geomet- ric and semantically consistent regions,

    S. Gould, R. Fulton, and D. Koller, “Decomposing a scene into geomet- ric and semantically consistent regions,” 2009 IEEE 12th International Conference on Computer Vision , pp. 1–8, 2009

  102. [110]

    Tetris: Scalable and efficient neural network acceleration with 3d memory,

    M. Gao, J. Pu, X. S. Yang, M. Horowitz, and C. E. Kozyrakis, “Tetris: Scalable and efficient neural network acceleration with 3d memory,” Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, 2017

  103. [111]

    Eyeriss: An energy- efficient reconfigurable accelerator for deep convolutional neural net- works,

    Y . hsin Chen, T. Krishna, J. S. Emer, and V . Sze, “Eyeriss: An energy- efficient reconfigurable accelerator for deep convolutional neural net- works,” IEEE Journal of Solid-State Circuits , vol. 52, pp. 127–138, 2016

  104. [112]

    Ambit: In- memory accelerator for bulk bitwise operations using commodity dram technology,

    V . Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. S. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Ambit: In- memory accelerator for bulk bitwise operations using commodity dram technology,”2017 50th Annual IEEE/ACM International Symposium on Microa...

  105. [113]

    Simdram: An end-to-end framework for bit-serial simd computing in dram,

    N. Hajinazar, G. F. Oliveira, S. Gregorio, J. D. Ferreira, N. M. Ghiasi, M. Patel, M. Alser, S. Ghose, J. G. Luna, and O. Mutlu, “Simdram: An end-to-end framework for bit-serial simd computing in dram,” ArXiv, vol. abs/2105.12839, 2021

  106. [114]

    Dracc: a dram based accelerator for accurate cnn inference,

    Q. Deng, L. Jiang, Y . Zhang, M. Zhang, and J. Yang, “Dracc: a dram based accelerator for accurate cnn inference,” 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC) , pp. 1–6, 2018

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.