Pith. sign in

REVIEW 4 major objections 5 minor 140 references

DSRC: Learning Density-insensitive and Semantic-aware Collaborative Representation against Corruptions

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read DSRC learns density-insensitive, semantic-aware collaborative representations that beat prior methods under all six tested LiDAR corruptions.

desk verdict A genuinely useful first corruption benchmark for collaborative perception, with a method that wins on it; the load-bearing uncertainty is whether the hand-picked corruption severities transfer to real deployments. read the letter →

arxiv 2412.10739 v2 pith:YPVGBSS6 submitted 2024-12-14 cs.CV

classification cs.CV
keywords collaborativeperception3DobjectdetectioncorruptionrobustnessknowledgedistillationLiDARpointcloudreconstructionV2Xbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that multi-agent collaborative perception, which is usually evaluated only on clean data, degrades sharply under realistic LiDAR corruptions, and that such failures can be reduced by training a student model to mimic a privileged teacher. It introduces a benchmark of six corruption types (beam missing, motion blur, fog, snow, crosstalk, and cross sensor) built on the OPV2V and DAIR-V2X datasets, and proposes DSRC, a teacher-student distillation framework. The teacher sees a multi-view dense point cloud whose object regions are painted with ground-truth semantic labels; the student sees only its own sparse single-view cloud and is aligned to the teacher after encoding, after fusion, and after prediction, with an additional feature-to-point cloud reconstruction loss. The paper reports that DSRC outperforms state-of-the-art collaborative perception methods on both clean and corrupted conditions, including a 4.9% AP@0.5 gain on cross-sensor corruption and a 4.02% AP@0.7 gain on crosstalk over the next best method on OPV2V. If the benchmark faithfully represents real-world degradation, the method would harden deployed V2X perception without adding inference cost.

What carries the argument

The load-bearing mechanism is a semantic-guided sparse-to-dense distillation loop. The teacher point cloud $\mathbb{P}^T$ is generated by replacing the object regions of the ego agent's sparse single-view cloud $\mathbb{P}^S$ with dense object points aggregated from multiple collaborative agents, then painting every point with a semantic indicator $s \in \{0,1\}$ derived from ground-truth boxes; features extracted from this denser, semantically labeled cloud are the supervision target. The student is aligned to the teacher at three depths: distillation after encoding ($L_d$, an $\ell^2$ distance on foreground-masked bird's-eye-view features), after fusion ($L_h$, an $\ell^2$ distance on fused features), and after prediction ($L_p$, a KL divergence on classification and regression outputs). A voxel-level feature-to-point cloud reconstruction module, predicting occupancy masks and point offsets, adds supervision on the fused features. All losses are computed on clean data only, and the teacher and reconstruction module are discarded at inference.

What would settle it

Collect real corrupted LiDAR scenes (physical fog or snow, actual beam dropout, crosstalk interference, or heterogeneous sensor rigs) with the same annotation format as OPV2V or DAIR-V2X, run DSRC and the strongest baseline on them, and compare AP@0.5 and AP@0.7; if the gains shrink below the reported 4.9% and 4.02% margins, or reverse, the robustness claim is falsified. A cheaper check is to re-run the existing benchmark across a severity sweep (for example, 8 versus 32 dropped beams, or 0.1 m versus 0.4 m jitter) and see whether DSRC remains best at all severities.

Watch

Extended reading notes

Core claim

The central claim is that corruption robustness in collaborative 3D detection can be learned from clean data alone by distilling from a teacher that has access to information absent at inference time. The teacher is built by fusing all agents' point clouds into a multi-view scene, replacing object regions in the sparse ego cloud with denser points contributed by multiple views, and appending a ground-truth semantic indicator to every point. DSRC then trains the student, which receives only its own sparse cloud, to match the teacher's bird's-eye-view features, its fused features, and its prediction logits, while a point cloud reconstruction head regularizes the fused representation by predicting voxel occupancy and point offsets. The paper reports that on OPV2V and DAIR-V2X the resulting student beats prior intermediate-fusion methods under clean conditions and under all six corruptions, with the largest margins on weather-related corruption, and that the teacher and reconstruction head are discarded at inference so the deployed model costs no more than the baseline detector.

Load-bearing premise

The load-bearing premise is that the six hand-simulated corruptions (16 dropped beams, 0.2 m jitter, 0.01-ratio crosstalk with 3 m noise, and fog and snow from published simulators) faithfully represent the corruptions a deployed V2X system will meet, so a method that wins on the simulated benchmark will also win on real corruption data.

Editorial extensions

If this is right

  • If the central claim holds, a collaborative detector can be made robust to fog, snow, beam dropout, jitter, crosstalk, and sensor heterogeneity without training on any corrupted point cloud.
  • The framework is fusion-agnostic: the paper's ablations show it also improves a position-wise maximum fusion baseline and even a no-collaboration single-vehicle mode, so the distillation can be layered onto existing fusion designs.
  • The benchmark supplies a six-corruption protocol with mCE and mAP reporting on OPV2V and DAIR-V2X, making future robustness claims in collaborative perception comparable.
  • Because only the student is retained, the robustness gain carries no additional communication or inference overhead in the deployed system.
  • Weather corruptions, especially fog and snow, emerge as the hardest cases, suggesting that future robustness work should concentrate on weather-induced sparsity and semantic degradation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the simulated corruptions match real sensor degradation, the same privileged-teacher recipe could transfer to other sparse-sensor tasks such as radar detection or low-beam-count LiDAR, since the teacher's advantage comes from cross-view density and semantic labels rather than from a specific corruption model.
  • The robustness ranking could be severity-dependent: the hand-chosen simulation parameters (16 dropped beams, 0.2 m jitter, 0.01 crosstalk ratio, and 3 m noise) set a difficulty level, and a benchmark with harsher or milder corruptions might reorder the baselines; reporting performance across severity sweeps would strengthen the comparison.
  • A clean-trained model that distills from ground-truth semantics is effectively using privileged information at train time, so combining the same scheme with corruption augmentation at training might push robustness further, a direction the paper does not explore.
  • The 'first comprehensive benchmark' claim should be read relative to this paper's definition of comprehensiveness; real corrupted LiDAR validation would be needed to confirm that the six simulated corruptions cover the failure modes that matter.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces a robustness benchmark for multi-agent collaborative perception under six LiDAR corruptions (beam missing, motion blur, fog, snow, crosstalk, and cross sensor) and proposes a method named DSRC based on teacher-student distillation. The teacher is trained on a dense multi-view point cloud painted with ground-truth semantic labels, and knowledge is transferred to a student through three-stage distillation (after encoding, after fusion, and after prediction), together with a feature-to-point-cloud reconstruction loss that regularizes collaborative feature fusion. Experiments on OPV2V and DAIR-V2X compare DSRC against no-collaboration, late fusion, and five intermediate-fusion baselines under clean and corrupted settings. Tables 1 and 2 report consistent gains, and ablations attribute the improvement to the three proposed components. The abstract claims that DSRC outperforms state-of-the-art collaborative perception methods in both clean and corrupted conditions.

Significance. If the result holds, the paper provides a useful corruption robustness benchmark for collaborative perception and a generally applicable distillation recipe. The strengths are the released code, the coverage of six physically motivated corruption types, the consistent gains across two datasets, and the clean ablation evidence for each component. The use of a privileged teacher that is discarded at inference is a standard and acceptable design. The significance is currently limited by three evidential gaps: all quantitative results are single-run without error bars or significance testing; the corruption benchmark uses hand-chosen severity parameters that are never swept; and no real corrupted multi-agent LiDAR data is used for validation. The headline claim of state-of-the-art robustness is therefore established only at one synthetic operating point.

major comments (4)
  1. [Table 1, OPV2V block] In Table 1 on OPV2V, the F-Cooper row reports identical AP values for Fog and Crosstalk (63.76/53.52) and identical values for Beam Missing and Cross Sensor (75.65/65.91). These duplicates are unlikely to be genuine measurement outcomes. Since F-Cooper is one of the baselines used for the claimed state-of-the-art gains and contributes to Figure 5, these entries must be corrected and the reported gains re-checked against the corrected numbers.
  2. [Appendix, 'More Details of Common Corruptions'] The corruption simulation parameters are fixed at a single severity level: 16 randomly dropped beams, 0.2 m jitter, 0.01 ratio with 3 m Gaussian noise for crosstalk, and the Hahner et al. fog and snow settings. The paper's central robustness claim is measured only at this operating point, and no severity sweep is reported. Because the relative ordering of methods can change with severity, please add sweeps (e.g., 8/16/32 beams, 0.1/0.2/0.4 m jitter, 0.005/0.01/0.02 crosstalk ratio, and several fog/snow densities) or justify the chosen severities with published sensor or weather models. Without this, the claim that DSRC outperforms SOTA under corruptions is not established as a general property.
  3. [Tables 1-4 and Figure 5] All reported performance numbers are from a single training run, with no error bars, standard deviations, or significance tests. Some differences are small (e.g., Clean AP@0.5 on OPV2V is 92.58 for DSRC versus 92.29 for V2VNet), and the ablation increments in Table 2 are also single-run. Given the SOTA claim, please report results from multiple seeds (at least three) as mean plus/minus standard deviation, or otherwise demonstrate that the observed ordering is stable across training runs.
  4. [Datasets and Evaluation Metrics; Introduction] The evaluation is performed on OPV2V, a clean CARLA simulation, and DAIR-V2X, a clean real-world dataset; the corrupted inputs are produced entirely by the authors' simulations. No real corrupted multi-agent LiDAR data is used. The abstract's reference to 'complex real-world environments' is therefore not directly evidenced. Please soften the claim or add validation on real degraded data (for example, adverse-weather sequences from V2V4Real if available), and discuss the expected transferability of simulated corruptions.
minor comments (5)
  1. [Equation (10)] The definitions of CE and mCE do not state whether they are computed at AP@0.5 or AP@0.7, and Figure 5 does not specify the threshold either; please clarify this in the text and in the figure caption.
  2. [Quantitative Evaluation, 'Comparison of corruption types'] The paragraph discussing Figure 4 contains a long run of unicode escape sequences (e.g., '/uni00000026/uni0000004f/...'), which is likely a PDF or LaTeX encoding artifact; this text must be repaired before submission.
  3. [Appendix, 'Metadata Sharing and Feature Extraction'] The sentence 'the features of the i-th agent is the features of the i-th agent is obtained as ...' contains a duplicated phrase and a subject-verb agreement error; please rewrite it.
  4. [References] The reference list contains many entries unrelated to the content of the paper (e.g., Clancey 1984, NASA 2015, and several duplicated entries), suggesting the bibliography was assembled from an external template; please trim it to only the works actually cited.
  5. [Introduction] The terms 'first comprehensive benchmark' and 'first study on the robustness' are stated without positioning against existing robustness benchmarks for single-agent 3D perception such as Robo3D or any prior collaborative robustness evaluation; please clarify the novelty relative to those lines of work.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: DSRC's teacher-student losses are training-time supervision, not disguised predictions, and all methods are compared under a shared corruption protocol.

full rationale

The paper's derivation chain is empirical rather than analytic, and no load-bearing step reduces to its own inputs by construction. The sparse-to-dense distillation uses a teacher point cloud built from multi-view aggregated points painted with ground-truth bounding boxes; however, the resulting distillation losses Ld, Lh, Lp and the reconstruction loss Lrec are training objectives only, and the paper explicitly states that during inference only the student model is retained and the teacher and reconstruction modules are discarded. Thus the GT-painted teacher is privileged training supervision, not a renamed prediction of the final model. The corruption benchmark is defined by fixed simulation rules (16 dropped beams, 0.2 m jitter, 0.01-ratio/3 m crosstalk, and Hahner et al. fog/snow), and every baseline is evaluated under the identical protocol, so the relative ranking is not forced by construction; no parameter is fitted to the corrupted test set, and the mCE metric of Eq. 10 is a standard normalized performance-drop measure rather than an objective optimized by DSRC. Self-citations (e.g., ERMVP as a baseline) are used for comparison only and are not load-bearing justification of the central claim. The more serious concern, that hand-chosen corruption severities may not reflect real deployment distributions, is an external-validity or soundness issue rather than circularity.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The empirical claims rest on the validity of the corruption simulations, the optimality of the GT-painted dense teacher signal, and the availability of precise inter-agent poses. Hand-chosen simulation and loss hyperparameters are the main free parameters. No new physical entities, forces, or particles are introduced; the semantic indicator is an input channel augmentation rather than an invented entity.

free parameters (6)
  • Distillation loss balance hyperparameters α, β, γ = α=1, β=1, γ=0.5
    Chosen in Implementation Details with no sensitivity analysis; the total distillation loss is L = αLd + βLh + γLp, so the reported gains could depend on these values.
  • Beam missing severity = 16 beams removed
    Hand-chosen in the Appendix to simulate the corruption; directly sets the difficulty of the beam missing robustness condition.
  • Motion blur jitter standard deviation = 0.2
    Hand-chosen in the Appendix; added to each point coordinate to simulate motion blur.
  • Crosstalk noise parameters = 0.01 subset, 3 m Gaussian noise
    Hand-chosen in the Appendix; the subset ratio and noise magnitude control how severe the crosstalk corruption is.
  • Fog and snow simulation severity = Not reported (inherited from Hahner et al. 2021 and 2022)
    The simulations are adopted from prior work, but the exact severity levels and parameters used in this benchmark are not specified, limiting reproducibility.
  • Inference thresholds = Confidence 0.25, NMS IoU 0.15
    Used for all methods, but these thresholds affect AP values and are not varied or justified in the paper.
assumptions (4)
  • domain assumption The teacher point cloud, built by replacing sparse object regions with multi-view dense points and painting them with ground truth boxes, defines the optimal observable feature F^o in Eq. (1).
    Methods section 'Teacher Point Cloud Generation'; the objective in Eq. (1) assumes that this GT-painted dense representation is the right distillation target, but no proof or empirical validation of optimality is given.
  • domain assumption The six simulated corruptions faithfully represent real-world natural corruptions for LiDAR in V2X deployment.
    Experiment Setup and Appendix 'More Details of Common Corruptions'; the robustness claims rest on this representativeness, which is not validated against real corrupted LiDAR data.
  • domain assumption Accurate inter-agent pose transformations Γj→i are available for aggregating multi-view point clouds and for fusion.
    Methods section 'Multi-view Construction of Object Point Cloud'; noisy or erroneous poses are not modeled, even though prior work such as CoAlign specifically addresses pose errors.
  • standard math Standard deep learning components (PointPillars encoder, anchor-based detection, focal loss, smooth L1 loss, Adam optimizer, cosine annealing) behave as commonly expected.
    Implementation Details; these are standard background methods used without proof or re-derivation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DSRC: Learning Density-insensitive and Semantic-aware Collaborative Representation against Corruptions." pith.science (2026). https://pith.science/paper/YPVGBSS6

@misc{pith2026241210739,
  author       = {Pith},
  title        = {Pith review of: DSRC: Learning Density-insensitive and Semantic-aware Collaborative Representation against Corruptions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YPVGBSS6}},
  note         = {Machine review of arXiv:2412.10739}
}
read the original abstract

As a potential application of Vehicle-to-Everything (V2X) communication, multi-agent collaborative perception has achieved significant success in 3D object detection. While these methods have demonstrated impressive results on standard benchmarks, the robustness of such approaches in the face of complex real-world environments requires additional verification. To bridge this gap, we introduce the first comprehensive benchmark designed to evaluate the robustness of collaborative perception methods in the presence of natural corruptions typical of real-world environments. Furthermore, we propose DSRC, a robustness-enhanced collaborative perception method aiming to learn Density-insensitive and Semantic-aware collaborative Representation against Corruptions. DSRC consists of two key designs: i) a semantic-guided sparse-to-dense distillation framework, which constructs multi-view dense objects painted by ground truth bounding boxes to effectively learn density-insensitive and semantic-aware collaborative representation; ii) a feature-to-point cloud reconstruction approach to better fuse critical collaborative representation across agents. To thoroughly evaluate DSRC, we conduct extensive experiments on real-world and simulated datasets. The results demonstrate that our method outperforms SOTA collaborative perception methods in both clean and corrupted conditions. Code is available at https://github.com/Terry9a/DSRC.

Figures

Figures reproduced from arXiv: 2412.10739 by the authors.

Figure 1
Figure 1. Visualization of typical corruption types in our [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. These scenarios encompass three distinct corrup [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The overall architecture of the proposed framework DSRC. It contains two branches with identical network structures: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figures from the paper (5 more)
Figure 3
Figure 3. Figure 3: Illustration of the proposed point cloud reconstruc [PITH_FULL_IMAGE:figures/full_fig_p005_3.png]
Figure 4
Figure 4. Figure 4: The average performance of all models under dif [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Benchmarking results of all models on the six ro [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: Ablation study on number of agents. Number of agents. In this part, we evaluate the impact of the number of collaborative agents on the performance of DSRC. As depicted in [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Visualization of collaborative perception results on a roadway in 3D view and BEV view, with [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

140 extracted references · 39 canonical work pages

  1. [1]

    N.; Seid, A

    Abishu, H. N.; Seid, A. M.; Yacob, Y. H.; Ayall, T.; Sun, G.; and Liu, G. 2021. Consensus mechanism for blockchain-enabled vehicle-to-vehicle energy trading in the internet of electric vehicles. IEEE Transactions on Vehicular Technology, 71(1): 946--960

  2. [2]

    L.; Kiros, J

    Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016. Layer normalization. arXiv preprint arXiv:1607.06450

  3. [3]

    J.; Liu, Y.; Sisbot, E

    Bai, Z.; Wu, G.; Barth, M. J.; Liu, Y.; Sisbot, E. A.; Oguchi, K.; and Huang, Z. 2022. A survey and framework of cooperative perception: From heterogeneous singleton to hierarchical cooperation. arXiv preprint arXiv:2208.10590

  4. [4]

    Bang, G.; Choi, K.; Kim, J.; Kum, D.; and Choi, J. W. 2024. RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features. arXiv preprint arXiv:2403.05061

  5. [5]

    Bengio, S.; Vinyals, O.; Jaitly, N.; and Shazeer, N. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems, 28

  6. [6]

    Bengio, Y.; Louradour, J.; Collobert, R.; and Weston, J. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, 41--48

  7. [7]

    Bintoro, K. B. Y. 2021. A Study of V2V Communication on VANET: Characteristic, Challenges and Research Trends. JISA (Jurnal Informatika dan Sains), 4(1): 46--58

  8. [8]

    P.; Anuradha, T.; Rao, P

    Bojjagani, S.; Reddy, Y. P.; Anuradha, T.; Rao, P. V.; Reddy, B. R.; and Khan, M. K. 2022. Secure Authentication and Key Management Protocol for Deployment of Internet of Vehicles (IoV) Concerning Intelligent Transport Systems. IEEE Transactions on Intelligent Transportation Systems, 23(12): 24698--24713

Show all 140 references
  1. [9]

    G.; Petrillo, A.; and Santini, S

    Caiazzo, B.; Lui, D. G.; Petrillo, A.; and Santini, S. 2022. Cooperative Finite-time Control for autonomous vehicles platoons with nonuniform V2V communication delays. IFAC-PapersOnLine, 55(36): 145--150

  2. [10]

    Chafii, M.; Naoumi, S.; Alami, R.; Almazrouei, E.; Bennis, M.; and Debbah, M. 2023. Emergent Communication in Multi-Agent Reinforcement Learning for Future Wireless Networks. arXiv preprint arXiv:2309.06021

  3. [11]

    Chen, M.-Y.; Chiang, H.-S.; and Yang, K.-J. 2022. Constructing cooperative intelligent transport systems for travel time prediction with deep learning approaches. IEEE Transactions on Intelligent Transportation Systems, 23(9): 16590--16599

  4. [12]

    Chen, Q.; Ma, X.; Tang, S.; Guo, J.; Yang, Q.; and Fu, S. 2019 a . F-cooper: Feature based cooperative perception for autonomous vehicle edge computing system using 3D point clouds. In Proceedings of the 4th ACM/IEEE Symposium on Edge Computing, 88--100

  5. [13]

    Chen, Q.; Tang, S.; Yang, Q.; and Fu, S. 2019 b . Cooper: Cooperative perception for connected autonomous vehicles based on 3d point clouds. In 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS), 514--524. IEEE

  6. [15]

    Chen, X.; Zhang, T.; Wang, Y.; Wang, Y.; and Zhao, H. 2022 b . Futr3d: A unified sensor fusion framework for 3d detection. [Online]. Available: https://arxiv.org/abs/2203.10642

  7. [16]

    Chen, Z.; Li, B.; Wu, S.; Ding, S.; and Zhang, W. 2023. Query-efficient decision-based black-box patch attack. IEEE Transactions on Information Forensics and Security

  8. [17]

    Chen, Z.; Li, B.; Wu, S.; Jiang, K.; Ding, S.; and Zhang, W. 2024. Content-based unrestricted adversarial attack. Advances in Neural Information Processing Systems, 36

  9. [18]

    Chen, Z.; Li, B.; Wu, S.; Xu, J.; Ding, S.; and Zhang, W. 2022 c . Shape matters: deformable patch attack. In European conference on computer vision, 529--548. Springer

  10. [19]

    Chen, Z.; Li, Z.; Zhang, S.; Fang, L.; Jiang, Q.; and Zhao, F. 2022 d . Bevdistill: Cross-modal bev distillation for multi-view 3d object detection. arXiv preprint arXiv:2211.09386

  11. [20]

    Clancey, W. J. 1979. Transfer of Rule-Based Expertise through a Tutorial Dialogue . Ph.D. diss., Dept.\ of Computer Science, Stanford Univ., Stanford, Calif

  12. [21]

    Clancey, W. J. 1983. Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education . In Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83) , 556--560. Menlo Park, Calif: IJCAI ...

  13. [22]

    Clancey, W. J. 1984. Classification Problem Solving . In Proceedings of the Fourth National Conference on Artificial Intelligence, 45--54. Menlo Park, Calif.: AAAI Press

  14. [23]

    Clancey, W. J. 2021. The Engineering of Qualitative Models . Forthcoming

  15. [24]

    S.; Sabaliauskaite, G.; and Zhou, F

    Cui, J.; Liew, L. S.; Sabaliauskaite, G.; and Zhou, F. 2019. A review on safety failures, security attacks, and available countermeasures for autonomous vehicles. Ad Hoc Networks, 90: 101823

  16. [25]

    Dai, Y.; Gieseke, F.; Oehmcke, S.; Wu, Y.; and Barnard, K. 2021. Attentional feature fusion. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 3560--3569

  17. [26]

    Dong, H.; Zhang, X.; Jiang, X.; Zhang, J.; Xu, J.; Ai, R.; Gu, W.; Lu, H.; Kannala, J.; and Chen, X. 2022. SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map Generation and Prediction. [Online]. Available: https://arxiv.org/abs/2211.15656

  18. [27]

    Dong, Y.; Kang, C.; Zhang, J.; Zhu, Z.; Wang, Y.; Yang, X.; Su, H.; Wei, X.; and Zhu, J. 2023. Benchmarking robustness of 3d object detection to common corruptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1022--1032

  19. [28]

    Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929

  20. [29]

    Dosovitskiy, A.; Ros, G.; Codevilla, F.; Lopez, A.; and Koltun, V. 2017. CARLA: An open urban driving simulator. In Conference on robot learning, 1--16. PMLR

  21. [30]

    El Zorkany, M.; Yasser, A.; and Galal, A. I. 2020. Vehicle to vehicle “V2V” communication: scope, importance, challenges, research directions and future. The Open Transportation Journal, 14(1)

  22. [31]

    Engelmore, R.; and Morgan, A., eds. 1986. Blackboard Systems. Reading, Mass.: Addison-Wesley

  23. [32]

    Fang, J.; Qiao, J.; Xue, J.; and Li, Z. 2023. Vision-Based Traffic Accident Detection and Anticipation: A Survey. IEEE Transactions on Circuits and Systems for Video Technology

  24. [33]

    Gao, S.-H.; Cheng, M.-M.; Zhao, K.; Zhang, X.-Y.; Yang, M.-H.; and Torr, P. 2019. Res2net: A new multi-scale backbone architecture. IEEE transactions on pattern analysis and machine intelligence, 43(2): 652--662

  25. [34]

    Goodfellow, I.; Bengio, Y.; and Courville, A. 2016. Deep learning. MIT press

  26. [35]

    Gu, J.; Zhang, J.; Zhang, M.; Meng, W.; Xu, S.; Zhang, J.; and Zhang, X. 2023. FeaCo: Reaching Robust Feature-Level Consensus in Noisy Pose Conditions. In Proceedings of the 31st ACM International Conference on Multimedia, MM '23, 3628–3636. New York, NY, USA: Association for ...

  27. [36]

    Hahner, M.; Sakaridis, C.; Bijelic, M.; Heide, F.; Yu, F.; Dai, D.; and Van Gool, L. 2022. Lidar snowfall simulation for robust 3d object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16364--16374

  28. [37]

    Hahner, M.; Sakaridis, C.; Dai, D.; and Van Gool, L. 2021. Fog simulation on real LiDAR point clouds for 3D object detection in adverse weather. In Proceedings of the IEEE/CVF international conference on computer vision, 15283--15292

  29. [38]

    Hasan, M.; Mohan, S.; Shimizu, T.; and Lu, H. 2020. Securing vehicle-to-everything (V2X) communication platforms. IEEE Transactions on Intelligent Vehicles, 5(4): 693--713

  30. [39]

    W.; Clancey, W

    Hasling, D. W.; Clancey, W. J.; and Rennels, G. 1984. Strategic explanations for a diagnostic consultation system. International Journal of Man-Machine Studies, 20(1): 3--19

  31. [40]

    W.; Clancey, W

    Hasling, D. W.; Clancey, W. J.; Rennels, G. R.; and Test, T. 1983. Strategic Explanations in Consultation---Duplicate . The International Journal of Man-Machine Studies, 20(1): 3--19

  32. [41]

    Hatamizadeh, A.; Yin, H.; Heinrich, G.; Kautz, J.; and Molchanov, P. 2023. Global context vision transformers. In International Conference on Machine Learning, 12633--12646. PMLR

  33. [42]

    He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770--778

  34. [43]

    He, S.; Yan, Q.; Wu, F.; Wang, L.; L \'e cuyer, M.; and Beschastnikh, I. 2023. GlueFL: Reconciling Client Sampling and Model Masking for Bandwidth Efficient Federated Learning. Proceedings of Machine Learning and Systems, 5

  35. [44]

    Hinton, G.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. arXiv:1503.02531

  36. [45]

    Hochreiter, S.; and Schmidhuber, J. 1997. Long short-term memory. Neural computation, 9(8): 1735--1780

  37. [46]

    Hong, S.; Liu, Y.; Li, Z.; Li, S.; and He, Y. 2024. Multi-agent Collaborative Perception via Motion-aware Robust Communication Network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15301--15310

  38. [47]

    Hu, G.; Zhu, Y.; Zhao, D.; Zhao, M.; and Hao, J. 2020. Event-Triggered Multi-agent Reinforcement Learning with Communication under Limited-bandwidth Constraint. [Online]. Available: https://arxiv.org/abs/2010.04978

  39. [48]

    Hu, Y.; Peng, J.; Liu, S.; Ge, J.; Liu, S.; and Chen, S. 2024. Communication-Efficient Collaborative Perception via Information Filling with Codebook. arXiv:2405.04966

  40. [49]

    F.; Zhong, Y.; and Chen, S

    Hu, Y.; snd Zixing Lei, S. F.; Zhong, Y.; and Chen, S. 2022. Where2comm: Communication-Efficient Collaborative Perception via Spatial Confidence Maps. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS)

  41. [50]

    Huang, J.; Huang, G.; Zhu, Z.; and Du, D. 2022. Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. [Online]. Available: https://arxiv.org/abs/2112.11790

  42. [51]

    Huang, T.; Zhang, Y.; Zheng, M.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2024. Knowledge diffusion for distillation. Advances in Neural Information Processing Systems, 36

  43. [52]

    Jiang, K.; Chen, Z.; Huang, H.; Wang, J.; Yang, D.; Li, B.; Wang, Y.; and Zhang, W. 2023. Efficient Decision-based Black-box Patch Attacks on Video Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 4379--4389

  44. [53]

    Ju, B.; Zou, Z.; Ye, X.; Jiang, M.; Tan, X.; Ding, E.; and Wang, J. 2022. Paint and distill: Boosting 3d object detection with semantic passing network. In Proceedings of the 30th ACM International Conference on Multimedia, 5639--5648

  45. [54]

    Kenney, J. B. 2011. Dedicated Short-Range Communications (DSRC) Standards in the United States. Proceedings of the IEEE, 99(7): 1162--1182

  46. [55]

    Khan, H.; Samarakoon, S.; and Bennis, M. 2020. Enhancing video streaming in vehicular networks via resource slicing. IEEE Transactions on Vehicular Technology, 69(4): 3513--3522

  47. [56]

    J.; Lee, T.; Son, K.; and Yi, Y

    Kim, D.; Moon, S.; Hostallero, D.; Kang, W. J.; Lee, T.; Son, K.; and Yi, Y. 2019. Learning to schedule communication in multi-agent reinforcement learning. [Online]. Available: https://arxiv.org/abs/1902.01554

  48. [57]

    P.; and Ba, J

    Kingma, D. P.; and Ba, J. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR)

  49. [58]

    Kong, L.; Liu, Y.; Li, X.; Chen, R.; Zhang, W.; Ren, J.; Pan, L.; Chen, K.; and Liu, Z. 2023. Robo3d: Towards robust and reliable 3d perception against corruptions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 19994--20006

  50. [59]

    Koopman, P.; and Wagner, M. 2017. Autonomous vehicle safety: An interdisciplinary challenge. IEEE Intelligent Transportation Systems Magazine, 9(1): 90--96

  51. [60]

    H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O

    Lang, A. H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O. 2019. Pointpillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12697--12705

  52. [61]

    Lei, Z.; Ren, S.; Hu, Y.; Zhang, W.; and Chen, S. 2022. Latency-aware collaborative perception. In European Conference on Computer Vision, 316--332. Springer

  53. [62]

    Li, Y.; An, Z.; Wang, Z.; Zhong, Y.; Chen, S.; and Feng, C. 2022 a . V2x-sim: A virtual collaborative perception dataset for autonomous driving. arXiv preprint arXiv:2202.08449

  54. [63]

    Li, Y.; Bhanu, B.; and Lin, W. 2010 a . Auction protocol for camera active control. In 2010 IEEE International Conference on Image Processing, 4325--4328. IEEE

  55. [64]

    Li, Y.; Bhanu, B.; and Lin, W. 2010 b . Auction protocol for camera active control. In 2010 IEEE International Conference on Image Processing, 4325--4328

  56. [65]

    H.; and Netravali, R

    Li, Y.; Padmanabhan, A.; Zhao, P.; Wang, Y.; Xu, G. H.; and Netravali, R. 2020. Reducto: On-camera filtering for resource-efficient real-time video analytics. In Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, t...

  57. [66]

    Li, Y.; Ren, S.; Wu, P.; Chen, S.; Feng, C.; and Zhang, W. 2021. Learning distilled collaboration graph for multi-agent perception. Advances in Neural Information Processing Systems, 34: 29541--29552

  58. [67]

    Li, Z.; Wang, W.; Li, H.; Xie, E.; Sima, C.; Lu, T.; Yu, Q.; and Dai, J. 2022 b . BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers. [Online]. Available: https://arxiv.org/abs/2203.17270

  59. [68]

    Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollar, P. 2017. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)

  60. [69]

    Liu, Y.; Liu, J.; Yang, K.; Ju, B.; Liu, S.; Wang, Y.; Yang, D.; Sun, P.; and Song, L. 2024. AMP-Net: Appearance-Motion Prototype Network Assisted Automatic Video Anomaly Detection System. IEEE Transactions on Industrial Informatics, 20(2): 2843--2855

  61. [70]

    Liu, Y.-C.; Tian, J.; Glaser, N.; and Kira, Z. 2020. When2com: Multi-agent perception via communication graph grouping. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, 4106--4115

  62. [71]

    Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, 10012--10022

  63. [72]

    Liu, Z.; Ning, J.; Cao, Y.; Wei, Y.; Zhang, Z.; Lin, S.; and Hu, H. 2022 a . Video swin transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 3202--3211

  64. [73]

    Liu, Z.; Tang, H.; Amini, A.; Yang, X.; Mao, H.; Rus, D.; and Han, S. 2022 b . BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation. [Online]. Available: https://arxiv.org/abs/2205.13542

  65. [74]

    Lu, Y.; Li, Q.; Liu, B.; Dianati, M.; Feng, C.; Chen, S.; and Wang, Y. 2023. Robust collaborative 3d object detection in presence of pose errors. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 4812--4818. IEEE

  66. [75]

    Meng, Z.; Xia, X.; Xu, R.; Liu, W.; and Ma, J. 2023. HYDRO-3D: Hybrid Object Detection and Tracking for Cooperative Perception Using 3D LiDAR. IEEE Transactions on Intelligent Vehicles

  67. [76]

    Nair, V.; and Hinton, G. E. 2010. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10), 807--814

  68. [77]

    NASA . 2015. Pluto: The 'Other' Red Planet. https://www.nasa.gov/nh/pluto-the-other-red-planet. Accessed: 2018-12-06

  69. [78]

    Ngo, H.; Fang, H.; and Wang, H. 2023. Cooperative Perception With V2V Communication for Autonomous Vehicles. IEEE Transactions on Vehicular Technology

  70. [79]

    R.; and Gombolay, M

    Niu, Y.; Paleja, R. R.; and Gombolay, M. C. 2021. Multi-Agent Graph-Attention Communication and Teaming. In AAMAS, 964--973

  71. [80]

    Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32

  72. [81]

    Qiu, H.; Huang, P.; Asavisanu, N.; Liu, X.; Psounis, K.; and Govindan, R. 2021. Autocast: Scalable infrastructure-less cooperative perception for distributed collaborative driving. [Online]. Available: https://arxiv.org/abs/2112.14947

  73. [82]

    Qureshi, F.; and Terzopoulos, D. 2008. Smart camera networks in virtual reality. Proceedings of the IEEE, 96(10): 1640--1656

  74. [83]

    Ren, J.; Pan, L.; and Liu, Z. 2022. Benchmarking and Analyzing Point Cloud Classification under Corruptions. arXiv:2202.03377

  75. [84]

    Rice, J. 1986. Poligon: A System for Parallel Problem Solving . Technical Report KSL-86-19, Dept.\ of Computer Science, Stanford Univ

  76. [85]

    Robinson, A. L. 1980 a . New Ways to Make Microcircuits Smaller. Science, 208(4447): 1019--1022

  77. [86]

    Robinson, A. L. 1980 b . New Ways to Make Microcircuits Smaller---Duplicate Entry . Science, 208: 1019--1026

  78. [87]

    Rodriguez, A.; and Laio, A. 2014. Clustering by fast search and find of density peaks. science, 344(6191): 1492--1496

  79. [88]

    Shang, C.; Li, H.; Meng, F.; Wu, Q.; Qiu, H.; and Wang, L. 2023. Incrementer: Transformer for class-incremental semantic segmentation with knowledge distillation focusing on old class. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7214--7224

  80. [89]

    Shannon, C. E. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3): 379--423

  81. [90]

    Shi, X.; Chen, Z.; Wang, H.; Yeung, D.-Y.; Wong, W.-K.; and Woo, W.-c. 2015. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Advances in neural information processing systems, 28

  82. [91]

    Song, Q.; Wang, C.; Jiang, Z.; Wang, Y.; Tai, Y.; Wang, C.; Li, J.; Huang, F.; and Wu, Y. 2021. Rethinking counting and localization in crowds: A purely point-based framework. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 3365--3374

  83. [93]

    Song, Z.; Wen, F.; Zhang, H.; and Li, J. 2022 b . An efficient and robust object-level cooperative perception framework for connected and automated driving. arXiv preprint arXiv:2210.06289

  84. [94]

    Tan, M. 1993. Multi-agent reinforcement learning: Independent vs. cooperative agents. In Proceedings of the tenth international conference on machine learning, 330--337

  85. [95]

    Tan, M.; and Le, Q. 2021. Efficientnetv2: Smaller models and faster training. In International conference on machine learning, 10096--10106. PMLR

  86. [96]

    Vadivelu, N.; Ren, M.; Tu, J.; Wang, J.; and Urtasun, R. 2021. Learning to communicate and correct pose errors. In Conference on Robot Learning, 1195--1210. PMLR

  87. [97]

    N.; Kaiser, L.; and Polosukhin, I

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention Is All You Need. arXiv:1706.03762

  88. [98]

    Verelst, T.; and Tuytelaars, T. 2021. BlockCopy: High-resolution video processing with block-sparse feature propagation and online policies. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 5158--5167

  89. [99]

    Wang, B.; Zhang, L.; Wang, Z.; Zhao, Y.; and Zhou, T. 2023 a . Core: Cooperative reconstruction for multi-agent perception. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8710--8720

  90. [100]

    Wang, C.; Deng, J.; He, J.; Zhang, T.; Zhang, Z.; and Zhang, Y. 2023 b . Long-short Range Adaptive Transformer with Dynamic Sampling for 3D Object Detection. IEEE Transactions on Circuits and Systems for Video Technology

  91. [101]

    Wang, J.; Shao, Y.; Ge, Y.; and Yu, R. 2019. A survey of vehicle to everything (V2X) testing. Sensors, 19(2): 334

  92. [102]

    Wang, L.; Liu, Y.; Du, P.; Ding, Z.; Liao, Y.; Qi, Q.; Chen, B.; and Liu, S. 2023 c . Object-aware distillation pyramid for open-vocabulary object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11186--11196

  93. [103]

    Wang, Q.; Wu, Y.; Yang, L.; Zuo, W.; and Hu, Q. 2024. Layer-Specific Knowledge Distillation for Class Incremental Semantic Segmentation. IEEE Transactions on Image Processing

  94. [104]

    Wang, R.; He, X.; Yu, R.; Qiu, W.; An, B.; and Rabinovich, Z. 2020 a . Learning efficient multi-agent communication: An information bottleneck approach. In International Conference on Machine Learning, 9908--9918. PMLR

  95. [105]

    Wang, T.; Chen, G.; Chen, K.; Liu, Z.; Zhang, B.; Knoll, A.; and Jiang, C. 2023 d . UMC: A Unified Bandwidth-efficient and Multi-resolution based Collaborative Perception Framework. arXiv preprint arXiv:2303.12400

  96. [106]

    Wang, T.; Hu, X.; Liu, Z.; and Fu, C.-W. 2022. Sparse2Dense: Learning to densify 3d features for 3d object detection. Advances in Neural Information Processing Systems, 35: 38533--38545

  97. [107]

    Wang, T.; Yuan, L.; Chen, Y.; Feng, J.; and Yan, S. 2021. Pnp-detr: Towards efficient visual analysis with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, 4661--4670

  98. [108]

    Wang, T.-H.; Manivasagam, S.; Liang, M.; Yang, B.; Zeng, W.; and Urtasun, R. 2020 b . V2vnet: Vehicle-to-vehicle communication for joint perception and prediction. In European Conference on Computer Vision, 605--621. Springer

  99. [109]

    Wang, X.; Yu, F.; Dou, Z.-Y.; Darrell, T.; and Gonzalez, J. E. 2018. Skipnet: Learning dynamic routing in convolutional networks. In Proceedings of the European Conference on Computer Vision (ECCV), 409--424

  100. [110]

    Wei, Y.; Wei, Z.; Rao, Y.; Li, J.; Zhou, J.; and Lu, J. 2022. Lidar distillation: Bridging the beam-induced domain gap for 3d object detection. In European Conference on Computer Vision, 179--195. Springer

  101. [111]

    Williams, R. J. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8: 229--256

  102. [112]

    Wu, Q.; Wang, K.; Li, K.; Zheng, J.; and Cai, J. 2023. Objectsdf++: Improved object-compositional neural implicit surfaces. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 21764--21774

  103. [113]

    Xia, X.; Hang, P.; Xu, N.; Huang, Y.; Xiong, L.; and Yu, Z. 2021. Advancing estimation accuracy of sideslip angle by fusing vehicle kinematics and dynamics information with fuzzy logic. IEEE Transactions on Vehicular Technology, 70(7): 6577--6590

  104. [114]

    Xu, R.; Chen, W.; Xiang, H.; Liu, L.; and Ma, J. 2023 a . Model-Agnostic Multi-Agent Perception Framework. arXiv:2203.13168

  105. [115]

    Xu, R.; Chen, W.; Xiang, H.; Xia, X.; Liu, L.; and Ma, J. 2023 b . Model-agnostic multi-agent perception framework. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 1471--1478. IEEE

  106. [116]

    Xu, R.; Guo, Y.; Han, X.; Xia, X.; Xiang, H.; and Ma, J. 2021. OpenCDA: an open cooperative driving automation framework integrated with co-simulation. In 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), 1155--1162. IEEE

  107. [117]

    Xu, R.; Tu, Z.; Xiang, H.; Shao, W.; Zhou, B.; and Ma, J. 2022 a . CoBEVT: Cooperative bird's eye view semantic segmentation with sparse transformers

  108. [118]

    Xu, R.; Xia, X.; Li, J.; Li, H.; Zhang, S.; Tu, Z.; Meng, Z.; Xiang, H.; Dong, X.; Song, R.; et al. 2023 c . V2v4real: A real-world large-scale dataset for vehicle-to-vehicle cooperative perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...

  109. [119]

    Xu, R.; Xiang, H.; Tu, Z.; Xia, X.; Yang, M.-H.; and Ma, J. 2022 b . V2X-ViT: Vehicle-to-everything cooperative perception with vision transformer. In European Conference on Computer Vision, 107–124. Springer

  110. [120]

    Xu, R.; Xiang, H.; Xia, X.; Han, X.; Li, J.; and Ma, J. 2022 c . Opv2v: An open benchmark dataset and fusion pipeline for perception with vehicle-to-vehicle communication. In 2022 International Conference on Robotics and Automation (ICRA), 2583--2589. IEEE

  111. [121]

    Yan, Y.; Mao, Y.; and Li, B. 2018. Second: Sparsely embedded convolutional detection. Sensors, 18(10): 3337

  112. [122]

    Yang, B.; Luo, W.; and Urtasun, R. 2018. Pixor: Real-time 3d object detection from point clouds. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 7652--7660

  113. [123]

    Yang, C.; Chen, Y.; Tian, H.; Tao, C.; Zhu, X.; Zhang, Z.; Huang, G.; Li, H.; Qiao, Y.; Lu, L.; Zhou, J.; and Dai, J. 2022 a . BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision. [Online]. Available: https://arxiv.org/abs/2...

  114. [124]

    Yang, K.; Sun, P.; Lin, J.; Boukerche, A.; and Song, L. 2022 b . A Novel Distributed Task Scheduling Framework for Supporting Vehicular Edge Intelligence. In 2022 IEEE 42nd International Conference on Distributed Computing Systems (ICDCS), 972--982

  115. [125]

    Yang, K.; Sun, P.; Lin, J.; Boukerche, A.; and Song, L. 2022 c . A novel distributed task scheduling framework for supporting vehicular edge intelligence. In 2022 IEEE 42nd International Conference on Distributed Computing Systems (ICDCS), 972--982. IEEE

  116. [126]

    Yang, K.; Yang, D.; Zhang, J.; Li, M.; Liu, Y.; Liu, J.; Wang, H.; Sun, P.; and Song, L. 2023 a . Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception. arXiv:2307.13929

  117. [127]

    Yang, K.; Yang, D.; Zhang, J.; Wang, H.; Sun, P.; and Song, L. 2023 b . What2comm: Towards Communication-Efficient Collaborative Perception via Feature Decoupling. In Proceedings of the 31st ACM International Conference on Multimedia, MM '23, 7686–7695. New York, NY, USA: Asso...

  118. [128]

    Y.; and Juang, B.-H

    Ye, H.; Li, G. Y.; and Juang, B.-H. F. 2019. Deep reinforcement learning based resource allocation for V2V communications. IEEE Transactions on Vehicular Technology, 68(4): 3163--3173

  119. [129]

    Yu, H.; Luo, Y.; Shu, M.; Huo, Y.; Yang, Z.; Shi, Y.; Guo, Z.; Li, H.; Hu, X.; Yuan, J.; et al. 2022. DAIR-V2X: A Large-Scale Dataset for Vehicle-Infrastructure Cooperative 3D Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  120. [130]

    Yu, H.; Yang, W.; Zhong, J.; Yang, Z.; Fan, S.; Luo, P.; and Nie, Z. 2024. End-to-End Autonomous Driving through V2X Cooperation. arXiv:2404.00717

  121. [131]

    Yuan, Z.; Song, X.; Bai, L.; Wang, Z.; and Ouyang, W. 2021. Temporal-channel transformer for 3d lidar-based video object detection for autonomous driving. IEEE Transactions on Circuits and Systems for Video Technology, 32(4): 2068--2078

  122. [132]

    Zhang, H.; Wu, C.; Zhang, Z.; Zhu, Y.; Lin, H.; Zhang, Z.; Sun, Y.; He, T.; Mueller, J.; Manmatha, R.; Li, M.; and Smola, A. 2022. ResNeSt: Split-Attention Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2736--2746

  123. [133]

    Zhang, J.; Yang, K.; Wang, H.; Sun, P.; and Song, L. 2024 a . Efficient Vehicular Collaborative Perception Based on Saptial-Temporal Feature Compression. IEEE Transactions on Vehicular Technology, 73(11): 16125--16133

  124. [134]

    Zhang, J.; Yang, K.; Wang, Y.; Wang, H.; Sun, P.; and Song, L. 2024 b . ERMVP: Communication-Efficient and Collaboration-Robust Multi-Vehicle Perception in Challenging Environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12575--12584

  125. [135]

    Zhang, L.; and Ma, K. 2023. Structured knowledge distillation for accurate and efficient object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence

  126. [136]

    Q.; Zhang, Q.; and Lin, J

    Zhang, S. Q.; Zhang, Q.; and Lin, J. 2020. Succinct and robust multi-agent communication with temporal message control. Advances in Neural Information Processing Systems, 33: 17271--17282

  127. [137]

    Zhang, Z.; and Fisac, J. F. 2021. Safe occlusion-aware autonomous driving via game-theoretic active perception. arXiv preprint arXiv:2105.08169

  128. [138]

    Zhou, B.; and Kr\"ahenb\"uhl, P. 2022. Cross-View Transformers for Real-Time Map-View Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13760--13769

  129. [139]

    Zhou, Y.; and Tuzel, O. 2018. Voxelnet: End-to-end learning for point cloud based 3d object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4490--4499

  130. [140]

    Zhu, C.; Li, L.; Wu, Y.; and Sun, Z. 2024. Saswot: Real-time semantic segmentation architecture search without training. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 7722--7730

  131. [141]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  132. [142]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.