REVIEW 4 major objections 6 minor 228 references
This survey claims that every semantic communication system for visual data can be classified by whether it preserves, expands, or refines transmitted semantics, and that this choice guides the entire system design.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 06:41 UTC pith:QVKPT7NT
load-bearing objection A broad, useful survey of SemCom-Vision whose central SPC/SEC/SRC taxonomy is intuitive but not tightly defined; the categories overlap under the paper's own information-theoretic gloss, and Ref [1] is mis-cited. the 4 major comments →
A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On its own terms, the survey's central claim is that SemCom-Vision approaches can be meaningfully categorized by communication goal as semantic preservation communication (SPC), semantic expansion communication (SEC), or semantic refinement communication (SRC). SPC aims to keep pixel- and semantic-level fidelity; SEC enriches received content with context or generative detail; SRC transmits only task-relevant semantics. The paper argues these goals are reflected in different semantic quantization schemes (entropy and mutual information, label embeddings, and knowledge graphs), different ML architectures and losses, and different knowledge structures, so the taxonomy doubles as a selection gu
What carries the argument
The load-bearing mechanism is the semantic quantization scheme—the way a system decides what counts as meaning—expressed three ways: information-theoretic (semantic entropy, mutual information, information bottleneck), labeling-based (vision-language embeddings, attention weights, classification probabilities), and knowledge-driven (nodes and edges of a knowledge graph). From these, the paper derives the SPC/SEC/SRC classification and uses it to organize encoder-decoder construction, loss functions, and knowledge-graph choices.
Load-bearing premise
The classification assumes every SemCom-Vision system can be assigned to exactly one of the three goals—preserve, expand, or refine semantics—but many real systems pursue more than one of these goals at once.
What would settle it
Take the systems surveyed here and ask independent reviewers to classify each as SPC, SEC, or SRC; the central claim fails if a substantial share of systems is placed in two categories or in none, since the paper provides no quantitative decision rule to resolve such ties.
If this is right
- Researchers can match encoder-decoder architectures to their communication goal: fidelity-oriented systems lean on convolutional and autoencoder designs, expansion-oriented systems on generative diffusion and GAN models, and refinement-oriented systems on transformers and attention-based filtering.
- Loss functions align with the same goals: pixel and perceptual losses for preservation, denoising and adversarial losses for expansion, and information-bottleneck, sparsity, and task-specific losses for refinement.
- Knowledge-graph structure choices become goal-dependent: triple stores suit refinement through logical reasoning, property graphs suit flexible semantic guidance, and hypergraphs suit complex multi-entity scenes.
- Application domains such as digital twins, metaverse, and wireless perception can be mapped to the three categories, so designers can select SPC, SRC, or SEC per application need.
- The survey's framework provides a shared vocabulary for comparing future SemCom-Vision systems and for spotting under-explored combinations, such as knowledge-driven quantization for visual tasks.
Where Pith is reading between the lines
- The three categories may be better viewed as regions on a spectrum than as discrete bins: a refinement system preserves the subset of semantics it deems task-relevant, so future work could formalize the boundaries using measurable changes in semantic entropy or mutual information.
- The same preserve-expand-refine logic likely extends beyond vision to multimodal and text-oriented semantic communication, where generative decoders similarly add or strip meaning.
- A testable design rule follows from the paper: given a fixed bandwidth and task, the optimal category depends on the task's tolerance for semantic loss, which could be quantified with downstream task accuracy rather than pixel fidelity.
- If the taxonomy were operationalized, it could feed into automated system selection—an agent could inspect the communication goal and choose architecture, loss, and knowledge structure without human intervention.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys semantic communication for visual data (SemCom-Vision). Its main contribution is a taxonomy that classifies existing systems into semantic preservation communication (SPC), semantic expansion communication (SEC), and semantic refinement communication (SRC) based on communication goals viewed through semantic quantization schemes. It then reviews ML-based encoder-decoder architectures, training losses, knowledge graph structures, and applications. The survey is broad and well-structured, but the taxonomy's boundaries are not sharp enough to support the claimed 'clear categorization'.
Significance. If the taxonomy could be made operational, it would provide a useful organizing principle for a rapidly growing and fragmented literature. The survey's coverage of ML architectures, loss functions, and especially the knowledge graph section is valuable as a reference. The authors' effort to connect information-theoretic quantities (semantic entropy, mutual information) to practical design choices is interesting and timely. The paper is also honest about several open problems in the field. However, the central classification currently lacks a decision rule; the overlap between SPC and SRC, in particular, means that Table III assignments are not uniquely determined. This is a solvable problem but requires either a formal criterion or an explicit treatment of hybrid systems.
major comments (4)
- [Section III-B, Table III, Section IV-C] The SPC/SEC/SRC categories are defined by qualitative goals, but the information-theoretic gloss in Section III-B does not make them exclusive. SRC 'deliberately reduces H(s)' and maximizes I(s; s_hat_T); any SRC system that faithfully transmits the task-relevant subset also preserves that subset's semantics and thus satisfies the SPC criterion on that subset. Table III assigns each reference to one category, and Section IV uses the categories to recommend architectures and losses (e.g., L9 in Eq. (11) reconstructs only Omega, yet the selected region is preserved). Without a formal decision rule or an explicit treatment of hybrid systems, the same paper can be placed in multiple buckets, so the claimed 'clear categorization' and the model-selection guidance are not uniquely determined.
- [Section III-A, Eqs. (1)-(2), Section III-B] The semantic entropy H(s) and mutual information I(s; s_hat) are used to differentiate SPC, SEC, and SRC, but the paper does not specify how p(s_k) or s_hat are obtained for a given system. SEC is said to 'intentionally increase the entropy beyond that of the source'; no criterion is given to decide when an added semantic is new versus a reconstruction artifact. Without operational definitions, the proposed 'semantic quantization schemes' cannot be used as a decision procedure to classify existing systems or to derive testable predictions. Please provide at least a minimal protocol (e.g., what random variables are defined, how the task target T is formalized) or restrict the claims to a qualitative taxonomy.
- [Table III, Section IV] The survey assigns dozens of references to SPC/SEC/SRC in Table III but provides no annotation protocol, no decision tree, and no inter-annotator reliability check. For example, [144] is placed in SEC because it uses a GAN to generate artistically styled images from segmentation maps, while [131] is placed in SPC despite also using a GAN and a discriminator; the boundary rests on a qualitative reading of the system's goal. Since the taxonomy is the paper's main contribution, the authors should explain how a reader can independently reproduce the assignments, or alternatively present the categories as overlapping design dimensions rather than a partition.
- [Section IV-A, Eq. (6)] The VAE training loss L4 is written as reconstruction error minus the KL-divergence term, i.e., L4 = MSE - (1/2) Sum(1 + log sigma^2 - mu^2 - sigma^2). The standard negative ELBO has a plus sign before the KL divergence; the minus sign here drives the encoder to maximize variance and is not the variational loss used in the cited works. Please correct the sign or clarify that this is a different objective.
minor comments (6)
- [Sections IV-A, IV-B, IV-C] The text repeatedly refers to 'Table IV' for the overview of SemCom encoder-decoder models, but the relevant table is Table III. This cross-reference error will confuse readers.
- [Section II-A] Typo: 'ReLu' should be 'ReLU'.
- [Figure 1 caption] Typo: 'Potiential Applications' should be 'Potential Applications'. Also the caption lists both 'V. Potiential Applications' and 'VI. Potential Applications', which appear to be duplicate labels.
- [Section V-A] Typo: 'TIV A-KG' should be 'TIVA-KG'.
- [Section III-A] Typo: 'UA V-photoed' should be 'UAV-photoed' (or with a space removed).
- [Eq. (9)] The expectation subscript is malformed: 'E s,epsilon~N(0,1)' should be written as E_{s, epsilon~N(0,1)}.
Circularity Check
No significant circularity: the SPC/SEC/SRC taxonomy is an interpretive classification, not a derivation that reduces to its inputs.
full rationale
The paper's central contribution is a survey taxonomy. SPC/SEC/SRC are introduced as definitions in Section III-B ('Semantic Preservation Communication (SPC): Transmit extracted semantics while mitigating channel-induced semantic loss...'; 'Semantic Expansion Communication (SEC): Extract semantics... enrich transmitted semantics...'; 'Semantic Refinement Communication (SRC): Extract and transmit only task-relevant semantics...'), and the rest of the paper maps existing systems onto these categories. This is an act of interpretation, not a derivation: no quantity is fitted and then reported as a prediction, and no mathematical claim is reduced by construction to its own assumptions. The information-theoretic gloss in Section III-B (SPC maximizes H(s)/I(s;s-hat), SRC applies IB to reduce H(s)) is a post-hoc characterization of the categories, not an input used to generate them. The self-citations used as examples ([20], [22], [25], [66], [84]) are illustrative and not load-bearing; the taxonomy does not depend on any uniqueness theorem or unverified prior claim by the authors. The objection that the categories lack a formal decision rule and may overlap (e.g., an SRC system that faithfully transmits the task-relevant subset also preserves mutual information over that subset) is a validity/usability limitation of the classification, but it is not circularity. Under the required standard of exhibiting a specific reduction, no circular step is present.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Semantics are independent: p(s) = ∏ p(s_k) (Eq. 1).
- ad hoc to paper All SemCom-Vision systems can be partitioned into SPC, SEC, SRC (Section III-B).
- domain assumption Semantic entropy is 'the fundamental measure of semantics' (Section III-A).
read the original abstract
Semantic communication (SemCom) emerges as a transformative paradigm for traffic-intensive visual data transmission, shifting focus from raw data to meaningful content transmission and relieving the increasing pressure on communication resources. However, to achieve SemCom, challenges are faced in accurate semantic quantization for visual data, robust semantic extraction and reconstruction under diverse tasks and goals, transceiver coordination with effective knowledge utilization, and adaptation to unpredictable wireless communication environments. In this paper, we present a systematic review of SemCom for visual data transmission (SemCom-Vision), wherein an interdisciplinary analysis integrating computer vision (CV) and communication engineering is conducted to provide comprehensive guidelines for the machine learning (ML)-empowered SemCom-Vision design. Specifically, this survey first elucidates the basics and key concepts of SemCom. Then, we introduce a novel classification perspective to categorize existing SemCom-Vision approaches as semantic preservation communication (SPC), semantic expansion communication (SEC), and semantic refinement communication (SRC) based on communication goals interpreted through semantic quantization schemes. Moreover, this survey articulates the ML-based encoder-decoder models and training algorithms for each SemCom-Vision category, followed by knowledge structure and utilization strategies. Finally, we discuss potential SemCom-Vision applications.
Figures
Reference graph
Works this paper leans on
-
[1]
The Synthesis Between Artificial Intelligence and Editing Stories of the Future,
S. Erdem, “The Synthesis Between Artificial Intelligence and Editing Stories of the Future,” inTransforming cinema with artificial intelli- gence. IGI Global Scientific Publishing, 2025, pp. 221–240. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 20
2025
-
[2]
Transmit what you need: task-adaptive semantic communications for visual information,
J. Park and S. W. Yoon, “Transmit what you need: task-adaptive semantic communications for visual information,”arXiv preprint arXiv:2412.13646, 2024
arXiv 2024
-
[3]
Vpr-bench: An open-source visual place recogni- tion evaluation framework with quantifiable viewpoint and appearance change,
M. Zaffar, S. Garg, M. Milford, J. Kooij, D. Flynn, K. McDonald- Maier, and S. Ehsan, “Vpr-bench: An open-source visual place recogni- tion evaluation framework with quantifiable viewpoint and appearance change,”International Journal of Computer Vision, vol. 129, no. 7, pp. 2136–2174, 2021
2021
-
[4]
A Survey on Semantic Communications for Intelligent Wireless Networks,
S. Iyer, R. Khanai, D. Torse, R. J. Pandya, K. M. Rabie, K. Pai, W. U. Khan, and Z. Fadlullah, “A Survey on Semantic Communications for Intelligent Wireless Networks,”Wireless Personal Communications, vol. 129, no. 1, pp. 569–611, 2023
2023
-
[5]
Semantic Communications for Future Internet: Fundamentals, Applications, and Challenges,
W. Yang, H. Du, Z. Q. Liew, W. Y . B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Shen, and C. Miao, “Semantic Communications for Future Internet: Fundamentals, Applications, and Challenges,”IEEE Communications Surveys & Tutorials, vol. 25, no. 1, pp. 213–250, 2022
2022
-
[6]
Semantics-Empowered Communications: A Tutorial-Cum-Survey,
Z. Lu, R. Li, K. Lu, X. Chen, E. Hossain, Z. Zhao, and H. Zhang, “Semantics-Empowered Communications: A Tutorial-Cum-Survey,” IEEE Communications Surveys & Tutorials, vol. 26, no. 1, pp. 41– 79, 2023
2023
-
[7]
Less Data, More Knowledge: Building Next Generation Semantic Communication Networks,
C. Chaccour, W. Saad, M. Debbah, Z. Han, and H. V . Poor, “Less Data, More Knowledge: Building Next Generation Semantic Communication Networks,”IEEE Communications Surveys & Tutorials, 2024
2024
-
[8]
Intellicise Wireless Networks from Semantic Communications: A Survey, Research Issues, and Challenges,
P. Zhang, W. Xu, Y . Liu, X. Qin, K. Niu, S. Cui, G. Shi, Z. Qin, X. Xu, F. Wanget al., “Intellicise Wireless Networks from Semantic Communications: A Survey, Research Issues, and Challenges,”IEEE Communications Surveys & Tutorials, 2024
2024
-
[9]
A Survey on Goal-Oriented Semantic Communication: Techniques, Challenges, and Future Direc- tions,
T. M. Getu, G. Kaddoum, and M. Bennis, “A Survey on Goal-Oriented Semantic Communication: Techniques, Challenges, and Future Direc- tions,”IEEE Access, 2024
2024
-
[10]
Semantic Communication: A Survey on Research Landscape, Challenges, and Future Directions,
——, “Semantic Communication: A Survey on Research Landscape, Challenges, and Future Directions,”Proceedings of the IEEE, 2025
2025
-
[11]
A Review of Convolutional Neural Networks in Computer Vision,
X. Zhao, L. Wang, Y . Zhang, X. Han, M. Deveci, and M. Parmar, “A Review of Convolutional Neural Networks in Computer Vision,” Artificial Intelligence Review, vol. 57, no. 4, p. 99, 2024
2024
-
[12]
Visual Analytics for Machine Learn- ing: A Data Perspective Survey,
J. Wang, S. Liu, and W. Zhang, “Visual Analytics for Machine Learn- ing: A Data Perspective Survey,”IEEE transactions on visualization and computer graphics, vol. 30, no. 12, pp. 7637–7656, 2024
2024
-
[13]
A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and Application,
B. Yuan and D. Zhao, “A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and Application,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
2024
-
[14]
E. Hosonuma, T. Yamazaki, T. Miyoshi, A. Taya, Y . Nishiyama, and K. Sezaki, “Image Generative Semantic Communication With Multi- Modal Similarity Estimation for Resource-Limited Networks,”arXiv preprint arXiv:2404.11280, 2024
Pith/arXiv arXiv 2024
-
[15]
X. Qiu and R. Miikkulainen, “Semantic Density: Uncertainty Quantifi- cation for Large Language Models Through Confidence Measurement in Semantic Space,”arXiv preprint arXiv:2405.13845, 2024
Pith/arXiv arXiv 2024
-
[16]
Feature Importance-Aware Task-Oriented Semantic Transmission and Optimization,
Y . Wang, S. Han, X. Xu, H. Liang, R. Meng, C. Dong, and P. Zhang, “Feature Importance-Aware Task-Oriented Semantic Transmission and Optimization,”IEEE Transactions on Cognitive Communications and Networking, 2024
2024
-
[17]
Task-Oriented Explainable Semantic Communications,
S. Ma, W. Qiao, Y . Wu, H. Li, G. Shi, D. Gao, Y . Shi, S. Li, and N. Al- Dhahir, “Task-Oriented Explainable Semantic Communications,”IEEE transactions on wireless communications, vol. 22, no. 12, pp. 9248– 9262, 2023
2023
-
[18]
Vector Quantized Semantic Communication System,
Q. Fu, H. Xie, Z. Qin, G. Slabaugh, and X. Tao, “Vector Quantized Semantic Communication System,”IEEE Wireless Communications Letters, vol. 12, no. 6, pp. 982–986, 2023
2023
-
[19]
Federated Semantic Learning Driven by Information Bottleneck for Task-Oriented Communications,
H. Wei, W. Ni, W. Xu, F. Wang, D. Niyato, and P. Zhang, “Federated Semantic Learning Driven by Information Bottleneck for Task-Oriented Communications,”IEEE Communications Letters, vol. 27, no. 10, pp. 2652–2656, 2023
2023
-
[20]
R. Cheng, Y . Sun, L. Zhang, L. Feng, L. Zhang, and M. A. Imran, “A Semantic Communication-Based Workload-Adjustable Transceiver for Wireless Ai-Generated Content (AIGC) Delivery,”arXiv preprint arXiv:2503.18874, 2025
Pith/arXiv arXiv 2025
-
[21]
Background Knowledge Aware Semantic Coding Model Selection,
F. Zhao, Y . Sun, R. Cheng, and M. A. Imran, “Background Knowledge Aware Semantic Coding Model Selection,” in2022 IEEE 22nd Inter- national Conference on Communication Technology (ICCT). IEEE, 2022, pp. 84–89
2022
-
[22]
Vista: Video Transmission Over a Semantic Communication Approach,
C. Liang, X. Deng, Y . Sun, R. Cheng, L. Xia, D. Niyato, and M. A. Imran, “Vista: Video Transmission Over a Semantic Communication Approach,” in2023 IEEE International Conference on Communica- tions Workshops (ICC Workshops). IEEE, 2023, pp. 1777–1782
2023
-
[23]
VISTA: A Semantic Communication Approach for Video Transmission,
C. Liang, X. Deng, Y . Sun, R. Cheng, L. Xia, D. Niyato, and M. Ali Imran, “VISTA: A Semantic Communication Approach for Video Transmission,”Wireless Semantic Communications: Concepts, Principles and Challenges, pp. 109–121, 2025
2025
-
[24]
Deep learning enabled semantic communication systems,
H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,”IEEE transactions on signal pro- cessing, vol. 69, pp. 2663–2675, 2021
2021
-
[25]
A wireless ai-generated content (aigc) provisioning framework em- powered by semantic communication,
R. Cheng, Y . Sun, D. Niyato, L. Zhang, L. Zhang, and M. Imran, “A wireless ai-generated content (aigc) provisioning framework em- powered by semantic communication,”IEEE Transactions on Mobile Computing, 2024
2024
-
[26]
When Wireless Communications Meet Computer Vision in Beyond 5G,
T. Nishio, Y . Koda, J. Park, M. Bennis, and K. Doppler, “When Wireless Communications Meet Computer Vision in Beyond 5G,”IEEE Communications Standards Magazine, vol. 5, no. 2, pp. 76–83, 2021
2021
-
[27]
Applying Deep-Learning-Based Computer Vision to Wireless Communications: Methodologies, Oppor- tunities, and Challenges,
Y . Tian, G. Pan, and M.-S. Alouini, “Applying Deep-Learning-Based Computer Vision to Wireless Communications: Methodologies, Oppor- tunities, and Challenges,”IEEE Open Journal of the Communications Society, vol. 2, pp. 132–143, 2020
2020
-
[28]
E. R. Davies,Computer Vision: Principles, Algorithms, Applications, Learning. Academic Press, 2017
2017
-
[29]
A Review on Machine Learning Styles in Computer Vision—Techniques and Future Directions,
S. V . Mahadevkar, B. Khemani, S. Patil, K. Kotecha, D. R. V ora, A. Abraham, and L. A. Gabralla, “A Review on Machine Learning Styles in Computer Vision—Techniques and Future Directions,”Ieee Access, vol. 10, pp. 107 293–107 329, 2022
2022
-
[30]
Deep Learning vs. Traditional Computer Vision,
N. O’Mahony, S. Campbell, A. Carvalho, S. Harapanahalli, G. V . Hernandez, L. Krpalkova, D. Riordan, and J. Walsh, “Deep Learning vs. Traditional Computer Vision,” inAdvances in computer vision: proceedings of the 2019 computer vision conference (CVC), volume 1 1. Springer, 2020, pp. 128–144
2019
-
[31]
S- ran: Semantic-aware radio access networks,
Y . Sun, L. Zhang, L. Guo, J. Li, D. Niyato, and Y . Fang, “S- ran: Semantic-aware radio access networks,”IEEE Communications Magazine, 2024
2024
-
[32]
Machine Learning Enabled Heterogeneous Semantic and Bit Communication,
M. Zhang, R. Zhong, X. Mu, and Y . Liu, “Machine Learning Enabled Heterogeneous Semantic and Bit Communication,”IEEE Transactions on Wireless Communications, 2024
2024
-
[33]
Semantic Entropy Can Simultaneously Benefit Transmission Efficiency and Channel Security of Wireless Se- mantic Communications,
Y . Rong, G. Nan, M. Zhang, S. Chen, S. Wang, X. Zhang, N. Ma, S. Gong, Z. Yang, Q. Cuiet al., “Semantic Entropy Can Simultaneously Benefit Transmission Efficiency and Channel Security of Wireless Se- mantic Communications,”IEEE Transactions on Information Forensics and Security, 2025
2025
-
[34]
Niu and P
K. Niu and P. Zhang,The Mathematical Theory of Semantic Commu- nication. Springer Nature, 2025
2025
-
[35]
Advanced Semantics for Commonsense Knowledge Extraction,
T.-P. Nguyen, S. Razniewski, and G. Weikum, “Advanced Semantics for Commonsense Knowledge Extraction,” inProceedings of the Web Conference 2021, 2021, pp. 2636–2647
2021
-
[36]
A Continual Learning Survey: Defy- ing Forgetting in Classification Tasks,
M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars, “A Continual Learning Survey: Defy- ing Forgetting in Classification Tasks,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 7, pp. 3366–3385, 2021
2021
-
[37]
Resource Management, Security, and Privacy Issues in Semantic Communications: A Survey,
D. Won, G. Woraphonbenjakul, A. B. Wondmagegn, A. T. Tran, D. Lee, D. S. Lakew, and S. Cho, “Resource Management, Security, and Privacy Issues in Semantic Communications: A Survey,”IEEE Communications Surveys & Tutorials, 2024
2024
-
[38]
Toward Intelli- gent Resource Allocation on Task-Oriented Semantic Communication,
H. Zhang, H. Wang, Y . Li, K. Long, and V . C. Leung, “Toward Intelli- gent Resource Allocation on Task-Oriented Semantic Communication,” IEEE Wireless Communications, vol. 30, no. 3, pp. 70–77, 2023
2023
-
[39]
A Survey on Semantic Communication Networks: Architecture, Security, and Privacy,
S. Guo, Y . Wang, N. Zhang, Z. Su, T. H. Luan, Z. Tian, and X. Shen, “A Survey on Semantic Communication Networks: Architecture, Security, and Privacy,”IEEE Communications Surveys & Tutorials, 2024
2024
-
[40]
What Is Semantic Communication? A View on Conveying Meaning in the Era of Machine Intelligence,
Q. Lan, D. Wen, Z. Zhang, Q. Zeng, X. Chen, P. Popovski, and K. Huang, “What Is Semantic Communication? A View on Conveying Meaning in the Era of Machine Intelligence,”Journal of Communica- tions and Information Networks, vol. 6, no. 4, pp. 336–371, 2021
2021
-
[41]
A Survey of Con- volutional Neural Networks: Analysis, Applications, and Prospects,
Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A Survey of Con- volutional Neural Networks: Analysis, Applications, and Prospects,” Ieee Transactions on Neural Networks and Learning Systems, vol. 33, no. 12, pp. 6999–7019, 2021
2021
-
[42]
An Equivalence of Fully Connected Layer and Convolutional Layer,
W. Ma and J. Lu, “An Equivalence of Fully Connected Layer and Convolutional Layer,”arXiv preprint arXiv:1712.01252, 2017
Pith/arXiv arXiv 2017
-
[43]
Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification,
S. S. Basha, S. R. Dubey, V . Pulabaigari, and S. Mukherjee, “Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification,”Neurocomputing, vol. 378, pp. 112–119, 2020
2020
-
[44]
The Unreasonable Effectiveness of Fully-Connected Layers for Low- Data Regimes,
P. Kocsis, P. S ´uken´ık, G. Bras´o, M. Nießner, L. Leal-Taix´e, and I. Elezi, “The Unreasonable Effectiveness of Fully-Connected Layers for Low- Data Regimes,”Advances in Neural Information Processing Systems, vol. 35, pp. 1896–1908, 2022. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 21
1908
-
[45]
Fully convolutional networks for semantic segmentation,
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440
2015
-
[46]
An Introduction to Convolutional Neural Networks,
K. O’shea and R. Nash, “An Introduction to Convolutional Neural Networks,”arXiv preprint arXiv:1511.08458, 2015
Pith/arXiv arXiv 2015
-
[47]
Understanding Convolution for Semantic Segmentation,
P. Wang, P. Chen, Y . Yuan, D. LIU, Z. Huang, X. Hou, and G. Cottrell, “Understanding Convolution for Semantic Segmentation,” in2018 Ieee Winter Conference on Applications of Computer Vision (Wacv). Ieee, 2018, pp. 1451–1460
2018
-
[48]
Activation Functions: Comparison of Trends in Practice and Research for Deep Learning,
C. Nwankpa, W. Ijomah, A. Gachagan, and S. Marshall, “Activation Functions: Comparison of Trends in Practice and Research for Deep Learning,”arXiv preprint arXiv:1811.03378, 2018
Pith/arXiv arXiv 2018
-
[49]
FinBack: Infiltrating Backdoors into Gradient Compressors on Federated Learning,
X. Tang, W. Yang, L. Peng, M. Shen, T. Zhang, Y . Weng, J. Kang, and D. Niyato, “FinBack: Infiltrating Backdoors into Gradient Compressors on Federated Learning,”IEEE Transactions on Information Forensics and Security, 2025
2025
-
[50]
ROBY: A Byzantine-Robust and Privacy-Preserving Serverless Federated Learning Framework,
X. Tang, M. Li, M. Shen, J. Kang, L. Zhu, Z. Liu, G. Yang, D. Niyato, and R. H. Deng, “ROBY: A Byzantine-Robust and Privacy-Preserving Serverless Federated Learning Framework,”IEEE Transactions on Information Forensics and Security, 2025
2025
-
[51]
Pile: Robust Privacy-Preserving Federated Learning via Verifiable Perturbations,
X. Tang, M. Shen, Q. Li, L. Zhu, T. Xue, and Q. Qu, “Pile: Robust Privacy-Preserving Federated Learning via Verifiable Perturbations,” IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 6, pp. 5005–5023, 2023
2023
-
[52]
Nonparametric Regression Using Deep Neu- ral Networks With Relu Activation Function,
J. SCHMIDT-HIEBER, “Nonparametric Regression Using Deep Neu- ral Networks With Relu Activation Function,”The Annals of Statistics, vol. 48, no. 4, pp. 1875–1897, 2020
2020
-
[53]
Why TanH Is a Hardware Friendly Activation Function for CNNs,
K. Abdelouahab, M. Pelcat, and F. Berry, “Why TanH Is a Hardware Friendly Activation Function for CNNs,” inProceedings of the 11th international conference on distributed smart cameras, 2017, pp. 199– 201
2017
-
[54]
The Most Used Activation Functions: Classic Versus Current,
M. A. Mercioni and S. Holban, “The Most Used Activation Functions: Classic Versus Current,” in2020 International Conference on Devel- opment and Application Systems (DAS). IEEE, 2020, pp. 141–145
2020
-
[55]
Mixed Pooling for Convolu- tional Neural Networks,
D. Yu, H. Wang, P. Chen, and Z. Wei, “Mixed Pooling for Convolu- tional Neural Networks,” inRough Sets and Knowledge Technology: 9th International Conference, RSKT 2014, Shanghai, China, October 24-26, 2014, Proceedings 9. Springer, 2014, pp. 364–375
2014
-
[56]
The Treasure Beneath Convolutional Layers: Cross-Convolutional-Layer Pooling for Image Classification,
L. Liu, C. Shen, and A. Van den Hengel, “The Treasure Beneath Convolutional Layers: Cross-Convolutional-Layer Pooling for Image Classification,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4749–4757
2015
-
[57]
Normalization Techniques in Training Dnns: Methodology, Analysis and Applica- tion,
L. Huang, J. Qin, Y . Zhou, F. Zhu, L. Liu, and L. Shao, “Normalization Techniques in Training Dnns: Methodology, Analysis and Applica- tion,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 8, pp. 10 173–10 196, 2023
2023
-
[58]
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,”arXiv preprint arXiv:1607.06450, 2016
Pith/arXiv arXiv 2016
-
[59]
A Review on the Attention Mechanism of Deep Learning,
Z. Niu, G. Zhong, and H. Yu, “A Review on the Attention Mechanism of Deep Learning,”Neurocomputing, vol. 452, pp. 48–62, 2021
2021
-
[60]
Attention Is All You Need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” Advances in neural information processing systems, vol. 30, 2017
2017
-
[61]
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gellyet al., “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,”arXiv preprint arXiv:2010.11929, 2020
Pith/arXiv arXiv 2010
-
[62]
Resnet in Resnet: Generalizing Residual Architectures,
S. Targ, D. Almeida, and K. Lyman, “Resnet in Resnet: Generalizing Residual Architectures,”arXiv preprint arXiv:1603.08029, 2016
Pith/arXiv arXiv 2016
-
[63]
Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks With Symmetric Skip Connections,
X. Mao, C. Shen, and Y .-B. Yang, “Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks With Symmetric Skip Connections,”Advances in neural information processing systems, vol. 29, 2016
2016
-
[64]
U-Net: Convolutional Networks for Biomedical Image Segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” inMedical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, pro- ceedings, part III 18. Springer, 2015, pp. 234–241
2015
-
[65]
Semantic Communi- cations: Principles and Challenges,
Z. Qin, X. Tao, J. Lu, W. Tong, and G. Y . Li, “Semantic Communi- cations: Principles and Challenges,”arXiv preprint arXiv:2201.01389, 2021
Pith/arXiv arXiv 2021
-
[66]
Knowledge-Assisted Privacy Preserving in Semantic Communication,
X. Liu, Y . Sun, R. Cheng, L. Xia, H. Abumarshoud, L. Zhang, and M. A. Imran, “Knowledge-Assisted Privacy Preserving in Semantic Communication,”IEEE Wireless Communications, vol. 32, no. 2, pp. 76–83, 2025
2025
-
[67]
Semantic Communications with Computer Vision Sensing for Edge Video Trans- mission,
Y . Peng, L. Xiang, K. Yang, K. Wang, and M. Debbah, “Semantic Communications with Computer Vision Sensing for Edge Video Trans- mission,”arXiv preprint arXiv:2503.07252, 2025
arXiv 2025
-
[68]
Lightweight Vision Model-based Multi-user Semantic Communication Systems,
F. Jiang, S. Tu, L. Dong, K. Wang, K. Yang, R. Liu, C. Pan, and J. Wang, “Lightweight Vision Model-based Multi-user Semantic Communication Systems,”arXiv preprint arXiv:2502.16424, 2025
Pith/arXiv arXiv 2025
-
[69]
Personalized Fed- erated Learning for GAI-Assisted Semantic Communications,
Y . Peng, F. Jiang, L. Dong, K. Wang, and K. Yang, “Personalized Fed- erated Learning for GAI-Assisted Semantic Communications,”IEEE Transactions on Cognitive Communications and Networking, 2025
2025
-
[70]
A Lite Distributed Semantic Communication System for Internet of Things,
H. Xie and Z. Qin, “A Lite Distributed Semantic Communication System for Internet of Things,”IEEE Journal on Selected Areas in Communications, vol. 39, no. 1, pp. 142–153, 2020
2020
-
[71]
Deep Joint Source-Channel Coding for Semantic Communications,
J. Xu, T.-Y . Tung, B. Ai, W. Chen, Y . Sun, and D. G ¨und¨uz, “Deep Joint Source-Channel Coding for Semantic Communications,”IEEE communications Magazine, vol. 61, no. 11, pp. 42–48, 2023
2023
-
[72]
Innovative Semantic Communication System,
C. Dong, H. Liang, X. Xu, S. Han, B. Wang, and P. Zhang, “Innovative Semantic Communication System,”arXiv preprint arXiv:2202.09595, 2022
Pith/arXiv arXiv 2022
-
[73]
Joint Source-Channel Cod- ing for Channel-Adaptive Digital Semantic Communications,
J. Park, Y . Oh, S. Kim, and Y .-S. Jeon, “Joint Source-Channel Cod- ing for Channel-Adaptive Digital Semantic Communications,”IEEE Transactions on Cognitive Communications and Networking, 2024
2024
-
[74]
Channel Coding Towards 6G: Technical Overview and Outlook,
M. Rowshan, M. Qiu, Y . Xie, X. Gu, and J. Yuan, “Channel Coding Towards 6G: Technical Overview and Outlook,”IEEE Open Journal of the Communications Society, 2024
2024
-
[75]
Improved Nonlinear Transform Source-Channel Coding to Catalyze Semantic Communications,
S. Wang, J. Dai, X. Qin, Z. Si, K. Niu, and P. Zhang, “Improved Nonlinear Transform Source-Channel Coding to Catalyze Semantic Communications,”IEEE Journal of Selected Topics in Signal Process- ing, vol. 17, no. 5, pp. 1022–1037, 2023
2023
-
[76]
Power Allocation for Throughput Maximization in NOMA-Based Semantic Communication System,
K. Ma, H. Abumarshoud, S. Hua, M. Imran, and Y . Sun, “Power Allocation for Throughput Maximization in NOMA-Based Semantic Communication System,” inICC 2025-IEEE International Conference on Communications. IEEE, 2025, pp. 4288–4293
2025
-
[77]
Deep Learning for Computer Vision: A Brief Review,
A. V oulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis, “Deep Learning for Computer Vision: A Brief Review,”Computational intelligence and neuroscience, vol. 2018, no. 1, p. 7068349, 2018
2018
-
[78]
Segmentation-guided semantic-aware self-supervised denoising for sar image,
Y . Yuan, Y . Wu, P. Feng, Y . Fu, and Y . Wu, “Segmentation-guided semantic-aware self-supervised denoising for sar image,”IEEE Trans- actions on Geoscience and Remote Sensing, vol. 61, pp. 1–16, 2023
2023
-
[79]
Deep learning in physical layer: Review on data driven end-to-end communication systems and their enabling semantic applications,
N. Islam and S. Shin, “Deep learning in physical layer: Review on data driven end-to-end communication systems and their enabling semantic applications,”IEEE Open Journal of the Communications Society, 2024
2024
-
[80]
A Mathematical Theory of Semantic Commu- nication,
K. Niu and P. Zhang, “A Mathematical Theory of Semantic Commu- nication,”arXiv preprint arXiv:2401.13387, 2024
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.