Pith. sign in

REVIEW 4 major objections 6 minor 228 references

This survey claims that every semantic communication system for visual data can be classified by whether it preserves, expands, or refines transmitted semantics, and that this choice guides the entire system design.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 06:41 UTC pith:QVKPT7NT

load-bearing objection A broad, useful survey of SemCom-Vision whose central SPC/SEC/SRC taxonomy is intuitive but not tightly defined; the categories overlap under the paper's own information-theoretic gloss, and Ref [1] is mis-cited. the 4 major comments →

arxiv 2601.22202 v2 pith:QVKPT7NT submitted 2026-01-29 eess.IV cs.CV

A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications

classification eess.IV cs.CV
keywords semantic communicationvisual data transmissionsemantic quantizationsemantic preservation communicationsemantic expansion communicationsemantic refinement communicationknowledge graphjoint source-channel coding
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper is a survey trying to establish a single organizing principle for the scattered field of semantic communication for visual data (SemCom-Vision): every system can be viewed as preserving, expanding, or refining the semantics it transmits, depending on the communication goal. It claims these three categories—semantic preservation, semantic expansion, and semantic refinement—emerge naturally when semantics are quantified through information-theoretic, labeling-based, or knowledge-driven schemes. A sympathetic reader would care because the taxonomy gives researchers a way to choose encoder-decoder models, loss functions, training algorithms, and knowledge-graph structures according to what they want the transmission to achieve. If the classification holds, it turns a sprawling literature into a design guide and helps identify gaps, for instance in applying knowledge-driven quantization to vision.

Core claim

On its own terms, the survey's central claim is that SemCom-Vision approaches can be meaningfully categorized by communication goal as semantic preservation communication (SPC), semantic expansion communication (SEC), or semantic refinement communication (SRC). SPC aims to keep pixel- and semantic-level fidelity; SEC enriches received content with context or generative detail; SRC transmits only task-relevant semantics. The paper argues these goals are reflected in different semantic quantization schemes (entropy and mutual information, label embeddings, and knowledge graphs), different ML architectures and losses, and different knowledge structures, so the taxonomy doubles as a selection gu

What carries the argument

The load-bearing mechanism is the semantic quantization scheme—the way a system decides what counts as meaning—expressed three ways: information-theoretic (semantic entropy, mutual information, information bottleneck), labeling-based (vision-language embeddings, attention weights, classification probabilities), and knowledge-driven (nodes and edges of a knowledge graph). From these, the paper derives the SPC/SEC/SRC classification and uses it to organize encoder-decoder construction, loss functions, and knowledge-graph choices.

Load-bearing premise

The classification assumes every SemCom-Vision system can be assigned to exactly one of the three goals—preserve, expand, or refine semantics—but many real systems pursue more than one of these goals at once.

What would settle it

Take the systems surveyed here and ask independent reviewers to classify each as SPC, SEC, or SRC; the central claim fails if a substantial share of systems is placed in two categories or in none, since the paper provides no quantitative decision rule to resolve such ties.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Researchers can match encoder-decoder architectures to their communication goal: fidelity-oriented systems lean on convolutional and autoencoder designs, expansion-oriented systems on generative diffusion and GAN models, and refinement-oriented systems on transformers and attention-based filtering.
  • Loss functions align with the same goals: pixel and perceptual losses for preservation, denoising and adversarial losses for expansion, and information-bottleneck, sparsity, and task-specific losses for refinement.
  • Knowledge-graph structure choices become goal-dependent: triple stores suit refinement through logical reasoning, property graphs suit flexible semantic guidance, and hypergraphs suit complex multi-entity scenes.
  • Application domains such as digital twins, metaverse, and wireless perception can be mapped to the three categories, so designers can select SPC, SRC, or SEC per application need.
  • The survey's framework provides a shared vocabulary for comparing future SemCom-Vision systems and for spotting under-explored combinations, such as knowledge-driven quantization for visual tasks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The three categories may be better viewed as regions on a spectrum than as discrete bins: a refinement system preserves the subset of semantics it deems task-relevant, so future work could formalize the boundaries using measurable changes in semantic entropy or mutual information.
  • The same preserve-expand-refine logic likely extends beyond vision to multimodal and text-oriented semantic communication, where generative decoders similarly add or strip meaning.
  • A testable design rule follows from the paper: given a fixed bandwidth and task, the optimal category depends on the task's tolerance for semantic loss, which could be quantified with downstream task accuracy rather than pixel fidelity.
  • If the taxonomy were operationalized, it could feed into automated system selection—an agent could inspect the communication goal and choose architecture, loss, and knowledge structure without human intervention.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper surveys semantic communication for visual data (SemCom-Vision). Its main contribution is a taxonomy that classifies existing systems into semantic preservation communication (SPC), semantic expansion communication (SEC), and semantic refinement communication (SRC) based on communication goals viewed through semantic quantization schemes. It then reviews ML-based encoder-decoder architectures, training losses, knowledge graph structures, and applications. The survey is broad and well-structured, but the taxonomy's boundaries are not sharp enough to support the claimed 'clear categorization'.

Significance. If the taxonomy could be made operational, it would provide a useful organizing principle for a rapidly growing and fragmented literature. The survey's coverage of ML architectures, loss functions, and especially the knowledge graph section is valuable as a reference. The authors' effort to connect information-theoretic quantities (semantic entropy, mutual information) to practical design choices is interesting and timely. The paper is also honest about several open problems in the field. However, the central classification currently lacks a decision rule; the overlap between SPC and SRC, in particular, means that Table III assignments are not uniquely determined. This is a solvable problem but requires either a formal criterion or an explicit treatment of hybrid systems.

major comments (4)
  1. [Section III-B, Table III, Section IV-C] The SPC/SEC/SRC categories are defined by qualitative goals, but the information-theoretic gloss in Section III-B does not make them exclusive. SRC 'deliberately reduces H(s)' and maximizes I(s; s_hat_T); any SRC system that faithfully transmits the task-relevant subset also preserves that subset's semantics and thus satisfies the SPC criterion on that subset. Table III assigns each reference to one category, and Section IV uses the categories to recommend architectures and losses (e.g., L9 in Eq. (11) reconstructs only Omega, yet the selected region is preserved). Without a formal decision rule or an explicit treatment of hybrid systems, the same paper can be placed in multiple buckets, so the claimed 'clear categorization' and the model-selection guidance are not uniquely determined.
  2. [Section III-A, Eqs. (1)-(2), Section III-B] The semantic entropy H(s) and mutual information I(s; s_hat) are used to differentiate SPC, SEC, and SRC, but the paper does not specify how p(s_k) or s_hat are obtained for a given system. SEC is said to 'intentionally increase the entropy beyond that of the source'; no criterion is given to decide when an added semantic is new versus a reconstruction artifact. Without operational definitions, the proposed 'semantic quantization schemes' cannot be used as a decision procedure to classify existing systems or to derive testable predictions. Please provide at least a minimal protocol (e.g., what random variables are defined, how the task target T is formalized) or restrict the claims to a qualitative taxonomy.
  3. [Table III, Section IV] The survey assigns dozens of references to SPC/SEC/SRC in Table III but provides no annotation protocol, no decision tree, and no inter-annotator reliability check. For example, [144] is placed in SEC because it uses a GAN to generate artistically styled images from segmentation maps, while [131] is placed in SPC despite also using a GAN and a discriminator; the boundary rests on a qualitative reading of the system's goal. Since the taxonomy is the paper's main contribution, the authors should explain how a reader can independently reproduce the assignments, or alternatively present the categories as overlapping design dimensions rather than a partition.
  4. [Section IV-A, Eq. (6)] The VAE training loss L4 is written as reconstruction error minus the KL-divergence term, i.e., L4 = MSE - (1/2) Sum(1 + log sigma^2 - mu^2 - sigma^2). The standard negative ELBO has a plus sign before the KL divergence; the minus sign here drives the encoder to maximize variance and is not the variational loss used in the cited works. Please correct the sign or clarify that this is a different objective.
minor comments (6)
  1. [Sections IV-A, IV-B, IV-C] The text repeatedly refers to 'Table IV' for the overview of SemCom encoder-decoder models, but the relevant table is Table III. This cross-reference error will confuse readers.
  2. [Section II-A] Typo: 'ReLu' should be 'ReLU'.
  3. [Figure 1 caption] Typo: 'Potiential Applications' should be 'Potential Applications'. Also the caption lists both 'V. Potiential Applications' and 'VI. Potential Applications', which appear to be duplicate labels.
  4. [Section V-A] Typo: 'TIV A-KG' should be 'TIVA-KG'.
  5. [Section III-A] Typo: 'UA V-photoed' should be 'UAV-photoed' (or with a space removed).
  6. [Eq. (9)] The expectation subscript is malformed: 'E s,epsilon~N(0,1)' should be written as E_{s, epsilon~N(0,1)}.

Circularity Check

0 steps flagged

No significant circularity: the SPC/SEC/SRC taxonomy is an interpretive classification, not a derivation that reduces to its inputs.

full rationale

The paper's central contribution is a survey taxonomy. SPC/SEC/SRC are introduced as definitions in Section III-B ('Semantic Preservation Communication (SPC): Transmit extracted semantics while mitigating channel-induced semantic loss...'; 'Semantic Expansion Communication (SEC): Extract semantics... enrich transmitted semantics...'; 'Semantic Refinement Communication (SRC): Extract and transmit only task-relevant semantics...'), and the rest of the paper maps existing systems onto these categories. This is an act of interpretation, not a derivation: no quantity is fitted and then reported as a prediction, and no mathematical claim is reduced by construction to its own assumptions. The information-theoretic gloss in Section III-B (SPC maximizes H(s)/I(s;s-hat), SRC applies IB to reduce H(s)) is a post-hoc characterization of the categories, not an input used to generate them. The self-citations used as examples ([20], [22], [25], [66], [84]) are illustrative and not load-bearing; the taxonomy does not depend on any uniqueness theorem or unverified prior claim by the authors. The objection that the categories lack a formal decision rule and may overlap (e.g., an SRC system that faithfully transmits the task-relevant subset also preserves mutual information over that subset) is a validity/usability limitation of the classification, but it is not circularity. Under the required standard of exhibiting a specific reduction, no circular step is present.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

The survey relies on a qualitative taxonomy rather than quantitative derivation. The main assumptions are the independence of semantics in the entropy definition and the mutual exclusivity/exhaustiveness of the three proposed categories. No free parameters or invented physical entities are involved.

axioms (3)
  • domain assumption Semantics are independent: p(s) = ∏ p(s_k) (Eq. 1).
    Used to define semantic entropy; independence is rarely true for visual scene semantics (objects are correlated), but the paper uses it without justification.
  • ad hoc to paper All SemCom-Vision systems can be partitioned into SPC, SEC, SRC (Section III-B).
    The taxonomy is introduced as a 'novel classification perspective' but no evidence of exhaustiveness or exclusivity is given.
  • domain assumption Semantic entropy is 'the fundamental measure of semantics' (Section III-A).
    Unproved; alternative measures (label similarity, graph distance) are then layered on, so the claim of fundamentality is not defended.

pith-pipeline@v1.3.0-alltime-deepseek · 40020 in / 9177 out tokens · 91835 ms · 2026-08-03T06:41:56.501983+00:00 · methodology

0 comments
read the original abstract

Semantic communication (SemCom) emerges as a transformative paradigm for traffic-intensive visual data transmission, shifting focus from raw data to meaningful content transmission and relieving the increasing pressure on communication resources. However, to achieve SemCom, challenges are faced in accurate semantic quantization for visual data, robust semantic extraction and reconstruction under diverse tasks and goals, transceiver coordination with effective knowledge utilization, and adaptation to unpredictable wireless communication environments. In this paper, we present a systematic review of SemCom for visual data transmission (SemCom-Vision), wherein an interdisciplinary analysis integrating computer vision (CV) and communication engineering is conducted to provide comprehensive guidelines for the machine learning (ML)-empowered SemCom-Vision design. Specifically, this survey first elucidates the basics and key concepts of SemCom. Then, we introduce a novel classification perspective to categorize existing SemCom-Vision approaches as semantic preservation communication (SPC), semantic expansion communication (SEC), and semantic refinement communication (SRC) based on communication goals interpreted through semantic quantization schemes. Moreover, this survey articulates the ML-based encoder-decoder models and training algorithms for each SemCom-Vision category, followed by knowledge structure and utilization strategies. Finally, we discuss potential SemCom-Vision applications.

Figures

Figures reproduced from arXiv: 2601.22202 by Ahmad Taha, David Flynn, Muhammad Ali Imran, Runze Cheng, Xuesong Liu, Yao Sun.

Figure 1
Figure 1. Figure 1: The construction of SemCom-Vision includes three major parts: 1) semantic quantization and SemCom classification, [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: In this section, we separately elaborate on the ML [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 2
Figure 2. Figure 2: Typical ML model architectures for encoder-decoder design. [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 2
Figure 2. Figure 2: In the SEC framework, the encoder extracts latent [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

228 extracted references · 35 linked inside Pith

  1. [1]

    The Synthesis Between Artificial Intelligence and Editing Stories of the Future,

    S. Erdem, “The Synthesis Between Artificial Intelligence and Editing Stories of the Future,” inTransforming cinema with artificial intelli- gence. IGI Global Scientific Publishing, 2025, pp. 221–240. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 20

  2. [2]

    Transmit what you need: task-adaptive semantic communications for visual information,

    J. Park and S. W. Yoon, “Transmit what you need: task-adaptive semantic communications for visual information,”arXiv preprint arXiv:2412.13646, 2024

  3. [3]

    Vpr-bench: An open-source visual place recogni- tion evaluation framework with quantifiable viewpoint and appearance change,

    M. Zaffar, S. Garg, M. Milford, J. Kooij, D. Flynn, K. McDonald- Maier, and S. Ehsan, “Vpr-bench: An open-source visual place recogni- tion evaluation framework with quantifiable viewpoint and appearance change,”International Journal of Computer Vision, vol. 129, no. 7, pp. 2136–2174, 2021

  4. [4]

    A Survey on Semantic Communications for Intelligent Wireless Networks,

    S. Iyer, R. Khanai, D. Torse, R. J. Pandya, K. M. Rabie, K. Pai, W. U. Khan, and Z. Fadlullah, “A Survey on Semantic Communications for Intelligent Wireless Networks,”Wireless Personal Communications, vol. 129, no. 1, pp. 569–611, 2023

  5. [5]

    Semantic Communications for Future Internet: Fundamentals, Applications, and Challenges,

    W. Yang, H. Du, Z. Q. Liew, W. Y . B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Shen, and C. Miao, “Semantic Communications for Future Internet: Fundamentals, Applications, and Challenges,”IEEE Communications Surveys & Tutorials, vol. 25, no. 1, pp. 213–250, 2022

  6. [6]

    Semantics-Empowered Communications: A Tutorial-Cum-Survey,

    Z. Lu, R. Li, K. Lu, X. Chen, E. Hossain, Z. Zhao, and H. Zhang, “Semantics-Empowered Communications: A Tutorial-Cum-Survey,” IEEE Communications Surveys & Tutorials, vol. 26, no. 1, pp. 41– 79, 2023

  7. [7]

    Less Data, More Knowledge: Building Next Generation Semantic Communication Networks,

    C. Chaccour, W. Saad, M. Debbah, Z. Han, and H. V . Poor, “Less Data, More Knowledge: Building Next Generation Semantic Communication Networks,”IEEE Communications Surveys & Tutorials, 2024

  8. [8]

    Intellicise Wireless Networks from Semantic Communications: A Survey, Research Issues, and Challenges,

    P. Zhang, W. Xu, Y . Liu, X. Qin, K. Niu, S. Cui, G. Shi, Z. Qin, X. Xu, F. Wanget al., “Intellicise Wireless Networks from Semantic Communications: A Survey, Research Issues, and Challenges,”IEEE Communications Surveys & Tutorials, 2024

  9. [9]

    A Survey on Goal-Oriented Semantic Communication: Techniques, Challenges, and Future Direc- tions,

    T. M. Getu, G. Kaddoum, and M. Bennis, “A Survey on Goal-Oriented Semantic Communication: Techniques, Challenges, and Future Direc- tions,”IEEE Access, 2024

  10. [10]

    Semantic Communication: A Survey on Research Landscape, Challenges, and Future Directions,

    ——, “Semantic Communication: A Survey on Research Landscape, Challenges, and Future Directions,”Proceedings of the IEEE, 2025

  11. [11]

    A Review of Convolutional Neural Networks in Computer Vision,

    X. Zhao, L. Wang, Y . Zhang, X. Han, M. Deveci, and M. Parmar, “A Review of Convolutional Neural Networks in Computer Vision,” Artificial Intelligence Review, vol. 57, no. 4, p. 99, 2024

  12. [12]

    Visual Analytics for Machine Learn- ing: A Data Perspective Survey,

    J. Wang, S. Liu, and W. Zhang, “Visual Analytics for Machine Learn- ing: A Data Perspective Survey,”IEEE transactions on visualization and computer graphics, vol. 30, no. 12, pp. 7637–7656, 2024

  13. [13]

    A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and Application,

    B. Yuan and D. Zhao, “A Survey on Continual Semantic Segmentation: Theory, Challenge, Method and Application,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  14. [14]

    Image Generative Semantic Communication With Multi- Modal Similarity Estimation for Resource-Limited Networks,

    E. Hosonuma, T. Yamazaki, T. Miyoshi, A. Taya, Y . Nishiyama, and K. Sezaki, “Image Generative Semantic Communication With Multi- Modal Similarity Estimation for Resource-Limited Networks,”arXiv preprint arXiv:2404.11280, 2024

  15. [15]

    Semantic Density: Uncertainty Quantifi- cation for Large Language Models Through Confidence Measurement in Semantic Space,

    X. Qiu and R. Miikkulainen, “Semantic Density: Uncertainty Quantifi- cation for Large Language Models Through Confidence Measurement in Semantic Space,”arXiv preprint arXiv:2405.13845, 2024

  16. [16]

    Feature Importance-Aware Task-Oriented Semantic Transmission and Optimization,

    Y . Wang, S. Han, X. Xu, H. Liang, R. Meng, C. Dong, and P. Zhang, “Feature Importance-Aware Task-Oriented Semantic Transmission and Optimization,”IEEE Transactions on Cognitive Communications and Networking, 2024

  17. [17]

    Task-Oriented Explainable Semantic Communications,

    S. Ma, W. Qiao, Y . Wu, H. Li, G. Shi, D. Gao, Y . Shi, S. Li, and N. Al- Dhahir, “Task-Oriented Explainable Semantic Communications,”IEEE transactions on wireless communications, vol. 22, no. 12, pp. 9248– 9262, 2023

  18. [18]

    Vector Quantized Semantic Communication System,

    Q. Fu, H. Xie, Z. Qin, G. Slabaugh, and X. Tao, “Vector Quantized Semantic Communication System,”IEEE Wireless Communications Letters, vol. 12, no. 6, pp. 982–986, 2023

  19. [19]

    Federated Semantic Learning Driven by Information Bottleneck for Task-Oriented Communications,

    H. Wei, W. Ni, W. Xu, F. Wang, D. Niyato, and P. Zhang, “Federated Semantic Learning Driven by Information Bottleneck for Task-Oriented Communications,”IEEE Communications Letters, vol. 27, no. 10, pp. 2652–2656, 2023

  20. [20]

    A Semantic Communication-Based Workload-Adjustable Transceiver for Wireless Ai-Generated Content (AIGC) Delivery,

    R. Cheng, Y . Sun, L. Zhang, L. Feng, L. Zhang, and M. A. Imran, “A Semantic Communication-Based Workload-Adjustable Transceiver for Wireless Ai-Generated Content (AIGC) Delivery,”arXiv preprint arXiv:2503.18874, 2025

  21. [21]

    Background Knowledge Aware Semantic Coding Model Selection,

    F. Zhao, Y . Sun, R. Cheng, and M. A. Imran, “Background Knowledge Aware Semantic Coding Model Selection,” in2022 IEEE 22nd Inter- national Conference on Communication Technology (ICCT). IEEE, 2022, pp. 84–89

  22. [22]

    Vista: Video Transmission Over a Semantic Communication Approach,

    C. Liang, X. Deng, Y . Sun, R. Cheng, L. Xia, D. Niyato, and M. A. Imran, “Vista: Video Transmission Over a Semantic Communication Approach,” in2023 IEEE International Conference on Communica- tions Workshops (ICC Workshops). IEEE, 2023, pp. 1777–1782

  23. [23]

    VISTA: A Semantic Communication Approach for Video Transmission,

    C. Liang, X. Deng, Y . Sun, R. Cheng, L. Xia, D. Niyato, and M. Ali Imran, “VISTA: A Semantic Communication Approach for Video Transmission,”Wireless Semantic Communications: Concepts, Principles and Challenges, pp. 109–121, 2025

  24. [24]

    Deep learning enabled semantic communication systems,

    H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,”IEEE transactions on signal pro- cessing, vol. 69, pp. 2663–2675, 2021

  25. [25]

    A wireless ai-generated content (aigc) provisioning framework em- powered by semantic communication,

    R. Cheng, Y . Sun, D. Niyato, L. Zhang, L. Zhang, and M. Imran, “A wireless ai-generated content (aigc) provisioning framework em- powered by semantic communication,”IEEE Transactions on Mobile Computing, 2024

  26. [26]

    When Wireless Communications Meet Computer Vision in Beyond 5G,

    T. Nishio, Y . Koda, J. Park, M. Bennis, and K. Doppler, “When Wireless Communications Meet Computer Vision in Beyond 5G,”IEEE Communications Standards Magazine, vol. 5, no. 2, pp. 76–83, 2021

  27. [27]

    Applying Deep-Learning-Based Computer Vision to Wireless Communications: Methodologies, Oppor- tunities, and Challenges,

    Y . Tian, G. Pan, and M.-S. Alouini, “Applying Deep-Learning-Based Computer Vision to Wireless Communications: Methodologies, Oppor- tunities, and Challenges,”IEEE Open Journal of the Communications Society, vol. 2, pp. 132–143, 2020

  28. [28]

    E. R. Davies,Computer Vision: Principles, Algorithms, Applications, Learning. Academic Press, 2017

  29. [29]

    A Review on Machine Learning Styles in Computer Vision—Techniques and Future Directions,

    S. V . Mahadevkar, B. Khemani, S. Patil, K. Kotecha, D. R. V ora, A. Abraham, and L. A. Gabralla, “A Review on Machine Learning Styles in Computer Vision—Techniques and Future Directions,”Ieee Access, vol. 10, pp. 107 293–107 329, 2022

  30. [30]

    Deep Learning vs. Traditional Computer Vision,

    N. O’Mahony, S. Campbell, A. Carvalho, S. Harapanahalli, G. V . Hernandez, L. Krpalkova, D. Riordan, and J. Walsh, “Deep Learning vs. Traditional Computer Vision,” inAdvances in computer vision: proceedings of the 2019 computer vision conference (CVC), volume 1 1. Springer, 2020, pp. 128–144

  31. [31]

    S- ran: Semantic-aware radio access networks,

    Y . Sun, L. Zhang, L. Guo, J. Li, D. Niyato, and Y . Fang, “S- ran: Semantic-aware radio access networks,”IEEE Communications Magazine, 2024

  32. [32]

    Machine Learning Enabled Heterogeneous Semantic and Bit Communication,

    M. Zhang, R. Zhong, X. Mu, and Y . Liu, “Machine Learning Enabled Heterogeneous Semantic and Bit Communication,”IEEE Transactions on Wireless Communications, 2024

  33. [33]

    Semantic Entropy Can Simultaneously Benefit Transmission Efficiency and Channel Security of Wireless Se- mantic Communications,

    Y . Rong, G. Nan, M. Zhang, S. Chen, S. Wang, X. Zhang, N. Ma, S. Gong, Z. Yang, Q. Cuiet al., “Semantic Entropy Can Simultaneously Benefit Transmission Efficiency and Channel Security of Wireless Se- mantic Communications,”IEEE Transactions on Information Forensics and Security, 2025

  34. [34]

    Niu and P

    K. Niu and P. Zhang,The Mathematical Theory of Semantic Commu- nication. Springer Nature, 2025

  35. [35]

    Advanced Semantics for Commonsense Knowledge Extraction,

    T.-P. Nguyen, S. Razniewski, and G. Weikum, “Advanced Semantics for Commonsense Knowledge Extraction,” inProceedings of the Web Conference 2021, 2021, pp. 2636–2647

  36. [36]

    A Continual Learning Survey: Defy- ing Forgetting in Classification Tasks,

    M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars, “A Continual Learning Survey: Defy- ing Forgetting in Classification Tasks,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 7, pp. 3366–3385, 2021

  37. [37]

    Resource Management, Security, and Privacy Issues in Semantic Communications: A Survey,

    D. Won, G. Woraphonbenjakul, A. B. Wondmagegn, A. T. Tran, D. Lee, D. S. Lakew, and S. Cho, “Resource Management, Security, and Privacy Issues in Semantic Communications: A Survey,”IEEE Communications Surveys & Tutorials, 2024

  38. [38]

    Toward Intelli- gent Resource Allocation on Task-Oriented Semantic Communication,

    H. Zhang, H. Wang, Y . Li, K. Long, and V . C. Leung, “Toward Intelli- gent Resource Allocation on Task-Oriented Semantic Communication,” IEEE Wireless Communications, vol. 30, no. 3, pp. 70–77, 2023

  39. [39]

    A Survey on Semantic Communication Networks: Architecture, Security, and Privacy,

    S. Guo, Y . Wang, N. Zhang, Z. Su, T. H. Luan, Z. Tian, and X. Shen, “A Survey on Semantic Communication Networks: Architecture, Security, and Privacy,”IEEE Communications Surveys & Tutorials, 2024

  40. [40]

    What Is Semantic Communication? A View on Conveying Meaning in the Era of Machine Intelligence,

    Q. Lan, D. Wen, Z. Zhang, Q. Zeng, X. Chen, P. Popovski, and K. Huang, “What Is Semantic Communication? A View on Conveying Meaning in the Era of Machine Intelligence,”Journal of Communica- tions and Information Networks, vol. 6, no. 4, pp. 336–371, 2021

  41. [41]

    A Survey of Con- volutional Neural Networks: Analysis, Applications, and Prospects,

    Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A Survey of Con- volutional Neural Networks: Analysis, Applications, and Prospects,” Ieee Transactions on Neural Networks and Learning Systems, vol. 33, no. 12, pp. 6999–7019, 2021

  42. [42]

    An Equivalence of Fully Connected Layer and Convolutional Layer,

    W. Ma and J. Lu, “An Equivalence of Fully Connected Layer and Convolutional Layer,”arXiv preprint arXiv:1712.01252, 2017

  43. [43]

    Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification,

    S. S. Basha, S. R. Dubey, V . Pulabaigari, and S. Mukherjee, “Impact of Fully Connected Layers on Performance of Convolutional Neural Networks for Image Classification,”Neurocomputing, vol. 378, pp. 112–119, 2020

  44. [44]

    The Unreasonable Effectiveness of Fully-Connected Layers for Low- Data Regimes,

    P. Kocsis, P. S ´uken´ık, G. Bras´o, M. Nießner, L. Leal-Taix´e, and I. Elezi, “The Unreasonable Effectiveness of Fully-Connected Layers for Low- Data Regimes,”Advances in Neural Information Processing Systems, vol. 35, pp. 1896–1908, 2022. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 21

  45. [45]

    Fully convolutional networks for semantic segmentation,

    J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3431–3440

  46. [46]

    An Introduction to Convolutional Neural Networks,

    K. O’shea and R. Nash, “An Introduction to Convolutional Neural Networks,”arXiv preprint arXiv:1511.08458, 2015

  47. [47]

    Understanding Convolution for Semantic Segmentation,

    P. Wang, P. Chen, Y . Yuan, D. LIU, Z. Huang, X. Hou, and G. Cottrell, “Understanding Convolution for Semantic Segmentation,” in2018 Ieee Winter Conference on Applications of Computer Vision (Wacv). Ieee, 2018, pp. 1451–1460

  48. [48]

    Activation Functions: Comparison of Trends in Practice and Research for Deep Learning,

    C. Nwankpa, W. Ijomah, A. Gachagan, and S. Marshall, “Activation Functions: Comparison of Trends in Practice and Research for Deep Learning,”arXiv preprint arXiv:1811.03378, 2018

  49. [49]

    FinBack: Infiltrating Backdoors into Gradient Compressors on Federated Learning,

    X. Tang, W. Yang, L. Peng, M. Shen, T. Zhang, Y . Weng, J. Kang, and D. Niyato, “FinBack: Infiltrating Backdoors into Gradient Compressors on Federated Learning,”IEEE Transactions on Information Forensics and Security, 2025

  50. [50]

    ROBY: A Byzantine-Robust and Privacy-Preserving Serverless Federated Learning Framework,

    X. Tang, M. Li, M. Shen, J. Kang, L. Zhu, Z. Liu, G. Yang, D. Niyato, and R. H. Deng, “ROBY: A Byzantine-Robust and Privacy-Preserving Serverless Federated Learning Framework,”IEEE Transactions on Information Forensics and Security, 2025

  51. [51]

    Pile: Robust Privacy-Preserving Federated Learning via Verifiable Perturbations,

    X. Tang, M. Shen, Q. Li, L. Zhu, T. Xue, and Q. Qu, “Pile: Robust Privacy-Preserving Federated Learning via Verifiable Perturbations,” IEEE Transactions on Dependable and Secure Computing, vol. 20, no. 6, pp. 5005–5023, 2023

  52. [52]

    Nonparametric Regression Using Deep Neu- ral Networks With Relu Activation Function,

    J. SCHMIDT-HIEBER, “Nonparametric Regression Using Deep Neu- ral Networks With Relu Activation Function,”The Annals of Statistics, vol. 48, no. 4, pp. 1875–1897, 2020

  53. [53]

    Why TanH Is a Hardware Friendly Activation Function for CNNs,

    K. Abdelouahab, M. Pelcat, and F. Berry, “Why TanH Is a Hardware Friendly Activation Function for CNNs,” inProceedings of the 11th international conference on distributed smart cameras, 2017, pp. 199– 201

  54. [54]

    The Most Used Activation Functions: Classic Versus Current,

    M. A. Mercioni and S. Holban, “The Most Used Activation Functions: Classic Versus Current,” in2020 International Conference on Devel- opment and Application Systems (DAS). IEEE, 2020, pp. 141–145

  55. [55]

    Mixed Pooling for Convolu- tional Neural Networks,

    D. Yu, H. Wang, P. Chen, and Z. Wei, “Mixed Pooling for Convolu- tional Neural Networks,” inRough Sets and Knowledge Technology: 9th International Conference, RSKT 2014, Shanghai, China, October 24-26, 2014, Proceedings 9. Springer, 2014, pp. 364–375

  56. [56]

    The Treasure Beneath Convolutional Layers: Cross-Convolutional-Layer Pooling for Image Classification,

    L. Liu, C. Shen, and A. Van den Hengel, “The Treasure Beneath Convolutional Layers: Cross-Convolutional-Layer Pooling for Image Classification,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 4749–4757

  57. [57]

    Normalization Techniques in Training Dnns: Methodology, Analysis and Applica- tion,

    L. Huang, J. Qin, Y . Zhou, F. Zhu, L. Liu, and L. Shao, “Normalization Techniques in Training Dnns: Methodology, Analysis and Applica- tion,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 8, pp. 10 173–10 196, 2023

  58. [58]

    Layer Normalization,

    J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer Normalization,”arXiv preprint arXiv:1607.06450, 2016

  59. [59]

    A Review on the Attention Mechanism of Deep Learning,

    Z. Niu, G. Zhong, and H. Yu, “A Review on the Attention Mechanism of Deep Learning,”Neurocomputing, vol. 452, pp. 48–62, 2021

  60. [60]

    Attention Is All You Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” Advances in neural information processing systems, vol. 30, 2017

  61. [61]

    An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gellyet al., “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,”arXiv preprint arXiv:2010.11929, 2020

  62. [62]

    Resnet in Resnet: Generalizing Residual Architectures,

    S. Targ, D. Almeida, and K. Lyman, “Resnet in Resnet: Generalizing Residual Architectures,”arXiv preprint arXiv:1603.08029, 2016

  63. [63]

    Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks With Symmetric Skip Connections,

    X. Mao, C. Shen, and Y .-B. Yang, “Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks With Symmetric Skip Connections,”Advances in neural information processing systems, vol. 29, 2016

  64. [64]

    U-Net: Convolutional Networks for Biomedical Image Segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” inMedical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, pro- ceedings, part III 18. Springer, 2015, pp. 234–241

  65. [65]

    Semantic Communi- cations: Principles and Challenges,

    Z. Qin, X. Tao, J. Lu, W. Tong, and G. Y . Li, “Semantic Communi- cations: Principles and Challenges,”arXiv preprint arXiv:2201.01389, 2021

  66. [66]

    Knowledge-Assisted Privacy Preserving in Semantic Communication,

    X. Liu, Y . Sun, R. Cheng, L. Xia, H. Abumarshoud, L. Zhang, and M. A. Imran, “Knowledge-Assisted Privacy Preserving in Semantic Communication,”IEEE Wireless Communications, vol. 32, no. 2, pp. 76–83, 2025

  67. [67]

    Semantic Communications with Computer Vision Sensing for Edge Video Trans- mission,

    Y . Peng, L. Xiang, K. Yang, K. Wang, and M. Debbah, “Semantic Communications with Computer Vision Sensing for Edge Video Trans- mission,”arXiv preprint arXiv:2503.07252, 2025

  68. [68]

    Lightweight Vision Model-based Multi-user Semantic Communication Systems,

    F. Jiang, S. Tu, L. Dong, K. Wang, K. Yang, R. Liu, C. Pan, and J. Wang, “Lightweight Vision Model-based Multi-user Semantic Communication Systems,”arXiv preprint arXiv:2502.16424, 2025

  69. [69]

    Personalized Fed- erated Learning for GAI-Assisted Semantic Communications,

    Y . Peng, F. Jiang, L. Dong, K. Wang, and K. Yang, “Personalized Fed- erated Learning for GAI-Assisted Semantic Communications,”IEEE Transactions on Cognitive Communications and Networking, 2025

  70. [70]

    A Lite Distributed Semantic Communication System for Internet of Things,

    H. Xie and Z. Qin, “A Lite Distributed Semantic Communication System for Internet of Things,”IEEE Journal on Selected Areas in Communications, vol. 39, no. 1, pp. 142–153, 2020

  71. [71]

    Deep Joint Source-Channel Coding for Semantic Communications,

    J. Xu, T.-Y . Tung, B. Ai, W. Chen, Y . Sun, and D. G ¨und¨uz, “Deep Joint Source-Channel Coding for Semantic Communications,”IEEE communications Magazine, vol. 61, no. 11, pp. 42–48, 2023

  72. [72]

    Innovative Semantic Communication System,

    C. Dong, H. Liang, X. Xu, S. Han, B. Wang, and P. Zhang, “Innovative Semantic Communication System,”arXiv preprint arXiv:2202.09595, 2022

  73. [73]

    Joint Source-Channel Cod- ing for Channel-Adaptive Digital Semantic Communications,

    J. Park, Y . Oh, S. Kim, and Y .-S. Jeon, “Joint Source-Channel Cod- ing for Channel-Adaptive Digital Semantic Communications,”IEEE Transactions on Cognitive Communications and Networking, 2024

  74. [74]

    Channel Coding Towards 6G: Technical Overview and Outlook,

    M. Rowshan, M. Qiu, Y . Xie, X. Gu, and J. Yuan, “Channel Coding Towards 6G: Technical Overview and Outlook,”IEEE Open Journal of the Communications Society, 2024

  75. [75]

    Improved Nonlinear Transform Source-Channel Coding to Catalyze Semantic Communications,

    S. Wang, J. Dai, X. Qin, Z. Si, K. Niu, and P. Zhang, “Improved Nonlinear Transform Source-Channel Coding to Catalyze Semantic Communications,”IEEE Journal of Selected Topics in Signal Process- ing, vol. 17, no. 5, pp. 1022–1037, 2023

  76. [76]

    Power Allocation for Throughput Maximization in NOMA-Based Semantic Communication System,

    K. Ma, H. Abumarshoud, S. Hua, M. Imran, and Y . Sun, “Power Allocation for Throughput Maximization in NOMA-Based Semantic Communication System,” inICC 2025-IEEE International Conference on Communications. IEEE, 2025, pp. 4288–4293

  77. [77]

    Deep Learning for Computer Vision: A Brief Review,

    A. V oulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis, “Deep Learning for Computer Vision: A Brief Review,”Computational intelligence and neuroscience, vol. 2018, no. 1, p. 7068349, 2018

  78. [78]

    Segmentation-guided semantic-aware self-supervised denoising for sar image,

    Y . Yuan, Y . Wu, P. Feng, Y . Fu, and Y . Wu, “Segmentation-guided semantic-aware self-supervised denoising for sar image,”IEEE Trans- actions on Geoscience and Remote Sensing, vol. 61, pp. 1–16, 2023

  79. [79]

    Deep learning in physical layer: Review on data driven end-to-end communication systems and their enabling semantic applications,

    N. Islam and S. Shin, “Deep learning in physical layer: Review on data driven end-to-end communication systems and their enabling semantic applications,”IEEE Open Journal of the Communications Society, 2024

  80. [80]

    A Mathematical Theory of Semantic Commu- nication,

    K. Niu and P. Zhang, “A Mathematical Theory of Semantic Commu- nication,”arXiv preprint arXiv:2401.13387, 2024

Showing first 80 references.