REVIEW 2 major objections 5 minor 187 references
Compact Visual Data Representation for Green Multimedia -- A Human Visual System Perspective
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A survey argues that the human visual system's roughly 100,000-fold compression, against VVC's roughly 1,000-fold, should push video coding toward compact, task-ready representations rather than pixel-perfect reconstruction.
desk verdict A solid survey with a shaky motivating statistic: the HVS-vs-VVC compression ratio is undefined and the 'impossible gap' claim goes beyond the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing machinery of the paper is the HVS-inspired tripartite framework that maps biological mechanisms onto coding architectures: perceptual coding based on the eye's nonuniform sensitivity and memory, compact feature representation based on task-driven selective attention, and collaborative or scalable coding based on hierarchical cortical processing. The most concrete anchor is the 'Digital Retina' architecture, which uses a dual representation of compact textures and compact semantic features, allowing receivers to use feature streams directly for analytics. These principles are used to classify the surveyed methods and to structure the paper's roadmap for future green multimedia technologies.
What would settle it
Measure the bitrate required to achieve a fixed level of a concrete task—for example, object detection on a standard benchmark at a target mean average precision—using a compact feature codec on one side and a VVC-reconstructed pixel codec on the other. If the resulting compression ratio relative to raw video is on the order of 1,000 rather than 100,000, the paper's motivating disparity is not supported by direct measurement.
Extended reading notes
Core claim
The paper's central claim is that the purpose of visual data representation should shift from signal reconstruction to knowledge extraction, motivated by the enormous efficiency of the human visual system. According to the authors, the HVS compresses visual information around 100,000 times while achieving high generalization and energy efficiency, whereas VVC achieves a compression ratio of only about 1,000 times for raw visual data. This disparity, they argue, makes it impossible to close the gap by improving codec efficiency alone, because the objectives differ: codecs reconstruct, brains understand. The survey accordingly maps a landscape of techniques that represent visual data compactly for machines—compact feature coding, end-to-end learned coding, generative and external-data-based compression, and scalable layered bitstreams that blend feature and texture coding—and positions these as the route to greener multimedia.
Load-bearing premise
The entire argument rests on the meaning of the claim that the human visual system compresses visual information about 100,000 times; if that number is not defined in terms of a comparable measure (such as bits per useful concept or task-accuracy-preserving bitrate), the gap between HVS and VVC loses its quantitative force.
Editorial extensions
If this is right
- If the survey's thesis is correct, future video coding standards should prioritize machine-consumable feature streams over pixel fidelity, potentially making Video Coding for Machines a mainstream direction.
- Layered or scalable bitstreams that carry a base feature layer for analytics and an enhancement layer for human viewing would become the default architecture, reducing bandwidth and decode energy for task-driven applications.
- End-to-end learned codecs and generative compression, which already show large bitrate savings at low rates, would be recognized not as niche tools but as core green multimedia technologies.
- Perceptual coding and just-noticeable-difference modeling would be applied more aggressively to save energy, not only to improve subjective quality.
- The connection to knowledge-centric networking and edge computing suggests that compact representation will be designed jointly with network and model updates, yielding systems that transmit only the information needed for a task.
Reading between the lines
- A testable extension of the survey's thesis: define compression ratio for a biological system by the bitrate needed to preserve task accuracy (for example, object detection at fixed mean average precision) and compare a compact feature codec against VVC-reconstructed pixels; the observed gap may be far smaller than 100,000x, which would weaken the quantitative foundation of the green-multimedia ar
- The survey implies but does not state that the comparison between HVS and VVC is not apples-to-apples: one is a task-accuracy measure and the other a signal-fidelity measure, so the 100,000x figure may conflate 'useful concepts' with 'bits'.
- A natural next step, beyond the paper, is to extend the HVS-inspired framework to AI-agent communication, where semantic communication among machines could achieve even higher compactness than human-oriented coding.
- The paper's own roadmap suggests that large vision-language models could become the ultimate consumers of compact visual data, which would make the design of bitstream syntax for semantic prompts a future standardization problem.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys compact visual data representation for green multimedia, framed by the claim that the human visual system (HVS) compresses visual information by roughly 100,000 times while the VVC video coding standard achieves about 1,000 times. The survey is organized into three areas: compact video compression (standards, end-to-end coding, perceptual coding, external-data/generative compression), compact feature compression (handcrafted and deep descriptors, CDVS/CDVA/VCM standardization), and unified representation for dynamic tasks (layered feature/texture coding, scalable coding). It closes with connections to AIGC, large vision-language models, knowledge-centric networking, edge computing, and future directions such as AI-agent communication and neuromorphic computing. The intended contribution is a research roadmap for HVS-inspired, task-oriented, energy-efficient visual representation.
Significance. If its motivating comparison were rigorously defined, this would be a timely and useful synthesis of an emerging area. The survey is broad and current, covering recent codec standards, learned compression, VCM, and joint feature/texture coding, and it explicitly connects these developments to sustainability. The taxonomy and the roadmap in Fig. 2 are reasonable organizational contributions. The paper does not present new experiments or derivations, and its value rests on the adequacy of its selected evidence and framing. The central quantitative motivation, however, is not operationalized, and several forward-looking claims in Section IV are asserted without support, which currently limits the paper's ability to justify the proposed research program over an alternative, qualitative framing.
major comments (2)
- [Section I and Section IV (also Abstract, Fig. 1)] The claim that the HVS compresses visual information 'around 100,000 times' while VVC achieves 'around 1,000 times' is cited to [15] without defining the compression ratio on either side. For a biological system, the numerator and denominator are unspecified (photoreceptor or retinal output bits versus conscious percept? bits of task-relevant content versus raw input?), and for VVC no bitrate or quality operating point is given. These are not commensurable quantities, so the 'notable disparity' and the green-multimedia research program built on it rest on an undefined benchmark. Please either provide an explicit bit model and operating points for both sides, or reframe the HVS comparison as qualitative biological inspiration rather than a quantitative compression gap.
- [Section IV, first paragraph] The statement that the reconstruction objective and the understanding objective 'ultimately make it impossible to bridge the performance gap' goes beyond what the preceding numbers establish. Different objectives do not by themselves imply an unbridgeable rate gap; many systems surveyed in Section III (layered feature/texture coding, VCM, generative compression, Digital Retina) are explicit attempts to serve both objectives simultaneously. Please soften this claim to reflect that the gap is not directly comparable or not yet quantified, rather than impossible.
minor comments (5)
- [Section II, paragraph 2] The sentence 'Central to this efficiency is the HVS's ability to operate at extremely low bitrate representations, particularly from the primary visual cortex (V1) to extrastriate cortical areas [21]' is supported only by a 1956 study on the speed of visual perception; this reference does not appear to establish a quantitative bitrate claim, so please clarify the basis or replace the citation.
- [Section III-A1, paragraph 1] The phrase 'their theoretical limits of compression efficiency [8] is being constantly approached' invokes Shannon's work without explaining how the lossless source-coding bound applies to lossy perceptual video coding; please qualify or cite a specific analysis for video.
- [Section IV, AI-agent communication paragraph] The assertion that AI-agent communication is 'without any doubt' greener than human-centric communication is unsupported: no energy model, bitrate comparison, or lifecycle assessment is provided, and the computational cost of semantic communication is not considered.
- [Section I, Green Metadata sentence] The sentence 'the reduction of decoding complexity could vastly decrease energy savings' appears to be a typo; it likely means 'increase energy savings' or 'decrease energy consumption.'
- [Various] There are several wording issues that should be corrected in a revision: 'envelopes' should be 'encompasses' (Section II), 'researches' should be 'research directions' (Section III-A3), and 'sheds on how' should be 'sheds light on how' (Section III-B).
Circularity Check
No significant circularity: the paper is a survey whose motivating benchmark is external and whose self-citations are descriptive, not load-bearing.
full rationale
This is a survey paper: it introduces no fitted parameters, no predictions, and no equations whose outputs are determined by their inputs. The central motivating comparison (HVS compresses visual information roughly 100,000 times while VVC compresses raw visual data roughly 1,000 times) is attributed to [15], the SRC Decadal Plan for Semiconductors, which is not authored by the present authors and is therefore an external benchmark rather than a self-citation; even if this ratio is undefined or non-comparable, that is a correctness or evidence concern, not a circularity. The paper's many citations to work by the same authors (e.g., [9], [22], [23] on Digital Retina; [117], [119], [164], [167] on compact and feature coding) are descriptive survey references to prior publications, not invoked as proof of a derivation; removing them would not reduce any derived claim. The Section IV statement that reconstruction and understanding objectives make the gap 'impossible to bridge' is an argument about objectives, not a consequence obtained by substituting a fitted parameter or renaming an input. Hence no circular step can be quoted and exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The HVS achieves a compression ratio of around 100,000 times for visual information and consumes only 12 to 25 watts, serving as an appropriate benchmark for green multimedia.
- domain assumption For machine vision, transmitting features instead of pixel textures preserves task performance at much lower bitrates.
Cite this review
Pith. "Pith review of Compact Visual Data Representation for Green Multimedia -- A Human Visual System Perspective." pith.science (2026). https://pith.science/paper/TDC3DN3I
@misc{pith2026241114135,
author = {Pith},
title = {Pith review of: Compact Visual Data Representation for Green Multimedia -- A Human Visual System Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/TDC3DN3I}},
note = {Machine review of arXiv:2411.14135}
}
read the original abstract
The Human Visual System (HVS), with its intricate sophistication, is capable of achieving ultra-compact information compression for visual signals. This remarkable ability is coupled with high generalization capability and energy efficiency. By contrast, the state-of-the-art Versatile Video Coding (VVC) standard achieves a compression ratio of around 1,000 times for raw visual data. This notable disparity motivates the research community to draw inspiration to effectively handle the immense volume of visual data in a green way. Therefore, this paper provides a survey of how visual data can be efficiently represented for green multimedia, in particular when the ultimate task is knowledge extraction instead of visual signal reconstruction. We introduce recent research efforts that promote green, sustainable, and efficient multimedia in this field. Moreover, we discuss how the deep understanding of the HVS can benefit the research community, and envision the development of future green multimedia technologies.
Figures
Reference graph
Works this paper leans on
-
[15]
Decadal plan for semiconductors
S. R. Corporation, “Decadal plan for semiconductors.”
-
[1]
Internet of video things: Next-generation iot with visual sensors,
C. W. Chen, “Internet of video things: Next-generation iot with visual sensors,” IEEE Internet of Things Journal, vol. 7, no. 8, pp. 6676–6685, 2020
2020
-
[3]
Overview of the H.264/A VC video coding standard,
T. Wiegand, G. J. Sullivan, G. Bjontegaard, and A. Luthra, “Overview of the H.264/A VC video coding standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, no. 7, pp. 560–576, 2003
2003
-
[4]
Overview of the high efficiency video coding (HEVC) standard,
G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the high efficiency video coding (HEVC) standard,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 22, no. 12, pp. 1649– 1668, 2012
2012
-
[5]
Overview of the versatile video coding (VVC) standard and its applications,
B. Bross, Y .-K. Wang, Y . Ye, S. Liu, J. Chen, G. J. Sullivan, and J.-R. Ohm, “Overview of the versatile video coding (VVC) standard and its applications,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 10, pp. 3736–3764, 2021
2021
-
[6]
An overview of core coding tools in the av1 video codec,
Y . Chen, D. Murherjee, J. Han, A. Grange, Y . Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi et al. , “An overview of core coding tools in the av1 video codec,” in 2018 Picture Coding Symposium (PCS) . IEEE, 2018, pp. 41–45
2018
-
[7]
Recent development of A VS video coding standard: A VS3,
J. Zhang, C. Jia, M. Lei, S. Wang, S. Ma, and W. Gao, “Recent development of A VS video coding standard: A VS3,” in 2019 Picture Coding Symposium (PCS) , 2019, pp. 1–5
2019
-
[8]
A mathematical theory of communication,
C. E. Shannon, “A mathematical theory of communication,” The Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948
1948
Show all 187 references
-
[9]
Digital retina: A way to make the city brain more efficient by visual coding,
W. Gao, S. Ma, L. Duan, Y . Tian, P. Xing, Y . Wang, S. Wang, H. Jia, and T. Huang, “Digital retina: A way to make the city brain more efficient by visual coding,” IEEE Transactions on Circuits and Systems for Video Technology, 2021
2021
-
[10]
Rethinking semantic image compression: Scalable representation with cross-modality transfer,
P. Zhang, S. Wang, M. Wang, J. Li, X. Wang, and S. Kwong, “Rethinking semantic image compression: Scalable representation with cross-modality transfer,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 8, pp. 4441–4445, 2023
2023
-
[11]
Extended signaling methods for reduced video decoder power con- sumption using green metadata,
C. Herglotz, M. Kr ¨anzler, X. Chu, E. Franc ¸ois, Y . He, and A. Kaup, “Extended signaling methods for reduced video decoder power con- sumption using green metadata,” IEEE Transactions on Circuits and Systems II: Express Briefs , vol. 71, no. 3, pp. 1141–1145, 2024
2024
-
[12]
Energy reduction opportuni- ties in HDR video encoding,
C. Herglotz, S. Le Moan, and A. Mercat, “Energy reduction opportuni- ties in HDR video encoding,” in 2024 IEEE International Conference on Image Processing (ICIP) . IEEE, 2024, pp. 3654–3660
2024
-
[13]
Complexity metrics for VVC decoder power reduction in green metadata,
C. Herglotz, M. Kr ¨anzler, R. Dai, and A. Kaup, “Complexity metrics for VVC decoder power reduction in green metadata,” in 2024 Picture Coding Symposium (PCS) . IEEE, 2024, pp. 1–5
2024
-
[14]
Deep hierarchies in the primate visual cortex: What can we learn for computer vision?
N. Kruger, P. Janssen, S. Kalkan, M. Lappe, A. Leonardis, J. Piater, A. J. Rodriguez-Sanchez, and L. Wiskott, “Deep hierarchies in the primate visual cortex: What can we learn for computer vision?” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 35, no. 8...
2013
-
[16]
A system hierarchy for brain-inspired computing,
Y . Zhang, P. Qu, Y . Ji, W. Zhang, G. Gao, G. Wang, S. Song, G. Li, W. Chen, W. Zheng et al. , “A system hierarchy for brain-inspired computing,” Nature, vol. 586, no. 7829, pp. 378–384, 2020
2020
-
[17]
Video coding for machines: A paradigm of collaborative compression and intelligent analytics,
L. Duan, J. Liu, W. Yang, T. Huang, and W. Gao, “Video coding for machines: A paradigm of collaborative compression and intelligent analytics,” IEEE Transactions on Image Processing, vol. 29, pp. 8680– 8695, 2020
2020
-
[18]
Generative visual compression: A review,
B. Chen, S. Yin, P. Chen, S. Wang, and Y . Ye, “Generative visual compression: A review,” in 2024 IEEE International Conference on Image Processing (ICIP) , 2024, pp. 3709–3715
2024
-
[19]
Modern image quality assessment,
Z. Wang and A. C. Bovik, “Modern image quality assessment,” Ph.D. dissertation, Springer, 2006
2006
-
[20]
L. A. Remington and D. Goodwin, Clinical anatomy of the visual system E-Book. Elsevier Health Sciences, 2011
2011
-
[21]
Some studies in the speed of visual perception,
G. Sziklai, “Some studies in the speed of visual perception,” IRE Transactions on Information Theory , vol. 2, no. 3, pp. 125–128, 1956
1956
-
[22]
Towards digital retina in smart cities: A model generation, utilization and communication paradigm,
Y . Lou, L.-Y . Duan, Y . Luo, Z. Chen, T. Liu, S. Wang, and W. Gao, “Towards digital retina in smart cities: A model generation, utilization and communication paradigm,” in 2019 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2019, pp. 19–24
2019
-
[23]
Towards efficient front-end visual sensing for digital retina: A model-centric paradigm,
——, “Towards efficient front-end visual sensing for digital retina: A model-centric paradigm,” IEEE Transactions on Multimedia , vol. 22, no. 11, pp. 3002–3013, 2020
2020
-
[24]
SVT-A V1 encoding bitrate estimation using motion search information,
L. Eicherm ¨uller, G. Chaudhari, I. Katsavounidis, Z. Lei, H. Tmar, C. Herglotz, and A. Kaup, “SVT-A V1 encoding bitrate estimation using motion search information,” in 2024 32nd European Signal Processing Conference (EUSIPCO). IEEE, 2024, pp. 937–941
2024
-
[25]
Extended quad-tree partitioning for future video coding,
M. Wang, J. Li, L. Zhang, K. Zhang, H. Liu, S. Wang, S. Kwong, and S. Ma, “Extended quad-tree partitioning for future video coding,” in 2019 Data Compression Conference (DCC) , 2019, pp. 300–309
2019
-
[26]
Extended coding unit partitioning for future video coding,
——, “Extended coding unit partitioning for future video coding,” IEEE Transactions on Image Processing, vol. 29, pp. 2931–2946, 2020
2020
-
[27]
Unified intra mode coding based on short and long range correlations,
J. Li, M. Wang, L. Zhang, K. Zhang, H. Liu, S. Wang, S. Ma, and W. Gao, “Unified intra mode coding based on short and long range correlations,” IEEE Transactions on Image Processing , vol. 29, pp. 7245–7260, 2020
2020
-
[28]
CE 3.3 related: Intra 67 modes coding with 3 MPM,
N. Choi, Y . Piao, K. Choi, and C. Kim, “CE 3.3 related: Intra 67 modes coding with 3 MPM,” Joint Video Exploration Team (JVET), doc. JVET-K0529, Jul. 2018
2018
-
[29]
Enhanced cross-component linear model for chroma intra-prediction in video coding,
K. Zhang, J. Chen, L. Zhang, X. Li, and M. Karczewicz, “Enhanced cross-component linear model for chroma intra-prediction in video coding,” IEEE Transactions on Image Processing , vol. 27, no. 8, pp. 3983–3997, 2018
2018
-
[30]
Sub-sampled cross-component prediction for chroma component coding,
J. Li, M. Wang, L. Zhang, K. Zhang, S. Wang, S. Wang, S. Ma, and W. Gao, “Sub-sampled cross-component prediction for chroma component coding,” in 2020 Data Compression Conference (DCC) , 2020, pp. 203–212
2020
-
[31]
Prediction with multi-cross component,
J. Li, L. Zhang, K. Zhang, H. Liu, M. Wang, S. Wang, S. Ma, and W. Gao, “Prediction with multi-cross component,” in 2020 IEEE International Conference on Multimedia Expo Workshops (ICMEW) , 2020, pp. 1–6
2020
-
[32]
Intra block copy in A VS3 video coding standard,
Y . Wang, X. Xu, and S. Liu, “Intra block copy in A VS3 video coding standard,” in 2020 IEEE International Conference on Multimedia Expo Workshops (ICMEW), 2020, pp. 1–6
2020
-
[33]
History based block vector predictor for intra block copy,
W. Yin, J. Xu, L. Zhang, K. Zhang, H. Liu, and X. Fan, “History based block vector predictor for intra block copy,” in2020 IEEE International Conference on Multimedia Expo Workshops (ICMEW) , 2020, pp. 1–6
2020
-
[34]
Analysis of palette mode on versatile video coding,
Y . Sun, J. Lou, Y . Chao, H. Wang, V . Seregin, and M. Karczewicz, “Analysis of palette mode on versatile video coding,” in 2019 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), 2019, pp. 455–458
2019
-
[35]
String prediction for 4:2:0 format screen content coding and its implementation in A VS3,
Q. Zhou, L. Zhao, K. Zhou, T. Lin, H. Wang, S. Wang, and M. Jiao, “String prediction for 4:2:0 format screen content coding and its implementation in A VS3,”IEEE Transactions on Multimedia , vol. 23, pp. 3867–3876, 2021
2021
-
[36]
Affine direct/skip mode with motion vector differences in video coding,
T. Fu, K. Zhang, H. Liu, L. Zhang, S. Wang, S. Ma, and W. Gao, “Affine direct/skip mode with motion vector differences in video coding,” in 2020 IEEE International Conference on Multimedia Expo Workshops (ICMEW), 2020, pp. 1–6
2020
-
[37]
An improved framework of affine motion compensation in video coding,
K. Zhang, Y . Chen, L. Zhang, W. Chien, and M. Karczewicz, “An improved framework of affine motion compensation in video coding,” IEEE Transactions on Image Processing, vol. 28, no. 3, pp. 1456–1469, 2019
2019
-
[38]
Adaptive motion vector resolution for affine-inter mode coding,
H. Liu, L. Zhang, K. Zhang, J. Xu, Y . Wang, J. Luo, and Y . He, “Adaptive motion vector resolution for affine-inter mode coding,” in 2019 Picture Coding Symposium (PCS) , 2019, pp. 1–4
2019
-
[39]
History-based motion vector prediction in versatile video coding,
L. Zhang, K. Zhang, H. Liu, H. C. Chuang, Y . Wang, J. Xu, P. Zhao, and D. Hong, “History-based motion vector prediction in versatile video coding,” in 2019 Data Compression Conference (DCC) , 2019, pp. 43–52
2019
-
[40]
Samsung’s response to the call for proposals on video compression technology,
K. McCann, W.-J. Han, I.-K. Kim, J.-H. Min, E. Alshina, A. Alshin, T. Lee, J. Chen, V . Seregin, S. Lee et al., “Samsung’s response to the call for proposals on video compression technology,” JCTVC-A124, pp. 1–42, 2010. 12
2010
-
[41]
Bi-directional optical flow for improving motion compensation,
A. Alshin, E. Alshina, and T. Lee, “Bi-directional optical flow for improving motion compensation,” in 28th Picture Coding Symposium . IEEE, 2010, pp. 422–425
2010
-
[42]
Merge mode with motion vector difference,
S. Jeong, Y . Piao, M. W. Park, M. Park, A. Tamse, N. Choi, K. Choi, W. Choi, and C. Kim, “Merge mode with motion vector difference,” in 2020 IEEE International Conference on Image Processing (ICIP) . IEEE, 2020, pp. 1157–1160
2020
-
[43]
Decoder- side motion vector refinement in VVC: Algorithm and hardware implementation considerations,
H. Gao, X. Chen, S. Esenlik, J. Chen, and E. Steinbach, “Decoder- side motion vector refinement in VVC: Algorithm and hardware implementation considerations,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 8, pp. 3197–3211, 2021
2021
-
[44]
Hybrid video coding with trellis-coded quantization,
H. Schwarz, T. Nguyen, D. Marpe, and T. Wiegand, “Hybrid video coding with trellis-coded quantization,” in 2019 Data Compression Conference (DCC), March 2019, pp. 182–191
2019
-
[45]
Joint sep- arable and non-separable transforms for next-generation video coding,
X. Zhao, J. Chen, M. Karczewicz, A. Said, and V . Seregin, “Joint sep- arable and non-separable transforms for next-generation video coding,” IEEE Transactions on Image Processing, vol. 27, no. 5, pp. 2514–2525, May 2018
2018
-
[46]
Implicit-selected transform in video coding,
Y . Zhang, K. Zhang, L. Zhang, H. Liu, Y . Wang, S. Wang, S. Ma, and W. Gao, “Implicit-selected transform in video coding,” in 2020 IEEE International Conference on Multimedia Expo Workshops (ICMEW) , 2020, pp. 1–6
2020
-
[47]
Overlapped block motion compen- sation: An estimation-theoretic approach,
M. T. Orchard and G. J. Sullivan, “Overlapped block motion compen- sation: An estimation-theoretic approach,” IEEE Transactions on Image Processing, vol. 3, no. 5, pp. 693–699, 1994
1994
-
[48]
High Throughput CABAC Entropy Coding in HEVC,
V . Sze and M. Budagavi, “High Throughput CABAC Entropy Coding in HEVC,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 12, pp. 1778–1791, Dec 2012
2012
-
[49]
A multi-grained parallel solution for hevc encoding on heterogeneous platforms,
B. Xiao, H. Wang, J. Wu, S. Kwong, and C.-C. J. Kuo, “A multi-grained parallel solution for hevc encoding on heterogeneous platforms,” IEEE Transactions on Multimedia , vol. 21, no. 12, pp. 2997–3009, 2019
2019
-
[50]
Filtered intra template matching prediction for future video coding,
R. G. Youvalari, D. B. Sansli, P. Astola, and J. Lainema, “Filtered intra template matching prediction for future video coding,” in 2023 31st European Signal Processing Conference (EUSIPCO) . IEEE, 2023, pp. 576–579
2023
-
[51]
Video compression beyond vvc: Quantitative analysis of intra cod- ing tools in enhanced compression model (ECM),
M. Abdoli, R. G. Youvalari, K. Naser, K. Reuz ´e, and F. L. L ´eannec, “Video compression beyond vvc: Quantitative analysis of intra cod- ing tools in enhanced compression model (ECM),” arXiv preprint arXiv:2404.07872, 2024
2024 arXiv
-
[52]
Intra template matching prediction with fusion techniques,
F. Pu, T. Lu, P. Yin, S. McCarthy, J. R. Arumugam, A. Natesan, V . Valvaiker, J. N. Shingala, X. Li, R.-L. Liao et al. , “Intra template matching prediction with fusion techniques,” in 2024 Data Compres- sion Conference (DCC) . IEEE, 2024, pp. 93–102
2024
-
[53]
An improvement to subblock- based temporal motion vector prediction beyond VVC,
R.-L. Liao, J. Chen, Y . Ye, and X. Li, “An improvement to subblock- based temporal motion vector prediction beyond VVC,” in 2024 Data Compression Conference (DCC) . IEEE, 2024, pp. 83–92
2024
-
[54]
Inter cross-component prediction merge mode for video coding beyond VVC,
Z. Deng, K. Zhang, and L. Zhang, “Inter cross-component prediction merge mode for video coding beyond VVC,” in 2024 Data Compres- sion Conference (DCC) . IEEE, 2024, pp. 551–551
2024
-
[55]
A discrete-mapping-based cross-component prediction paradigm for screen content coding,
B. Vishwanath, K. Zhang, and L. Zhang, “A discrete-mapping-based cross-component prediction paradigm for screen content coding,” IEEE Transactions on Image Processing , vol. 33, pp. 16–26, 2023
2023
-
[56]
A survey of surface reconstruction from point clouds,
M. Berger et al, “A survey of surface reconstruction from point clouds,” in Computer Graphics Forum, vol. 36, no. 1, 2017, pp. 301–329
2017
-
[57]
Emerging mpeg standards for point cloud compression,
S. Schwarz, M. Preda, V . Baroncini, M. Budagavi, P. Cesar, P. A. Chou, R. A. Cohen, M. Krivoku ´ca, S. Lasserre, Z. Li et al., “Emerging mpeg standards for point cloud compression,”IEEE Journal on Emerging and Selected Topics in Circuits and Systems , vol. 9, no. 1, pp. 133–148, 2018
2018
-
[58]
3d mesh com- pression: Survey, comparisons, and emerging trends,
A. Maglo, G. Lavou ´e, F. Dupont, and C. Hudelot, “3d mesh com- pression: Survey, comparisons, and emerging trends,” ACM Computing Surveys (CSUR), vol. 47, no. 3, pp. 1–41, 2015
2015
-
[59]
Video coding optimiza- tion for virtual reality 360-degree source,
Y . Zhou, L. Tian, C. Zhu, X. Jin, and Y . Sun, “Video coding optimiza- tion for virtual reality 360-degree source,” IEEE Journal of Selected Topics in Signal Processing , vol. 14, no. 1, pp. 118–129, 2019
2019
-
[60]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
-
[61]
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.” ACM Transactions on Graphics, vol. 42, no. 4, pp. 139–1, 2023
2023
-
[62]
Video-based point- cloud-compression standard in MPEG: From evidence collection to committee draft [standards in a nutshell],
E. S. Jang, M. Preda, K. Mammou, A. M. Tourapis, J. Kim, D. B. Graziosi, S. Rhyu, and M. Budagavi, “Video-based point- cloud-compression standard in MPEG: From evidence collection to committee draft [standards in a nutshell],” IEEE Signal Processing Magazine, vol. 36, no. 3, p...
2019
-
[63]
Efficient projected frame padding for video-based point cloud compression,
L. Li, Z. Li, S. Liu, and H. Li, “Efficient projected frame padding for video-based point cloud compression,” IEEE Transactions on Multimedia, vol. 23, pp. 2806–2819, 2020
2020
-
[64]
Advanced 3D motion prediction for video-based dynamic point cloud compression,
L. Li, Z. Li, V . Zakharchenko, J. Chen, and H. Li, “Advanced 3D motion prediction for video-based dynamic point cloud compression,” IEEE Trans. Image Process. , vol. 29, pp. 289–302, 2019
2019
-
[65]
Video-based point cloud compression artifact removal,
A. Akhtar, W. Gao, L. Li, Z. Li, W. Jia, and S. Liu, “Video-based point cloud compression artifact removal,” IEEE Transactions on Multimedia, vol. 24, pp. 2866–2876, 2022
2022
-
[66]
An overview of ongoing point cloud compression standardization activities: Video-based (V-PCC) and geometry-based (G-PCC),
D. Graziosi, O. Nakagami, S. Kuma, A. Zaghetto, T. Suzuki, and A. Tabatabai, “An overview of ongoing point cloud compression standardization activities: Video-based (V-PCC) and geometry-based (G-PCC),” APSIPA Transactions on Signal and Information Processing, vol. 9, 2020
2020
-
[67]
An improved enhancement layer for octree based point cloud compression with plane projection approximation,
K. Ainala et al, “An improved enhancement layer for octree based point cloud compression with plane projection approximation,” in Appl. digit. image process. XXXIX , vol. 9971. SPIE, 2016, pp. 223–231
2016
-
[68]
Immersive video postprocessing for efficient video coding,
A. Dziembowski, D. Mieloch, J. Y . Jeong, and G. Lee, “Immersive video postprocessing for efficient video coding,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 8, pp. 4349– 4361, 2023
2023
-
[69]
End-to-end optimized image compression,
J. Ball ´e, V . Laparra, and E. Simoncelli, “End-to-end optimized image compression,” in International Conference on Learning Representa- tions, 2016
2016
-
[70]
End-to-end optimization of nonlinear transform codes for perceptual quality,
J. Ball ´e, V . Laparra, and E. P. Simoncelli, “End-to-end optimization of nonlinear transform codes for perceptual quality,” in 2016 Picture Coding Symposium (PCS) . IEEE, 2016, pp. 1–5
2016
-
[71]
Vari- ational image compression with a scale hyperprior,
J. Ball ´e, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston, “Vari- ational image compression with a scale hyperprior,” in International Conference on Learning Representations , 2018
2018
-
[72]
Learned image com- pression with discretized gaussian mixture likelihoods and attention modules,
Z. Cheng, H. Sun, M. Takeuchi, and J. Katto, “Learned image com- pression with discretized gaussian mixture likelihoods and attention modules,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 7939–7948
2020
-
[74]
Joint graph attention and asymmetric convolutional neural network for deep image compression,
Z. Tang, H. Wang, X. Yi, Y . Zhang, S. Kwong, and C.-C. J. Kuo, “Joint graph attention and asymmetric convolutional neural network for deep image compression,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 1, pp. 421–433, 2023
2023
-
[75]
A neural-network enhanced video coding framework beyond ECM,
Y . Zhao, W. He, C. Jia, Q. Wang, J. Li, Y . Li, C. Lin, K. Zhang, L. Zhang, and S. Ma, “A neural-network enhanced video coding framework beyond ECM,” in 2024 Data Compression Conference (DCC), 2024, pp. 605–605
2024
-
[76]
DVC: An end-to-end deep video compression framework,
G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao, “DVC: An end-to-end deep video compression framework,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 10 998–11 007
2019
-
[77]
An end- to-end learning framework for video compression,
G. Lu, X. Zhang, W. Ouyang, L. Chen, Z. Gao, and D. Xu, “An end- to-end learning framework for video compression,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 43, no. 10, pp. 3292–3308, 2021
2021
-
[78]
Learning for video compression with hierarchical quality and recurrent enhancement,
R. Yang, F. Mentzer, L. V . Gool, and R. Timofte, “Learning for video compression with hierarchical quality and recurrent enhancement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 6628–6637
2020
-
[79]
Deep contextual video compression,
J. Li, B. Li, and Y . Lu, “Deep contextual video compression,” Advances in Neural Information Processing Systems , vol. 34, pp. 18 114–18 125, 2021
2021
-
[80]
Temporal context min- ing for learned video compression,
X. Sheng, J. Li, B. Li, L. Li, D. Liu, and Y . Lu, “Temporal context min- ing for learned video compression,” IEEE Transactions on Multimedia, vol. 25, pp. 7311–7322, 2022
2022
-
[81]
Hybrid spatial-temporal entropy modelling for neural video compression,
J. Li, B. Li, and Y . Lu, “Hybrid spatial-temporal entropy modelling for neural video compression,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 1503–1511
2022
-
[82]
Neural video compression with diverse contexts,
——, “Neural video compression with diverse contexts,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 22 616–22 626
2023
-
[83]
Neural video compression with feature modulation,
——, “Neural video compression with feature modulation,” in Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 099–26 108
2024
-
[84]
NVC-1B: A large neural video coding model,
X. Sheng, C. Tang, L. Li, D. Liu, and F. Wu, “NVC-1B: A large neural video coding model,” arXiv preprint arXiv:2407.19402 , 2024
2024 arXiv
-
[85]
EVC: Towards real-time neural image compression with mask decay,
W. Guo-Hua, J. Li, B. Li, and Y . Lu, “EVC: Towards real-time neural image compression with mask decay,” in The Eleventh International 13 Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=XUxad2Gj40n
2023
-
[86]
Perceptual quality-oriented rate allocation via distillation from end-to-end image compression,
R. Yang, D. Liu, S. Ma, F. Wu, and W. Gao, “Perceptual quality-oriented rate allocation via distillation from end-to-end image compression,” ACM Trans. Multimedia Comput. Commun. Appl., vol. 20, no. 7, Apr. 2024. [Online]. Available: https: //doi.org/10.1145/3650034
2024 doi
-
[87]
Towards hybrid-optimization video coding,
S. Huo, D. Liu, H. Zhang, L. Li, S. Ma, F. Wu, and W. Gao, “Towards hybrid-optimization video coding,” ACM Comput. Surv., vol. 56, no. 9, Apr. 2024. [Online]. Available: https://doi.org/10.1145/3652148
2024 doi
-
[88]
Power-aware content-adaptive H.264 video encoding,
A. K. Kannur and B. Li, “Power-aware content-adaptive H.264 video encoding,” in 2009 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2009, pp. 925–928
2009
-
[89]
Game as video: Bit rate reduction through adaptive object encoding,
M. Hemmati, A. Javadtalab, A. A. Nazari Shirehjini, S. Shirmoham- madi, and T. Arici, “Game as video: Bit rate reduction through adaptive object encoding,” in Proceeding of the 23rd ACM Workshop on Network and Operating Systems Support for Digital Audio and Video , 2013, pp. 7–12
2013
-
[90]
Perceptual compression for video storage and processing systems,
A. Mazumdar, B. Haynes, M. Balazinska, L. Ceze, A. Cheung, and M. Oskin, “Perceptual compression for video storage and processing systems,” in Proceedings of the ACM Symposium on Cloud Computing, 2019, pp. 179–192
2019
-
[91]
Perceptual visual signal compression and transmission,
H. R. Wu, A. R. Reibman, W. Lin, F. Pereira, and S. S. Hemami, “Perceptual visual signal compression and transmission,” Proceedings of the IEEE , vol. 101, no. 9, pp. 2025–2043, 2013
2025
-
[92]
A survey on perceptually optimized video coding,
Y . Zhang, L. Zhu, G. Jiang, S. Kwong, and C.-C. J. Kuo, “A survey on perceptually optimized video coding,” ACM Computing Surveys , vol. 55, no. 12, pp. 1–37, 2023
2023
-
[93]
FSIM: A feature similarity index for image quality assessment,
L. Zhang, L. Zhang, X. Mou, and D. Zhang, “FSIM: A feature similarity index for image quality assessment,” IEEE transactions on Image Processing, vol. 20, no. 8, pp. 2378–2386, 2011
2011
-
[94]
Characterizing perceptual artifacts in compressed video streams,
K. Zeng, T. Zhao, A. Rehman, and Z. Wang, “Characterizing perceptual artifacts in compressed video streams,” in Human vision and electronic imaging XIX, vol. 9014. SPIE, 2014, pp. 173–182
2014
-
[95]
Toward a practical perceptual video quality metric,
Z. Li, A. Anne, K. Ioannis, M. Anush, and M. Megha, “Toward a practical perceptual video quality metric,” 2016. [Online]. Available: https://netflixtechblog.com/ toward-a-practical-perceptual-video-quality-metric-653f208b9652
2016
-
[96]
VMAF oriented perceptual coding based on piecewise metric coupling,
Z. Luo, C. Zhu, Y . Huang, R. Xie, L. Song, and C.-C. J. Kuo, “VMAF oriented perceptual coding based on piecewise metric coupling,” IEEE Transactions on Image Processing , vol. 30, pp. 5109–5121, 2021
2021
-
[97]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595
2018
-
[98]
Motion-compensated residue preprocessing in video coding based on just-noticeable- distortion profile,
X. Yang, W. Lin, Z. Lu, E. Ong, and S. Yao, “Motion-compensated residue preprocessing in video coding based on just-noticeable- distortion profile,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 15, no. 6, pp. 742–752, 2005
2005
-
[99]
Learning-based just-noticeable- quantization-distortion modeling for perceptual video coding,
S. Ki, S.-H. Bae, M. Kim, and H. Ko, “Learning-based just-noticeable- quantization-distortion modeling for perceptual video coding,” IEEE Transactions on Image Processing, vol. 27, no. 7, pp. 3178–3193, 2018
2018
-
[100]
Just noticeable difference level prediction for perceptual image compression,
T. Tian, H. Wang, L. Zuo, C.-C. J. Kuo, and S. Kwong, “Just noticeable difference level prediction for perceptual image compression,” IEEE Transactions on Broadcasting, vol. 66, no. 3, pp. 690–700, 2020
2020
-
[101]
BL- JUNIPER: A CNN-assisted framework for perceptual video coding leveraging block-level jnd,
S. Nami, F. Pakdaman, M. R. Hashemi, and S. Shirmohammadi, “BL- JUNIPER: A CNN-assisted framework for perceptual video coding leveraging block-level jnd,” IEEE Transactions on Multimedia, vol. 25, pp. 5077–5092, 2022
2022
-
[102]
Signal compression based on models of human perception,
N. Jayant, J. Johnston, and R. Safranek, “Signal compression based on models of human perception,” Proceedings of the IEEE , vol. 81, no. 10, pp. 1385–1422, 1993
1993
-
[103]
Saliency-aware video compression,
H. Hadizadeh and I. V . Baji ´c, “Saliency-aware video compression,” IEEE Transactions on Image Processing , vol. 23, no. 1, pp. 19–33, 2013
2013
-
[104]
Ssim-motivated rate-distortion optimization for video coding,
S. Wang, A. Rehman, Z. Wang, S. Ma, and W. Gao, “Ssim-motivated rate-distortion optimization for video coding,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, no. 4, pp. 516–529, 2011
2011
-
[105]
Proxiqa: A proxy approach to perceptual optimization of learned image compression,
L.-H. Chen, C. G. Bampis, Z. Li, A. Norkin, and A. C. Bovik, “Proxiqa: A proxy approach to perceptual optimization of learned image compression,” IEEE Transactions on Image Processing, vol. 30, pp. 360–373, 2020
2020
-
[106]
Perceptually adaptive lagrangian multiplier for hevc guided rate-distortion optimization,
K. Rouis, M.-C. Larabi, and J. B. Tahar, “Perceptually adaptive lagrangian multiplier for hevc guided rate-distortion optimization,” IEEE Access, vol. 6, pp. 33 589–33 603, 2018
2018
-
[107]
Mazumdar, Perceptual Optimizations for Video Capture, Processing, and Storage Systems
A. Mazumdar, Perceptual Optimizations for Video Capture, Processing, and Storage Systems . University of Washington, 2020
2020
-
[108]
Ed: Perceptually tuned enhanced compression model,
P. Philippe, T. Ladune, S. Davenet, and T. Leguay, “Ed: Perceptually tuned enhanced compression model,” arXiv preprint arXiv:2401.02145, 2024
2024 arXiv
-
[109]
Deep perceptual preprocessing for video coding,
A. Chadha and Y . Andreopoulos, “Deep perceptual preprocessing for video coding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 852–14 861
2021
-
[110]
Gop-based deep preprocessing for video coding,
D. Arai, S. Iwamura, K. Iguchi, and A. Ichigaya, “Gop-based deep preprocessing for video coding,” in 2024 Picture Coding Symposium (PCS). IEEE, 2024, pp. 1–5
2024
-
[111]
Cloud-based image coding for mobile devices—toward thousands to one compression,
H. Yue, X. Sun, J. Yang, and F. Wu, “Cloud-based image coding for mobile devices—toward thousands to one compression,” IEEE Transactions on Multimedia , vol. 15, no. 4, pp. 845–857, 2013
2013
-
[112]
Joint compression of near- duplicate videos,
H. Wang, T. Tian, M. Ma, and J. Wu, “Joint compression of near- duplicate videos,” IEEE Transactions on Multimedia , vol. 19, no. 5, pp. 908–920, 2017
2017
-
[113]
Peering into the sketch: Ultra-low bitrate face compression for joint human and machine perception,
Y . Mao, P. Chen, S. Wang, S. Wang, and D. Wu, “Peering into the sketch: Ultra-low bitrate face compression for joint human and machine perception,” in Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 2564–2572
2023
-
[114]
Low bandwidth video- chat compression using deep generative models,
M. Oquab, P. Stock, D. Haziza, T. Xu, P. Zhang, O. Celebi, Y . Hasson, P. Labatut, B. Bose-Kolanu, T. Peyronel et al., “Low bandwidth video- chat compression using deep generative models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2...
2021
-
[115]
Generative compression for face video: A hybrid scheme,
A. Tang, Y . Huang, J. Ling, Z. Zhang, Y . Zhang, R. Xie, and L. Song, “Generative compression for face video: A hybrid scheme,” in 2022 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2022, pp. 1–6
2022
-
[116]
Towards ultra low bit-rate digital human character communication via compact 3d face descriptors,
B. Li, B. Chen, Z. Wang, S. Wang, and Y . Ye, “Towards ultra low bit-rate digital human character communication via compact 3d face descriptors,” in 2022 Data Compression Conference (DCC) . IEEE, 2022, pp. 461–461
2022
-
[117]
Compact temporal trajectory representation for talking face video compression,
B. Chen, Z. Wang, B. Li, S. Wang, and Y . Ye, “Compact temporal trajectory representation for talking face video compression,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 33, no. 11, pp. 7009–7023, 2023
2023
-
[118]
Conceptual compression via deep structure and texture synthesis,
J. Chang, Z. Zhao, C. Jia, S. Wang, L. Yang, Q. Mao, J. Zhang, and S. Ma, “Conceptual compression via deep structure and texture synthesis,” IEEE Transactions on Image Processing, vol. 31, pp. 2809– 2823, 2022
2022
-
[119]
Dynamic multi-reference gen- erative prediction for face video compression,
Z. Wang, B. Chen, Y . Ye, and S. Wang, “Dynamic multi-reference gen- erative prediction for face video compression,” in IEEE International Conference on Image Processing (ICIP) , 2022, pp. 896–900
2022
-
[120]
When video coding meets multimodal large language models: A unified paradigm for video coding,
P. Zhang, J. Li, M. Wang, N. Sebe, S. Kwong, and S. Wang, “When video coding meets multimodal large language models: A unified paradigm for video coding,” arXiv preprint arXiv:2408.08093 , 2024
2024 arXiv
-
[121]
Compress-then-analyze versus analyze-then-compress: What is best in visual sensor networks?
A. Redondi, L. Baroffio, L. Bianchi, M. Cesana, and M. Tagliasacchi, “Compress-then-analyze versus analyze-then-compress: What is best in visual sensor networks?” IEEE Transactions on Mobile Computing , vol. 15, no. 12, pp. 3000–3013, 2016
2016
-
[122]
Spectral hashing,
Y . Weiss, A. Torralba, and R. Fergus, “Spectral hashing,” Advances in Neural Information Processing Systems, vol. 21, pp. 1753–1760, 2008
2008
-
[123]
Transform coding of image feature descriptors,
V . Chandrasekhar, G. Takacs, D. Chen, S. S. Tsai, J. Singh, and B. Girod, “Transform coding of image feature descriptors,” in Visual Communications and Image Processing 2009, vol. 7257. International Society for Optics and Photonics, 2009, p. 725710
2009
-
[124]
Product quantization for nearest neighbor search,
H. J ´egou, M. Douze, and C. Schmid, “Product quantization for nearest neighbor search,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, no. 1, pp. 117–128, 2011
2011
-
[125]
Optimized product quantization,
T. Ge, K. He, Q. Ke, and J. Sun, “Optimized product quantization,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 36, no. 4, pp. 744–755, 2014
2014
-
[126]
BRIEF: Computing a local binary descriptor very fast,
M. Calonder, V . Lepetit, M. Ozuysal, T. Trzcinski, C. Strecha, and P. Fua, “BRIEF: Computing a local binary descriptor very fast,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 34, no. 7, pp. 1281–1298, 2012
2012
-
[127]
ORB: An efficient alternative to sift or surf,
E. Rublee, V . Rabaud, K. Konolige, and G. Bradski, “ORB: An efficient alternative to sift or surf,” in 2011 International Conference on Computer Vision , 2011, pp. 2564–2571
2011
-
[128]
BRISK: Binary robust invariant scalable keypoints,
S. Leutenegger, M. Chli, and R. Y . Siegwart, “BRISK: Binary robust invariant scalable keypoints,” in 2011 International Conference on Computer Vision, 2011, pp. 2548–2555
2011
-
[129]
USB: ultrashort binary descriptor for fast visual matching and retrieval,
S. Zhang, Q. Tian, Q. Huang, W. Gao, and Y . Rui, “USB: ultrashort binary descriptor for fast visual matching and retrieval,” IEEE Trans- actions on Image Processing , vol. 23, no. 8, pp. 3671–3683, 2014. 14
2014
-
[130]
Aggregating local de- scriptors into a compact image representation,
H. J ´egou, M. Douze, C. Schmid, and P. P ´erez, “Aggregating local de- scriptors into a compact image representation,” in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2010, pp. 3304–3311
2010
-
[131]
Large-scale image retrieval with compressed fisher vectors,
F. Perronnin, Y . Liu, J. S ´anchez, and H. Poirier, “Large-scale image retrieval with compressed fisher vectors,” in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition . IEEE, 2010, pp. 3384–3391
2010
-
[132]
Residual enhanced visual vector as a compact signature for mobile visual search,
D. Chen, S. Tsai, V . Chandrasekhar, G. Takacs, R. Vedantham, R. Grzeszczuk, and B. Girod, “Residual enhanced visual vector as a compact signature for mobile visual search,” Signal Processing, vol. 93, no. 8, pp. 2316–2327, 2013
2013
-
[133]
Residual enhanced visual vectors for on-device image matching,
D. Chen, S. Tsai, V . Chandrasekhar, G. Takacs, H. Chen, R. Vedantham, R. Grzeszczuk, and B. Girod, “Residual enhanced visual vectors for on-device image matching,” in 2011 Conference Record of the Forty Fifth Asilomar Conference on Signals, Systems and Computers (ASILOMAR). I...
2011
-
[134]
Rate- adaptive compact fisher codes for mobile visual search,
J. Lin, L.-Y . Duan, Y . Huang, S. Luo, T. Huang, and W. Gao, “Rate- adaptive compact fisher codes for mobile visual search,” IEEE Signal Processing Letters, vol. 21, no. 2, pp. 195–198, 2014
2014
-
[135]
Supervised hashing with kernels,
W. Liu, J. Wang, R. Ji, Y .-G. Jiang, and S.-F. Chang, “Supervised hashing with kernels,” in 2012 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2012, pp. 2074–2081
2012
-
[136]
Supervised discrete hashing,
F. Shen, C. Shen, W. Liu, and H. T. Shen, “Supervised discrete hashing,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 37–45
2015
-
[137]
A survey on deep hashing methods,
X. Luo, H. Wang, D. Wu, C. Chen, M. Deng, J. Huang, and X.-S. Hua, “A survey on deep hashing methods,”ACM Transactions on Knowledge Discovery from Data , vol. 17, no. 1, pp. 1–50, 2023
2023
-
[138]
One-bit deep hashing: Towards resource-efficient hashing model with binary neural network,
L. He, Z. Huang, C. Liu, R. Li, R. Wu, Q. Liu, and E. Chen, “One-bit deep hashing: Towards resource-efficient hashing model with binary neural network,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 7162–7171
2024
-
[139]
Entropy-optimized deep weighted product quantization for image retrieval,
L. Gu, J. Liu, X. Liu, W. Wan, and J. Sun, “Entropy-optimized deep weighted product quantization for image retrieval,” IEEE Transactions on Image Processing , vol. 33, pp. 1162–1174, 2024
2024
-
[140]
Two-step discrete hashing for cross-modal retrieval,
J. Tu, X. Liu, Y . Hao, R. Hong, and M. Wang, “Two-step discrete hashing for cross-modal retrieval,” IEEE Transactions on Multimedia , vol. 26, pp. 8730–8741, 2024
2024
-
[141]
HNIP: Compact deep invariant representations for video matching, localization, and retrieval,
J. Lin, L.-Y . Duan, S. Wang, Y . Bai, Y . Lou, V . Chandrasekhar, T. Huang, A. Kot, and W. Gao, “HNIP: Compact deep invariant representations for video matching, localization, and retrieval,” IEEE Transactions on Multimedia , vol. 19, no. 9, pp. 1968–1983, 2017
1968
-
[142]
Rate-performance- loss optimization for inter-frame deep feature coding from videos,
L. Ding, Y . Tian, H. Fan, Y . Wang, and T. Huang, “Rate-performance- loss optimization for inter-frame deep feature coding from videos,” IEEE Transactions on Image Processing , vol. 26, no. 12, pp. 5743– 5757, 2017
2017
-
[143]
Joint coding of local and global deep features in videos for visual search,
L. Ding, Y . Tian, H. Fan, C. Chen, and T. Huang, “Joint coding of local and global deep features in videos for visual search,” IEEE Transactions on Image Processing , vol. 29, pp. 3734–3749, 2020
2020
-
[144]
Neurosurgeon: Collaborative intelligence between the cloud and mobile edge,
Y . Kang, J. Hauswald, C. Gao, A. Rovinski, T. Mudge, J. Mars, and L. Tang, “Neurosurgeon: Collaborative intelligence between the cloud and mobile edge,” ACM SIGARCH Computer Architecture News, vol. 45, no. 1, pp. 615–629, 2017
2017
-
[145]
Tensor completion methods for collaborative intelligence,
L. Bragilevsky and I. V . Baji ´c, “Tensor completion methods for collaborative intelligence,” IEEE Access , vol. 8, pp. 41 162–41 174, 2020
2020
-
[146]
Multi-task learning with compressible features for collaborative intelligence,
S. R. Alvar and I. V . Baji ´c, “Multi-task learning with compressible features for collaborative intelligence,” in 2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019, pp. 1705–1709
2019
-
[147]
Bit allocation for multi-task collaborative intelligence,
——, “Bit allocation for multi-task collaborative intelligence,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 4342–4346
2020
-
[148]
Lightweight compression of neural network feature tensors for collaborative intelligence,
R. A. Cohen, H. Choi, and I. V . Baji ´c, “Lightweight compression of neural network feature tensors for collaborative intelligence,” in 2020 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2020, pp. 1–6
2020
-
[149]
Toward intelligent sensing: Intermediate deep feature compression,
Z. Chen, K. Fan, S. Wang, L. Duan, W. Lin, and A. C. Kot, “Toward intelligent sensing: Intermediate deep feature compression,” IEEE Transactions on Image Processing , vol. 29, pp. 2230–2243, 2019
2019
-
[150]
Data repre- sentation in hybrid coding framework for feature maps compression,
Z. Chen, L.-Y . Duan, S. Wang, W. Lin, and A. C. Kot, “Data repre- sentation in hybrid coding framework for feature maps compression,” in 2020 IEEE International Conference on Image Processing (ICIP) . IEEE, 2020, pp. 3094–3098
2020
-
[151]
End-to-end compression to- wards machine vision: Network architecture design and optimization,
S. Wang, Z. Wang, S. Wang, and Y . Ye, “End-to-end compression to- wards machine vision: Network architecture design and optimization,” IEEE Open Journal of Circuits and Systems, vol. 2, pp. 675–685, 2021
2021
-
[152]
Image coding for machines: an end-to-end learned approach,
N. Le, H. Zhang, F. Cricri, R. Ghaznavi-Youvalari, and E. Rahtu, “Image coding for machines: an end-to-end learned approach,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 1590–1594
2021
-
[153]
A coding framework and benchmark towards low-bitrate video understanding,
Y . Tian, G. Lu, Y . Yan, G. Zhai, L. Chen, and Z. Gao, “A coding framework and benchmark towards low-bitrate video understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 8, pp. 5852–5872, 2024
2024
-
[154]
Collaborative intelligence: Challenges and opportunities,
I. V . Baji´c, W. Lin, and Y . Tian, “Collaborative intelligence: Challenges and opportunities,” in 2021 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP) , 2021, pp. 8493–8497
2021
-
[155]
Pareto-optimal bit allocation for collabora- tive intelligence,
S. R. Alvar and I. V . Baji´c, “Pareto-optimal bit allocation for collabora- tive intelligence,” IEEE Trans. Image Process., vol. 30, pp. 3348–3361, 2021
2021
-
[156]
Transtic: Transferring transformer-based image compression from human perception to machine perception,
Y .-H. Chen, Y .-C. Weng, C.-H. Kao, C. Chien, W.-C. Chiu, and W.- H. Peng, “Transtic: Transferring transformer-based image compression from human perception to machine perception,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 23 297–23 307
2023
-
[157]
Overview of the MPEG-CDVS standard,
L.-Y . Duan, V . Chandrasekhar, J. Chen, J. Lin, Z. Wang, T. Huang, B. Girod, and W. Gao, “Overview of the MPEG-CDVS standard,”IEEE Transactions on Image Processing , vol. 25, no. 1, pp. 179–194, 2015
2015
-
[158]
A low complexity interest point detector,
J. Chen, L.-Y . Duan, F. Gao, J. Cai, A. C. Kot, and T. Huang, “A low complexity interest point detector,” IEEE Signal Processing Letters , vol. 22, no. 2, pp. 172–176, 2014
2014
-
[159]
CDVS CE2: Local descriptor compression proposal,
S. Paschalakis, K. Wnukowicz, M. Bober, A. Mosca, and M. Mattel- liano, “CDVS CE2: Local descriptor compression proposal,” ISO/IEC JTC1/SC29/WG11 M, vol. 25929, 2012
2012
-
[160]
Location coding for mobile image retrieval,
S. S. Tsai, D. Chen, G. Takacs, V . Chandrasekhar, J. P. Singh, and B. Girod, “Location coding for mobile image retrieval,” in Proceedings of the 5th International ICST Mobile Multimedia Communications Conference, 2009, pp. 1–7
2009
-
[161]
Improved coding for image feature location information,
S. S. Tsai, D. Chen, G. Takacs, V . Chandrasekhar, M. Makar, R. Grzeszczuk, and B. Girod, “Improved coding for image feature location information,” in Applications of Digital Image Processing XXXV, vol. 8499. International Society for Optics and Photonics, 2012, p. 84991E
2012
-
[162]
Compact descriptors for video analysis: The emerging MPEG standard,
L.-Y . Duan, Y . Lou, Y . Bai, T. Huang, W. Gao, V . Chandrasekhar, J. Lin, S. Wang, and A. C. Kot, “Compact descriptors for video analysis: The emerging MPEG standard,” IEEE MultiMedia, vol. 26, no. 2, pp. 44– 54, 2018
2018
-
[163]
Video coding for machines: Compact visual representation compression for intelligent collaborative analytics,
W. Yang, H. Huang, Y . Hu, L.-Y . Duan, and J. Liu, “Video coding for machines: Compact visual representation compression for intelligent collaborative analytics,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 7, pp. 5174–5191, 2024
2024
-
[164]
Scalable facial image compression with deep feature reconstruction,
S. Wang, S. Wang, X. Zhang, S. Wang, S. Ma, and W. Gao, “Scalable facial image compression with deep feature reconstruction,” in 2019 IEEE International Conference on Image Processing (ICIP) . IEEE, 2019, pp. 2691–2695
2019
-
[165]
Deepsvc: Deep scalable video coding for both machine and human vision,
H. Lin, B. Chen, Z. Zhang, J. Lin, X. Wang, and T. Zhao, “Deepsvc: Deep scalable video coding for both machine and human vision,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 9205–9214
2023
-
[166]
Real-time action recognition with enhanced motion vector cnns,
B. Zhang, L. Wang, Z. Wang, Y . Qiao, and H. Wang, “Real-time action recognition with enhanced motion vector cnns,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 2718–2726
2016
-
[167]
Joint feature and texture coding: Toward smart video representation via front- end intelligence,
S. Ma, X. Zhang, S. Wang, X. Zhang, C. Jia, and S. Wang, “Joint feature and texture coding: Toward smart video representation via front- end intelligence,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 29, no. 10, pp. 3095–3105, 2018
2018
-
[168]
A joint compression scheme of video feature descriptors and visual content,
X. Zhang, S. Ma, S. Wang, X. Zhang, H. Sun, and W. Gao, “A joint compression scheme of video feature descriptors and visual content,” IEEE Transactions on Image Processing , vol. 26, no. 2, pp. 633–647, 2016
2016
-
[169]
Joint rate-distortion optimization for simultaneous texture and deep feature compression of facial images,
Y . Li, C. Jia, S. Wang, X. Zhang, S. Wang, S. Ma, and W. Gao, “Joint rate-distortion optimization for simultaneous texture and deep feature compression of facial images,” in 2018 IEEE Fourth International Conference on Multimedia Big Data (BigMM) . IEEE, 2018, pp. 1–5
2018
-
[170]
Semantic structured image coding framework for multiple intelligent applications,
S. Sun, T. He, and Z. Chen, “Semantic structured image coding framework for multiple intelligent applications,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 31, no. 9, pp. 3631– 3642, 2020
2020
-
[171]
Towards analysis-friendly face representation with scalable feature and texture compression,
S. Wang, S. Wang, W. Yang, X. Zhang, S. Wang, S. Ma, and W. Gao, “Towards analysis-friendly face representation with scalable feature and texture compression,” IEEE Transactions on Multimedia , vol. 24, pp. 3169–3181, 2021. 15
2021
-
[172]
Neo- cortex saves energy by reducing coding precision during food scarcity,
Z. Padamsey, D. Katsanevaki, N. Dupuy, and N. L. Rochefort, “Neo- cortex saves energy by reducing coding precision during food scarcity,” Neuron, vol. 110, no. 2, pp. 280–296, 2022
2022
-
[173]
Scalable image coding for humans and machines,
H. Choi and I. V . Baji ´c, “Scalable image coding for humans and machines,” IEEE Transactions on Image Processing, vol. 31, pp. 2739– 2754, 2022
2022
-
[174]
Task-driven video compression for humans and machines: Framework design and opti- mization,
X. Yi, H. Wang, S. Kwong, and C.-C. Jay Kuo, “Task-driven video compression for humans and machines: Framework design and opti- mization,” IEEE Transactions on Multimedia , vol. 25, pp. 8091–8102, 2023
2023
-
[175]
Vnvc: A versatile neural video coding framework for efficient human-machine vision,
X. Sheng, L. Li, D. Liu, and H. Li, “Vnvc: A versatile neural video coding framework for efficient human-machine vision,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 7, pp. 4579–4596, 2024
2024
-
[176]
Task-aware encoder control for deep video compression,
X. Ge, J. Luo, X. Zhang, T. Xu, G. Lu, D. He, J. Geng, Y . Wang, J. Zhang, and H. Qin, “Task-aware encoder control for deep video compression,” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , 2024, pp. 26 036–26 045
2024
-
[177]
Human–machine collaborative image compres- sion method based on implicit neural representations,
H. Li and X. Zhang, “Human–machine collaborative image compres- sion method based on implicit neural representations,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems , vol. 14, no. 2, pp. 198–208, 2024
2024
-
[178]
Unified and scalable deep image compression framework for human and machine,
G. Zhang, X. Zhang, and L. Tang, “Unified and scalable deep image compression framework for human and machine,” ACM Trans. Multimedia Comput. Commun. Appl. , vol. 20, no. 10, Oct. 2024. [Online]. Available: https://doi.org/10.1145/3678472
2024 doi
-
[179]
A survey of ai-generated content (aigc),
Y . Cao, S. Li, Y . Liu, Z. Yan, Y . Dai, P. Yu, and L. Sun, “A survey of ai-generated content (aigc),” ACM Comput. Surv. , Dec. 2024, just Accepted. [Online]. Available: https://doi.org/10.1145/3704262
2024 doi
-
[180]
Aigc for various data modalities: A survey,
L. G. Foo, H. Rahmani, and J. Liu, “Aigc for various data modalities: A survey,” arXiv preprint arXiv:2308.14177 , 2023
2023 arXiv
-
[181]
Unleashing the power of edge-cloud generative AI in mobile networks: A survey of AIGC services,
M. Xu, H. Du, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, A. Jamalipour, D. I. Kim, X. Shen, V . C. M. Leung, and H. V . Poor, “Unleashing the power of edge-cloud generative AI in mobile networks: A survey of AIGC services,” IEEE Communications Surveys & Tutorials, vol. 26, ...
2024
-
[182]
High efficiency image compres- sion for large visual-language models,
B. Li, S. Wang, S. Wang, and Y . Ye, “High efficiency image compres- sion for large visual-language models,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2024
2024
-
[183]
Vision and challenges for knowledge centric networking,
D. Wu, Z. Li, J. Wang, Y . Zheng, M. Li, and Q. Huang, “Vision and challenges for knowledge centric networking,” IEEE Wireless Communications, vol. 26, no. 4, pp. 117–123, 2019
2019
-
[184]
Edge computing: Vision and challenges,
W. Shi, J. Cao, Q. Zhang, Y . Li, and L. Xu, “Edge computing: Vision and challenges,” IEEE Internet of Things Journal , vol. 3, no. 5, pp. 637–646, 2016
2016
-
[185]
On data-driven Saak transform,
C.-C. J. Kuo and Y . Chen, “On data-driven Saak transform,” Journal of Visual Communication and Image Representation , vol. 50, pp. 237– 246, 2018
2018
-
[186]
Semantic communications: Overview, open issues, and future research directions,
X. Luo, H.-H. Chen, and Q. Guo, “Semantic communications: Overview, open issues, and future research directions,” IEEE Wireless Communications, vol. 29, no. 1, pp. 210–219, 2022
2022
-
[187]
Physics for neuromorphic computing,
D. Markovi ´c, A. Mizrahi, D. Querlioz, and J. Grollier, “Physics for neuromorphic computing,” Nature Reviews Physics , vol. 2, no. 9, pp. 499–510, 2020
2020
-
[188]
Photonics for artificial intelligence and neuromorphic computing,
B. J. Shastri, A. N. Tait, T. Ferreira de Lima, W. H. Pernice, H. Bhaskaran, C. D. Wright, and P. R. Prucnal, “Photonics for artificial intelligence and neuromorphic computing,” Nature Photonics, vol. 15, no. 2, pp. 102–114, 2021
2021
-
[2021]
Available: https://www.src.org/about/decadal-plan/ decadal-plan-full-report.pdf
[Online]. Available: https://www.src.org/about/decadal-plan/ decadal-plan-full-report.pdf
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.