{"id":"67d4a6e2-d9ea-41ab-9a78-507ab7710d5c","arxiv_id":"2411.14135","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of compact visual data representation for green multimedia, structured around human visual system principles.","lead":"This paper surveys ways to make video and image storage, transmission, and analysis use less energy, inspired by how the human visual system compresses information. It organizes recent work on perceptual coding, feature compression, and unified machine-and-human representation into a roadmap for green multimedia.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The motivating HVS-versus-VVC compression gap (100,000x vs 1,000x) is cited without defining what 'compression ratio' means for a biological system, so the paper's central quantitative foundation is unestablished.","rationale":"I agree with the reader that the weakest point is the HVS-versus-VVC comparison. The survey itself is a competent organization of the field, and the Digital Retina, VCM, and layered coding work give the research program independent value regardless of the headline number. The concern is not that the paper's recommendations are wrong; it is that the quantitative anchor used to motivate them is not operationalized, so the claimed two-orders-of-magnitude gap is currently unsupported. The Section IV statement that the gap is 'impossible to bridge' because reconstruction and understanding have different objectives is a further overclaim, since the paper's own surveyed approaches are designed to combine feature and texture coding. These issues justify a conditional verdict: the survey should clarify or soften the quantitative framing, but the body of work reviewed still supports the general direction. Because the reader already arrived at CONDITIONAL, my read does not change the verdict.","tokens_in":26001,"tokens_out":6207,"duration_ms":59442,"concrete_test":"Open the cited SRC Decadal Plan [15], locate the exact sentence behind '100,000 times,' and check whether it provides an input/output bit model for the HVS and a task definition. If the source presents the figure as an illustrative estimate rather than an operational compression ratio, the paper should replace the opening disparity with a qualitative statement; as a second check, encode a standard 1080p or 4K raw sequence with VTM at one typical bitrate and compare the resulting bits-per-pixel ratio against the claimed 1,000x figure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative premise, repeated in the Abstract, Section I, Fig. 1, and Section IV, is that the HVS compresses visual information by about 100,000x while VVC achieves roughly 1,000x. The 100,000x figure is attributed to the SRC Decadal Plan [15], but the paper never defines the numerator or denominator for a biological system. There is no stated bit model for the HVS input (photoreceptor activation rates? retinal ganglion cell spike rates? bits of 'useful concept'?) or output (conscious percept? task accuracy? memory?), so the ratio is not comparable to a codec's lossy pixel-reconstruction ratio at a specified rate-distortion operating point. The VVC side is likewise used without a bitrate or quality anchor. Because the 'notable disparity' motivates the entire green multimedia research program and is reused in Section IV to argue that the reconstruction objective makes the gap 'impossible to bridge,' the central argument rests on an undefined benchmark. The 'impossible' claim also goes beyond what the numbers establish: different objectives do not by themselves imply an unbridgeable rate gap, and the survey's own examples (generative compression, layered feature/texture coding, VCM) are attempts to serve both objectives at once. This is not an internal contradiction, but an underspecified empirical premise that needs either a definition or a downgrade to qualitative motivation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys compact visual data representation for green multimedia, framed by the claim that the human visual system (HVS) compresses visual information by roughly 100,000 times while the VVC video coding standard achieves about 1,000 times. The survey is organized into three areas: compact video compression (standards, end-to-end coding, perceptual coding, external-data/generative compression), compact feature compression (handcrafted and deep descriptors, CDVS/CDVA/VCM standardization), and unified representation for dynamic tasks (layered feature/texture coding, scalable coding). It closes with connections to AIGC, large vision-language models, knowledge-centric networking, edge computing, and future directions such as AI-agent communication and neuromorphic computing. The intended contribution is a research roadmap for HVS-inspired, task-oriented, energy-efficient visual representation.","tokens_in":26295,"tokens_out":4763,"duration_ms":46672,"significance":"If its motivating comparison were rigorously defined, this would be a timely and useful synthesis of an emerging area. The survey is broad and current, covering recent codec standards, learned compression, VCM, and joint feature/texture coding, and it explicitly connects these developments to sustainability. The taxonomy and the roadmap in Fig. 2 are reasonable organizational contributions. The paper does not present new experiments or derivations, and its value rests on the adequacy of its selected evidence and framing. The central quantitative motivation, however, is not operationalized, and several forward-looking claims in Section IV are asserted without support, which currently limits the paper's ability to justify the proposed research program over an alternative, qualitative framing.","major_comments":[{"comment":"The claim that the HVS compresses visual information 'around 100,000 times' while VVC achieves 'around 1,000 times' is cited to [15] without defining the compression ratio on either side. For a biological system, the numerator and denominator are unspecified (photoreceptor or retinal output bits versus conscious percept? bits of task-relevant content versus raw input?), and for VVC no bitrate or quality operating point is given. These are not commensurable quantities, so the 'notable disparity' and the green-multimedia research program built on it rest on an undefined benchmark. Please either provide an explicit bit model and operating points for both sides, or reframe the HVS comparison as qualitative biological inspiration rather than a quantitative compression gap.","section":"Section I and Section IV (also Abstract, Fig. 1)"},{"comment":"The statement that the reconstruction objective and the understanding objective 'ultimately make it impossible to bridge the performance gap' goes beyond what the preceding numbers establish. Different objectives do not by themselves imply an unbridgeable rate gap; many systems surveyed in Section III (layered feature/texture coding, VCM, generative compression, Digital Retina) are explicit attempts to serve both objectives simultaneously. Please soften this claim to reflect that the gap is not directly comparable or not yet quantified, rather than impossible.","section":"Section IV, first paragraph"}],"minor_comments":[{"comment":"The sentence 'Central to this efficiency is the HVS's ability to operate at extremely low bitrate representations, particularly from the primary visual cortex (V1) to extrastriate cortical areas [21]' is supported only by a 1956 study on the speed of visual perception; this reference does not appear to establish a quantitative bitrate claim, so please clarify the basis or replace the citation.","section":"Section II, paragraph 2"},{"comment":"The phrase 'their theoretical limits of compression efficiency [8] is being constantly approached' invokes Shannon's work without explaining how the lossless source-coding bound applies to lossy perceptual video coding; please qualify or cite a specific analysis for video.","section":"Section III-A1, paragraph 1"},{"comment":"The assertion that AI-agent communication is 'without any doubt' greener than human-centric communication is unsupported: no energy model, bitrate comparison, or lifecycle assessment is provided, and the computational cost of semantic communication is not considered.","section":"Section IV, AI-agent communication paragraph"},{"comment":"The sentence 'the reduction of decoding complexity could vastly decrease energy savings' appears to be a typo; it likely means 'increase energy savings' or 'decrease energy consumption.'","section":"Section I, Green Metadata sentence"},{"comment":"There are several wording issues that should be corrected in a revision: 'envelopes' should be 'encompasses' (Section II), 'researches' should be 'research directions' (Section III-A3), and 'sheds on how' should be 'sheds light on how' (Section III-B).","section":"Various"}],"recommendation":"major_revision","confidential_remarks":"The paper is closely aligned with the authors' own research program (Digital Retina, joint feature/texture coding, compact feature compression), which is natural for a survey in a niche area, but the manuscript would benefit from an independent critical assessment of those lines of work. The 100,000x HVS figure is taken from an industry decadal plan [15] rather than a peer-reviewed measurement; given that the figure is load-bearing for the paper's framing, the editor may wish to ensure that the revision addresses this provenance explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a competent survey and the taxonomy holds up, but the motivating number is a problem.\n\nWhat it does well: it organizes a sprawling area into three coherent tracks—signal-level perceptual coding, compact feature coding for machines, and unified/scalable coding for both—and it covers the relevant standards (VVC, AV1, AVS3, CDVS, CDVA, VCM) plus recent learning-based work. The roadmap figure is genuinely useful for someone entering the field. The authors also connect the work to AIGC, LVLMs, edge computing, and neuromorphic ideas, which makes it a useful entry point.\n\nThe soft spot is the central motivation. The abstract and intro claim the HVS compresses visual information by ~100,000x while VVC achieves ~1,000x, cited to an industry decadal plan, without defining what 'compression ratio' means for a biological system. Bits per conscious percept? Task accuracy? Photoreceptor vs cortical spike rates? The number is not comparable to a codec's rate-distortion operating point, and Section IV leans on it to say the gap is 'impossible to bridge'—a claim that goes beyond the evidence and sits awkwardly with the survey's own examples of layered feature/texture coding and VCM that are explicitly trying to serve both objectives at once. There's also a minor overstatement in Section IV that AI-agent communication is 'without any doubt' greener than human-centric communication.\n\nNone of this sinks the survey. The taxonomy and coverage are valuable regardless of the HVS benchmark, and the authors could fix the issue by either defining the ratio operationally or presenting it as qualitative inspiration. The citation pattern looks fine; self-citations are natural in this niche.\n\nI'd send it to peer review. A careful referee can ask for the benchmark to be clarified and Section IV tempered, and the result would be a solid reference survey. I'd cite it if I were writing in this area.","headline":"A solid survey with a shaky motivating statistic: the HVS-vs-VVC compression ratio is undefined and the 'impossible gap' claim goes beyond the evidence.","tokens_in":26753,"tokens_out":2420,"would_cite":true,"duration_ms":23021,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey argues that the human visual system's roughly 100,000-fold compression, against VVC's roughly 1,000-fold, should push video coding toward compact, task-ready representations rather than pixel-perfect reconstruction.","keywords":["green multimedia","human visual system","compact visual representation","video coding for machines","perceptual coding","feature compression","collaborative intelligence","semantic communication"],"falsifier":"Measure the bitrate required to achieve a fixed level of a concrete task—for example, object detection on a standard benchmark at a target mean average precision—using a compact feature codec on one side and a VVC-reconstructed pixel codec on the other. If the resulting compression ratio relative to raw video is on the order of 1,000 rather than 100,000, the paper's motivating disparity is not supported by direct measurement.","tokens_in":25833,"feed_emoji":"🧠","tokens_out":2840,"duration_ms":28246,"temperature":0.7,"pith_summary":"This paper is a survey with a central thesis: the human visual system compresses visual information by a factor of about 100,000, while the state-of-the-art VVC video codec compresses raw visual data by only about 1,000 times. The authors argue this gap exists because the brain extracts knowledge from visual input, whereas digital codecs aim to reconstruct the signal itself. They therefore advocate for a research program in which visual data are represented compactly for downstream tasks—machines, perception, or both—rather than for human viewing alone. The survey organizes current work into three directions: perceptual coding, compact feature representation, and unified or layered representations that serve both human and machine vision.","feed_headline":"Vision compresses 100,000x; today's codecs only 1,000x","feed_subtitle":"A survey argues green multimedia should transmit task-ready features, not reconstruct pixels.","key_machinery":"The organizing machinery of the paper is the HVS-inspired tripartite framework that maps biological mechanisms onto coding architectures: perceptual coding based on the eye's nonuniform sensitivity and memory, compact feature representation based on task-driven selective attention, and collaborative or scalable coding based on hierarchical cortical processing. The most concrete anchor is the 'Digital Retina' architecture, which uses a dual representation of compact textures and compact semantic features, allowing receivers to use feature streams directly for analytics. These principles are used to classify the surveyed methods and to structure the paper's roadmap for future green multimedia technologies.","core_discovery":"The paper's central claim is that the purpose of visual data representation should shift from signal reconstruction to knowledge extraction, motivated by the enormous efficiency of the human visual system. According to the authors, the HVS compresses visual information around 100,000 times while achieving high generalization and energy efficiency, whereas VVC achieves a compression ratio of only about 1,000 times for raw visual data. This disparity, they argue, makes it impossible to close the gap by improving codec efficiency alone, because the objectives differ: codecs reconstruct, brains understand. The survey accordingly maps a landscape of techniques that represent visual data compactly for machines—compact feature coding, end-to-end learned coding, generative and external-data-based compression, and scalable layered bitstreams that blend feature and texture coding—and positions these as the route to greener multimedia.","pith_inferences":["A testable extension of the survey's thesis: define compression ratio for a biological system by the bitrate needed to preserve task accuracy (for example, object detection at fixed mean average precision) and compare a compact feature codec against VVC-reconstructed pixels; the observed gap may be far smaller than 100,000x, which would weaken the quantitative foundation of the green-multimedia ar","The survey implies but does not state that the comparison between HVS and VVC is not apples-to-apples: one is a task-accuracy measure and the other a signal-fidelity measure, so the 100,000x figure may conflate 'useful concepts' with 'bits'.","A natural next step, beyond the paper, is to extend the HVS-inspired framework to AI-agent communication, where semantic communication among machines could achieve even higher compactness than human-oriented coding.","The paper's own roadmap suggests that large vision-language models could become the ultimate consumers of compact visual data, which would make the design of bitstream syntax for semantic prompts a future standardization problem."],"forward_implications":["If the survey's thesis is correct, future video coding standards should prioritize machine-consumable feature streams over pixel fidelity, potentially making Video Coding for Machines a mainstream direction.","Layered or scalable bitstreams that carry a base feature layer for analytics and an enhancement layer for human viewing would become the default architecture, reducing bandwidth and decode energy for task-driven applications.","End-to-end learned codecs and generative compression, which already show large bitrate savings at low rates, would be recognized not as niche tools but as core green multimedia technologies.","Perceptual coding and just-noticeable-difference modeling would be applied more aggressively to save energy, not only to improve subjective quality.","The connection to knowledge-centric networking and edge computing suggests that compact representation will be designed jointly with network and model updates, yielding systems that transmit only the information needed for a task."],"supporting_citations":[{"why":"Supplies the central benchmark that the HVS compresses visual information around 100,000 times, which anchors the entire motivation.","marker":"[15]"},{"why":"Defines the VVC standard, the 1,000x compression-ratio baseline that the paper contrasts against the HVS.","marker":"[5]"},{"why":"Introduces Video Coding for Machines, the paradigm of collaborative compression and intelligent analytics that the survey advances as a core direction.","marker":"[17]"},{"why":"Presents the Digital Retina framework, the central architecture exemplifying dual signal- and semantic-based compact representation.","marker":"[9]"},{"why":"Provides the first comprehensive end-to-end deep video compression framework, a key pillar of the compact-representation research program.","marker":"[76]"},{"why":"Establishes deep contextual video coding, a learned approach that replaces residual coding with conditional coding and achieves state-of-the-art efficiency.","marker":"[79]"},{"why":"Demonstrates conceptual compression of images, an ultra-low-bitrate approach that the paper groups under external-data-based and generative coding.","marker":"[118]"},{"why":"Formulates Collaborative Intelligence, the scheme for splitting computation between edge devices and the cloud that underlies compact intermediate-feature coding.","marker":"[154]"}],"fun_headline_variants":["Brains compress 100,000x; codecs only 1,000x","Forget pixel-perfect: send task-ready features","Green video: brains do 100,000x compression, codecs 1,000x","Why transmit pixels? Just send the meaning","HVS-inspired coding: 100x more compression than VVC"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument rests on the meaning of the claim that the human visual system compresses visual information about 100,000 times; if that number is not defined in terms of a comparable measure (such as bits per useful concept or task-accuracy-preserving bitrate), the gap between HVS and VVC loses its quantitative force.","fun_headline_variants_meta":{"raw":{"variants":["Brains compress 100,000x; codecs only 1,000x","Forget pixel-perfect: send task-ready features","Green video: brains do 100,000x compression, codecs 1,000x","Why transmit pixels? Just send the meaning","HVS-inspired coding: 100x more compression than VVC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000558,"raw_usage":{"total_tokens":2608,"prompt_tokens":855,"completion_tokens":1753,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":1660}},"tokens_in":471,"tokens_out":1753,"duration_ms":14358,"temperature":1.0,"reasoning_tokens":1660,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:29:30.350728+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the bitrate required to achieve a fixed level of a concrete task—for example, object detection on a standard benchmark at a target mean average precision—using a compact feature codec on one side and a VVC-reconstructed pixel codec on the other. If the resulting compression ratio relative to raw video is on the order of 1,000 rather than 100,000, the paper's motivating disparity is not supported by direct measurement.","supporting_citations":[{"cited_title":"Collaborative intelligence: Challenges and opportunities,","cited_arxiv_id":null,"evidence_quote":"Formulates Collaborative Intelligence, the scheme for splitting computation between edge devices and the cloud that underlies compact intermediate-feature coding."}],"review_version":1}