{"id":"d30531e6-b4e8-45eb-b724-b27912f305f7","arxiv_id":"1908.02039","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review that classifies video watermarking methods by application-driven trade-offs, embedding location, and transform domain.","lead":"This paper surveys the state of the art in digital video watermarking and proposes a framework for choosing a scheme based on the balance among invisibility, robustness, and efficiency. It organizes hundreds of cited techniques by embedding location, insertion domain, and threat model.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper promises a property-based deduction of watermarking choices but never supplies the decision procedure; 'efficiency' is defined but quantified with bitrate, not runtime.","rationale":"The reader's weakest_assumption focused on the incompleteness of the invisibility-robustness-efficiency triple, noting that capacity, blindness, and time complexity also affect the optimal choice. My concern is adjacent but more fundamental: even for the three advertised properties, the paper never defines a measurable 'efficiency' nor provides any rule for deducing location and method from the equilibrium. Section 2.2.6 conflates time complexity with BIR, a bitrate overhead metric, which means the efficiency axis is not consistently defined. Sections 4 and 5 are descriptive taxonomies without a synthesis. This supports the reader's UNVERDICTED verdict because the central claim is a non-testable promise in a review paper, but it goes beyond the reader's weakest_assumption by identifying that the missing piece is not only additional properties but the complete absence of the claimed deduction mechanism. Since the paper is a survey rather than a novel technical contribution, the appropriate verdict remains UNVERDICTED; however, the finding sharpens the reason: the survey's organizing claim is unfulfilled, not merely incomplete in its property list. The concrete test I propose would settle whether the framework can be applied at all, which is the crux of the concern.","tokens_in":28870,"tokens_out":4146,"duration_ms":46160,"concrete_test":"Take the real-time communication scenario from Section 7 and attempt to use the paper's criteria (invisibility, robustness, efficiency) to select a scheme from Sections 4 and 5. If the paper provides no way to evaluate embedding latency and no decision rule, the claimed deduction fails. Alternatively, build a property matrix from the surveyed schemes; if the 'efficiency' column cannot be filled from the paper because only BIR appears and no runtime values, the equilibrium cannot be instantiated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract, is that a designer can first define the equilibrium point among invisibility, robustness, and efficiency, then deduce from that balance the best embedding location and method. The body of the survey never actually performs or specifies this deduction. Section 2.2.6 defines time complexity as the embedding/extraction delay, but then says a common quantification is the Bit Increased Rate (BIR), which is a bitrate overhead measure, not a runtime measure. No surveyed scheme in Sections 4 or 5 reports actual embedding or extraction time, so the 'efficiency' axis cannot be populated. Sections 4 and 5 provide a taxonomy of locations (e.g., I-frames, HVS criteria, packet fields) and methods (spatial, DCT, DWT, network channels), but they contain no decision rule, scoring function, or worked example that maps property weights to a recommended scheme. The conclusion merely asserts that the developer 'can select a precise location' based on visibility and robustness, without showing how. Thus the promised deduction is unsupported: the survey catalogs properties and schemes side by side but never connects them into a procedure. The load-bearing assertion that the survey organization maps directly onto a design procedure therefore fails as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a survey of digital video watermarking, organizing the field into watermarking properties (fidelity, blindness, robustness, capacity, time complexity), applications (copyright protection, tampering identification, clandestine communication, traffic analysis, access control), threat models, embedding locations (visual-object level vs. network-flow level), and embedding methods (spatial, DCT, DST, DWT, DFT, SVD, hybrid, and other domains). The stated goal, in the abstract and introduction, is to enable a designer to first define a desired equilibrium among invisibility, robustness, and efficiency and then deduce the best embedding location and method from that balance. The paper compiles a large set of references and provides a useful taxonomy, but the promised decision procedure is never actually specified, and several property definitions contain factual errors.","tokens_in":29021,"tokens_out":5522,"duration_ms":54103,"significance":"If the claimed design framework were actually delivered, the paper would provide a valuable decision aid for practitioners choosing a video watermarking scheme. The survey's strengths include its broad compilation of roughly 170 references, a clear separation between visual-object watermarking and network-flow watermarking (Section 4.2), a systematic threat-model taxonomy (Section 3), and useful discussions of hybrid transform domains (Section 5.2.6) and scalability-aware embedding (Section 4.1.5). However, the central assertion that a property-based equilibrium can determine the best location and method is not substantiated: sections 4 and 5 present taxonomies without a decision rule or worked example, and the efficiency axis is defined inconsistently. The paper is therefore a reasonably organized survey whose main advertised contribution remains unrealized.","major_comments":[{"comment":"The text states: \"The higher the correlation is, the more distorted the signal is.\" This is reversed. The formula C(X, X̂) = cov(X, X̂) / sqrt(var(X) · var(X̂)) is a similarity measure: high correlation means the watermarked signal closely matches the original, i.e., low distortion. Because medium fidelity/invisibility is one of the three axes in the abstract's equilibrium point, this inversion corrupts the metric basis of the proposed framework and must be corrected.","section":"Section 2.2.1, Correlation Coefficient"},{"comment":"Time complexity is defined as the embedding/extraction delay, but the quantified measure, BIR = (R_X̂ − R_X)/R_X × 100, is a bit-rate overhead, not a runtime measure. The survey never reports actual embedding or extraction times for any scheme in Sections 4 or 5, so the \"efficiency\" axis of the invisibility/robustness/efficiency equilibrium cannot be instantiated. This conflation also undermines the conclusion's claim that embedding speed is overlooked, since BIR-based studies do not measure speed.","section":"Section 2.2.6, Time Complexity and BIR"},{"comment":"The abstract promises that, given an equilibrium among invisibility, robustness, and efficiency, the designer can \"deduce the best location of the information embedding as well as the method used to embed it.\" The body never performs or specifies this deduction. Sections 4 and 5 catalog locations and methods side by side but contain no decision rule, scoring function, or worked example that maps property weights to a recommended scheme. Section 7 merely asserts that the developer \"can select a precise location\" based on visibility and robustness. The load-bearing claim is therefore unsupported as stated.","section":"Abstract and Sections 4-5, 7"},{"comment":"The abstract's equilibrium triad of invisibility, robustness, and efficiency omits other properties that the paper itself identifies as application-dependent: capacity (Section 2.2.5), blindness (Section 2.2.3), and robustness level (fragile vs. semi-fragile vs. robust, Section 2.2.4). For steganography, capacity and undetectability dominate; for access control, reversibility matters; for tamper detection, fragile behavior is decisive. The promised deduction from a three-way equilibrium is therefore incomplete unless these additional properties are folded into the selection procedure.","section":"Section 2.2 and Section 2.3"},{"comment":"The title promises a \"Review of the State-Of-The-Art,\" and Section 7 makes an explicit gap claim (\"Robust watermark scheme allowing a watermark embedding in a such short time is still to be developed\"), but the paper gives no search methodology, inclusion criteria, or coverage analysis. Without this, the completeness of the survey and the real-time gap claim cannot be assessed; notable recent directions such as deep-learning-based watermarking are absent from the reviewed schemes.","section":"Title, Abstract, and Section 7"}],"minor_comments":[{"comment":"The text refers to \"the flowchart of Figure 4.1.5\"; the referenced figure is numbered Figure 6.","section":"Section 4.1.5"},{"comment":"The sentence \"the magnitude yields much more information about the spatial structure of the image\" is backwards for image processing; the phase of the Fourier coefficients carries most structural information. This should be corrected to avoid misleading readers about DFT-domain watermarking.","section":"Section 5.2.4, DFT domain"},{"comment":"Minor typo: \"the most important step of the all scheme\" should read \"of the whole scheme.\"","section":"Section 2.1"},{"comment":"References [81] and [95] appear to cite the same paper by Nakano-Miyatake and Perez-Meana in two different bibliographic forms; please consolidate.","section":"References [81] and [95]"},{"comment":"The terminology \"steganography is watermark-oriented\" followed by \"we define watermarking focused on the medium as carrier signal-oriented\" is confusing and appears to reverse the intended distinction; please clarify the definitions.","section":"Section 2.3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a survey rather than a technical contribution, and its main advertised contribution is a design framework that is not actually delivered. The factual errors in the fidelity and efficiency metrics (Section 2.2.1 and 2.2.6) and the missing decision procedure are fixable in revision, but they are load-bearing for the central claim. I recommend major revision rather than rejection because the survey's taxonomy and reference collection have genuine value. I also note that the absence of a search methodology makes the \"state-of-the-art\" completeness claim difficult to accept, and the real-time gap claim in Section 7 goes beyond what the surveyed literature supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a literature review, not a research contribution. Its value is coverage and organization, and for someone entering video watermarking it's a decent map. The abstract promises more than the body delivers: it says you can define an equilibrium among invisibility, robustness, and efficiency and then deduce the best embedding location and method. That deduction never happens. The paper catalogs locations in Section 4 and methods in Section 5 but gives no decision rule, scoring function, or worked example. The conclusion restates the idea without showing how. So the load-bearing framing is unsupported.\n\nWhat it does well: the survey is genuinely broad—applications, threat models, visual and network-level embedding, spatial and frequency methods, plus a bit on hybrid and emerging tech. The network watermarking section is a useful addition you don't see in every video-watermarking survey. The HVS-based location breakdown is clear, and the citation list gives a good starting point.\n\nSoft spots, in proportion: the efficiency angle is muddled. Section 2.2.6 defines time complexity as embedding/extraction delay but then adopts Bit Increased Rate as the common quantification. BIR is bitrate overhead, not runtime. No surveyed scheme reports actual embedding time, so the efficiency axis of the equilibrium triangle is empty. There's also a factual slip in 2.2.1: higher correlation means more similarity, hence less distortion, not more. The survey method is unspecified—no search strategy or inclusion criteria—so the \"state of the art\" claim is not reproducible. The real-time embedding gap asserted in the conclusion is an opinion, not a demonstrated absence in the literature. Minor issues: a duplicate citation and a broken figure reference here and there.\n\nNone of this is fatal for a survey. If the abstract were toned down to say \"we organize the field and identify trade-offs\" rather than \"we enable deduction,\" the paper would be honest. With those fixes, plus correcting the correlation statement and adding a paragraph on survey methodology, I'd take it as a reasonable review. As it stands, it's a useful catalog with an oversized claim.\n\nRecommendation: send it to peer review, but the referee should push for abstract/body alignment and the factual corrections. It's not a desk reject, but it needs revision.","headline":"A broad, useful survey of video watermarking whose abstract overpromises a design framework the body never delivers; fixable with honest framing and a few corrections.","tokens_in":29568,"tokens_out":2968,"would_cite":false,"duration_ms":29471,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"To pick a video watermarking scheme, first fix the trade-off among invisibility, robustness, and efficiency; the survey argues the right location and method follow from that balance.","keywords":["digital watermarking","video watermarking","state-of-the-art review","invisibility","robustness","efficiency","embedding location","threat model"],"falsifier":"Find two applications with identical weights on invisibility, robustness, and efficiency but different capacity or blindness needs, and exhibit that the optimal location/method differs; this would contradict the claim that the equilibrium point alone determines the choice. The paper's own examples, a fragile tampering watermark versus a robust copyright watermark, suggest such a pair exists.","tokens_in":28619,"feed_emoji":"🎬","tokens_out":4996,"duration_ms":46368,"temperature":0.7,"pith_summary":"This paper is a state-of-the-art review that tries to make video watermarking scheme selection systematic. The authors argue that every application can be characterized by a desired equilibrium point among invisibility, robustness, and efficiency, and that this point should drive two further decisions: the embedding location and the embedding method. They survey threat models, embedding locations, embedding methods, and combinations with other technologies, then draw out a gap: robust watermark embedding fast enough for real-time communication has not been demonstrated. If the framework holds, a designer can pick a scheme from the literature by matching the application's required trade-off instead of re-deriving trade-offs for each new use case.","feed_headline":"Three-way trade-off maps video watermarking choices","feed_subtitle":"A review that turns scheme selection into a location-and-method decision driven by a three-property equilibrium.","key_machinery":"The carrying device is a two-level taxonomy. At the property level, watermarking schemes are characterized by medium fidelity (invisibility), watermark fidelity and recognizability, blindness, robustness, capacity, and time complexity, with the first three in explicit tension. At the design level, the paper decomposes any scheme into location (videos as visual objects with HVS-based block selection, or videos as network data with physical-layer, storage-channel, timing-channel, and application-protocol embedding) and method (spatial versus frequency domain, including DCT, DST, DWT, DFT, SVD, and hybrid domains, plus motion-vector embedding). The survey's central claim is that fixing the equilibrium point on the three headline properties, together with the threat model, is enough to walk this taxonomy and reach the best location and method.","core_discovery":"On its own terms, this paper claims that the sprawling design space of video watermarking can be navigated in three decisions: first fix the desired equilibrium among invisibility (how little distortion the watermark causes), robustness (how hard it is to remove), and efficiency (embedding and extraction cost); then deduce where to embed (metadata or packet level, whole frames, or selected regions of frames chosen by human visual system criteria); then choose how to embed (spatial-domain methods such as least-significant-bit and linear masking, or frequency-domain methods such as DCT, DWT, DFT, SVD and hybrids). It organizes the surveyed literature as responses to these criteria and links each property to applications: copyright protection, tampering identification, clandestine communication, traffic analysis, and access control. The survey also finds that no robust scheme has been shown to embed within the time budget of real-time communication, which points to a concrete open problem.","pith_inferences":["The three-way equilibrium is probably only part of the decision: capacity, blindness, and reversibility can change the optimal location and method even when invisibility, robustness, and efficiency weights are fixed, so the framework is best read as a starting point rather than a complete decision procedure.","The survey's organization suggests a testable quantitative version: assemble reported PSNR, bit-error rate, and embedding-time measurements and plot surveyed schemes in the equilibrium space to see whether location/method clusters actually separate by application.","The real-time gap may be closed by embedding in the compressed domain during encoding, for example in P-frame coefficients, since raw-frame spatial methods would not fit the time budget; this is an extension the paper points to but does not test."],"forward_implications":["Application needs and threat model narrow the candidate set to a small number of location/method pairs, so a developer can select a surveyed scheme without building a custom comparison.","The taxonomy predicts that robustness-oriented applications such as copyright protection should prefer frequency-domain methods and HVS-selected blocks, while capacity-oriented steganography can accept spatial-domain embedding at lower robustness.","The identified gap implies a concrete research target: a robust watermarking scheme whose embedding completes within the frame acquisition budget, 33 ms at 30 fps and 16 ms at 60 fps.","Combining watermarking with homomorphic encryption, machine-learning detection, blockchain provenance, or quantum transforms extends the same location/method taxonomy to new carriers without changing its structure."],"supporting_citations":[{"why":"Defines the property set (fidelity, blindness, robustness, capacity, complexity) that the equilibrium framework operates on.","marker":"[18]"},{"why":"Supplies the distortion and recognizability measures used to quantify invisibility and detection accuracy.","marker":"[21]"},{"why":"Provides the four watermark lifecycles and the network-flow watermarking context that structure the survey's transmission and threat analysis.","marker":"[10]"},{"why":"Supplies the standard attack-simulation benchmark used to measure robustness.","marker":"[37, 38]"},{"why":"Introduces GOP, I/P/B-frame and inter-prediction concepts that drive the video-as-visual-object embedding-location discussion.","marker":"[92]"},{"why":"Defines SVC scalability, which the survey uses to characterize scalability-resistant watermarking.","marker":"[96]"},{"why":"Provides the TCP/IP protocol model underlying the four network-watermarking location families: physical, storage, timing, and application protocol.","marker":"[100]"},{"why":"Surveys network flow watermarking and supports the packet-level view of video watermarking.","marker":"[107]"}],"fun_headline_variants":["Trade-off triangle guides video watermarking choices","Invisibility, robustness, efficiency: the watermark triangle","Watermarking video: trade-off picks location and method","Real-time robust watermarking remains an open problem"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole deduction rests on the premise that the desired balance among invisibility, robustness, and efficiency is the main determinant of the best embedding location and method; if capacity, blindness, or reversibility dominates for a given application, the promised deduction is incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Trade-off triangle guides video watermarking choices","Invisibility, robustness, efficiency: the watermark triangle","Watermarking video: trade-off picks location and method","Real-time robust watermarking remains an open problem"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000958,"raw_usage":{"total_tokens":4078,"prompt_tokens":936,"completion_tokens":3142,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":3081}},"tokens_in":552,"tokens_out":3142,"duration_ms":27171,"temperature":1.0,"reasoning_tokens":3081,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:55:18.130009+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find two applications with identical weights on invisibility, robustness, and efficiency but different capacity or blindness needs, and exhibit that the optimal location/method differs; this would contradict the claim that the equilibrium point alone determines the choice. The paper's own examples, a fragile tampering watermark versus a robust copyright watermark, suggest such a pair exists.","supporting_citations":[{"cited_title":"Overview of the scalable video coding extension of the h","cited_arxiv_id":null,"evidence_quote":"Defines SVC scalability, which the survey uses to characterize scalability-resistant watermarking."},{"cited_title":"TCP/IP protocol suite","cited_arxiv_id":null,"evidence_quote":"Provides the TCP/IP protocol model underlying the four network-watermarking location families: physical, storage, timing, and application protocol."},{"cited_title":"Network ﬂow watermarking: A survey","cited_arxiv_id":null,"evidence_quote":"Surveys network flow watermarking and supports the packet-level view of video watermarking."}],"review_version":1}