{"id":"9e93547e-84ab-4f41-88d3-7a47f0cfdfa6","arxiv_id":"2502.08830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of 33 APT campaigns shows command-and-control traffic overwhelmingly uses HTTP(S) and DNS, with over half of campaigns splitting traffic across multiple servers to evade volume-based detection.","lead":"This paper surveys 33 Advanced Persistent Threat campaigns over 22 years and identifies the network protocols and evasion techniques they use to avoid detection, finding heavy reliance on HTTP(S), DNS, and multi-server fallback channels. It is a reference for security researchers and defenders who need a structured picture of how APT command-and-control traffic hides in plain sight.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection and report-completeness bias in the 33-campaign sample undermines the headline percentages, but the qualitative TTP findings survive; one targeted check would settle the matter.","rationale":"The reader identified the same load-bearing assumption I did: the 33-campaign selection and the assumed completeness of 118 vendor reports. I agree with that as the central concern, and my proposed test is a sensitivity analysis that can be run entirely from the paper's own cited sources, without new data collection. I did not find a deeper internal inconsistency that would invalidate the qualitative claims; the Skula/Sakula naming, the 'xxx domains' placeholder, the absent Stuxnet entry, and the abstract/body HTTPS discrepancy are real but editorial, and the reader already flagged them. The core reasoning of the paper — APTs rely on C&C, C&C is mostly addressed via DNS, and the dominant payload protocols are HTTP(S) with DNS — is well supported by the cited literature and by the paper's own traffic plots in Figures 8–13. The single most load-bearing concern is therefore not the existence of those techniques but the generalizability of the specific percentages that the abstract elevates to headline findings. If the sensitivity test shows stability, the CONDITIONAL verdict could stand as ACCEPT; if it shows instability, the percentages should be reframed as descriptive statistics for a non-random sample. Either way the verdict stays CONDITIONAL until the authors supply the campaign-level data behind Tables I–III, which the paper currently does not provide.","tokens_in":32119,"tokens_out":1659,"duration_ms":15572,"concrete_test":"Rebuild Tables I–III from the 118 cited reports but restricted to the campaigns that the paper itself lists as selected on availability grounds, then recompute Figure 14 percentages under two perturbations: (1) drop the five least-documented campaigns; (2) add the five campaigns that MITRE ATT&CK lists as most active but that the paper excluded. If 81% HTTPS, 60.6% fallback channels, or 54.5% multi-level encryption move by more than 10 percentage points in either perturbation, the headline numbers are sample-dependent and the paper should report them as descriptive of the selected campaigns, not of APTs generally.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claims — 81% HTTPS, 45% DNS, 60.6% fallback channels, 54.5% multi-level encryption, and all Figure 14 percentages — are computed over 33 campaigns selected by 'fair distribution' plus 'availability of data from industry vendors' plus 'the one that causes major damage' (Section III). That is a non-random, damage-maximizing sample drawn from a population (MITRE ATT&CK lists over 100 groups). If the selection correlates with protocol usage or TTP sophistication — e.g., major-damage campaigns are disproportionately the ones with large infrastructures and fallback channels — then every headline percentage is an upper-biased estimate, not a population statistic. The paper also treats a TTP as absent when no vendor report mentions it, which is an assumption of report completeness that is not stated or defended anywhere in Sections III or V. This matters because the abstract frames the numbers as facts about 'APTs' generally ('the most popular protocol... with 81% of APT campaigns'), and the reader's verdict already conditions on this. The paper's own text concedes the underlying data asymmetry: 'the primary source of information comes from industry, due to the relative monopoly on information related to APT campaigns' (Section III). That is not a flaw by itself, but it makes the percentages a statement about what vendors document for high-impact campaigns, not about the APT population. The qualitative taxonomy (Figure 5) and the traffic-based TTP analyses in Sections VII–IX do not depend on the percentages and remain valuable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys 33 APT campaigns documented in 118 industry reports, supplemented by the authors' own traffic datasets, with the goal of characterizing network-based tactics, techniques, and procedures (TTPs) for command-and-control (C&C) communication. It presents quantitative claims about protocol usage (81% HTTPS, 45% DNS), channel-based evasion (60.6% fallback channels, 54.5% multi-level encryption), and DNS-based TTPs (27.2% dynamic DNS, 24.2% DGA/DNS tunneling), and it develops a taxonomy of 13 network-based TTPs. The paper concludes with recommendations for next-generation NIDS, emphasizing HTTP(S) and DNS inspection and the need to treat fallback channels and multi-hop proxies as first-class detection signals.","tokens_in":32304,"tokens_out":4333,"duration_ms":42792,"significance":"The qualitative contribution is valuable: the TTP taxonomy (Figure 5), the per-campaign tables, and the real-traffic examples of domain fronting, protocol impersonation, fallback channels, and raw-TCP/ICMP channels give concrete, actionable evidence for detection research. The paper also draws on a large corpus of industry reports and the authors' own datasets, and several figures (Figures 8–13) provide traffic-level grounding that is often missing in survey papers. If the quantitative claims were properly scoped, they would provide useful evidence for prioritizing HTTP(S) and DNS inspection. However, the headline percentages are load-bearing and currently rest on an unstated representativeness assumption, and at least one internal inconsistency affects the paper's central protocol claim.","major_comments":[{"comment":"The selection of the 33 campaigns is non-random: the methodology states that campaigns are chosen for 'fair distribution' over 22 years, 'availability of data from industry vendors,' and 'the one that causes major damage' (Section III). This is a damage-maximizing, availability-biased sample, yet the abstract and Section X present the resulting percentages (81% HTTPS, 45% DNS, 60.6% fallback channels, 54.5% multi-level encryption) as facts about APTs generally. The authors should either reframe these numbers as descriptive statistics of the selected sample with an explicit limitations paragraph, or justify representativeness and provide uncertainty estimates. As written, the central quantitative claims overgeneralize.","section":"Section III"},{"comment":"The abstract states that 'the most popular protocol to deploy evasion techniques is using HTTP(S) with 81% of APT campaigns,' but Section X and Figure 14.b state that all campaigns use HTTP and that 81% rely on HTTPS to bypass NIDS. These statements conflict: if HTTP is used by all campaigns, then the 81% figure cannot refer to HTTP(S) as a combined category. This internal inconsistency affects the paper's most prominent protocol claim and must be corrected.","section":"Abstract vs. Section X"},{"comment":"The analysis treats a TTP as absent when no vendor report mentions it, effectively assuming that the 118 reports are complete descriptions of each campaign. This assumption is never stated or defended; the paper acknowledges only that industry is the primary information source. Because the abstract frames the numbers as facts about APT campaigns, the authors should explicitly state that the percentages describe what is documented in the selected reports, not necessarily what occurred, and discuss the implications of report incompleteness.","section":"Sections V and X"},{"comment":"The percentages in Figure 14 are said to be based on the 33 campaigns in Table III, but the paper does not provide the per-campaign counts that map Table III's checkmarks to the reported figures. A reader cannot verify 60.6%, 54.5%, 51.5%, 27.2%, or 24.2% from the material as presented. Please include the underlying counts (e.g., a supplementary table listing, for each campaign, which TTPs were counted) or a machine-readable version of Table III.","section":"Figure 14 and Table III"}],"minor_comments":[{"comment":"In the HEALP dataset description, the text reads 'we collect xxx domains'; the placeholder 'xxx' should be replaced with the actual number of domains collected, as this is part of the dataset description needed for reproducibility.","section":"Section III"},{"comment":"The caption describes 'Stealthy malicious APT StringPity' while the surrounding text refers to njRAT; the mismatch should be resolved.","section":"Figure 11"},{"comment":"The caption uses 'GRIFFON' while the text and Section II use 'Griffon'; please standardize the spelling.","section":"Figure 13"},{"comment":"The sentence 'Figure 14.b shows that all APT campaigns have continued using HTTP since 2001' should be checked against Table III, since some rows appear to have no HTTP checkmark; if those rows do use HTTP, the claim should be clarified to avoid overstatement.","section":"Section X"},{"comment":"The dataset references [4], [25], and [127] point to prior work and a GitHub repository, but the paper does not clearly state how the reader can access the APTracePlus and MCFP datasets used for the traffic figures; consider adding a data-availability note.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is best viewed as a field survey with a useful qualitative taxonomy and real traffic examples, rather than a novel measurement study with rigorous statistical inference. The main weakness is the gap between the ambitious general claims in the abstract and the non-representative sample used to compute the percentages. This is fixable by reframing the scope, adding a limitations section, and correcting the HTTP(S)/HTTPS inconsistency. I would not reject the paper, but the authors need to address the sampling and completeness issues before publication. The repeated citation of the authors' own prior work is reasonable given that the datasets are their own, but the paper should make the dataset artifacts more accessible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a useful consolidation of APT network TTPs, built from 118 industry reports and 33 campaigns, with a clean taxonomy and some genuinely illustrative real-traffic examples. The individual techniques are not new—DGA, DNS tunneling, domain fronting, fallback channels are all in MITRE ATT&CK and in Lemay et al. and Ussath et al. But the cross-campaign aggregation over two decades and the traffic-based walkthroughs (GRIFFON splitting over four C&C servers, Mivast/Sakula faking TLS, njRAT staying under a volume threshold) give NIDS designers a concrete evidence base for what to prioritize. I'd want this on my shelf as a reference.\n\nNow the soft spots. The headline numbers—81% HTTPS, 45% DNS, 60.6% fallback channels—are computed over a non-random sample. Section III says campaigns were chosen by 'fair distribution,' vendor-report availability, and 'the one that causes major damage.' That is a convenience sample tilted toward high-visibility, well-documented operations. On top of that, the count treats any technique not mentioned in a vendor report as absent. So the percentages are statements about what vendors document for impactful campaigns, not robust population statistics. The abstract overreaches by presenting them as facts about 'APTs' generally. The qualitative taxonomy in Figure 5 and the traffic examples do not rest on those percentages; they survive. This is the main fix the paper needs: reframe the numbers as sample statistics and discuss the sampling and completeness biases, or do a sensitivity analysis.\n\nPolish is also not there yet. Section III has a literal 'xxx domains' placeholder. The conclusion mentions Stuxnet and Taidoor, neither of which is in the campaign tables. The malware is called Skula in the introduction and Sakula in Section VIII. The abstract says HTTP(S) 81% while the body says HTTPS 81%. None of these are deep problems, but they add up to a manuscript that needs a careful revision pass.\n\nBottom line: this deserves peer review, not desk rejection. A serious referee can push the authors to fix the methodology framing and the consistency issues. The final version would be a solid reference for threat-intel and NIDS researchers. I'd bring it to reading group to talk about sample bias in threat-intel synthesis.","headline":"A useful, messy consolidation of APT C&C TTPs whose headline percentages rest on a convenience sample; worth refereeing, but the numbers need re-framing.","tokens_in":32956,"tokens_out":3343,"would_cite":true,"duration_ms":32792,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"APTs cannot operate without C&C servers, and HTTP(S) is the evasion channel in 81% of campaigns.","keywords":["Advanced persistent threats","command and control","evasion techniques","DNS tunneling","HTTPS C&C","fallback channels","network intrusion detection","TTP analysis"],"falsifier":"Compile a fresh sample of every APT campaign disclosed by several major vendors between 2023 and 2025, code the same TTP features without excluding any campaign, and compare the HTTPS share (81%), DNS share (45%), fallback-channel share (60.6%), and multi-hop proxy share (51.5%); a substantial shortfall would show that the paper's percentages are artifacts of its campaign-selection rule.","tokens_in":31819,"feed_emoji":"🌐","tokens_out":8556,"duration_ms":72594,"temperature":0.7,"pith_summary":"This paper tries to establish what advanced persistent threats actually do at the network layer, from the adversary's point of view, by reading 22 years of industry reports and real traffic artifacts. The central claim is that APT operations are impossible without command-and-control (C&C) servers, that those servers are mostly reached through DNS, and that HTTP(S) is the dominant vehicle for evading detection, used by 81% of campaigns. A sympathetic reader should care because the percentages point at specific design priorities for next-generation network intrusion detection systems (NIDS): inspect HTTP(S) and DNS in context, treat fallback channels and multi-hop proxies as first-class signals, and stop relying on traffic volume alone.","feed_headline":"81% of APT campaigns hide command traffic in HTTPS","feed_subtitle":"A 22-year study of 33 campaigns says DNS and fallback channels are the other must-watch signals","key_machinery":"The carrying object is a network-level TTP (tactics, techniques, and procedures) taxonomy with three branches: DNS-based techniques (DGA, dynamic DNS, FQDN hijacking, DNS tunneling), traffic-based techniques (web-protocol misuse, data obfuscation through protocol impersonation, non-application protocols such as raw TCP or ICMP, stealthy bursts, low profile), and channel-based techniques (multi-layer encrypted channels and fallback channels). The taxonomy is filled in by reading the vendor reports into per-campaign feature tables and then validating selected techniques against collected traffic and domain datasets, so the percentage claims come from campaign-level binary features rather than from individual packet labels.","core_discovery":"The paper claims that no APT can function without C&C infrastructure, and that infrastructure is reachable mostly through DNS-addressed servers. Across 33 campaigns drawn from 118 industry technical reports spanning 22 years, HTTP(S) is the workhorse evasion protocol: 81% of campaigns use HTTPS, essentially all use HTTP at some point, and 45% use DNS for resolution or tunneling. The defensive consequence is that fallback channels are not an edge case: 60.6% of campaigns split traffic across multiple C&C servers to defeat volume-based detection, 54.5% layer encryption so that decrypting one layer still leaves another, and 51.5% use domain fronting or multi-hop proxies. The paper also reports that 84.8% of campaigns use backdoors and 78.7% use remote-access trojans, with botnets rare, and it frames these findings as requirements for the next generation of network-based APT detection.","pith_inferences":["The 81% and 60.6% figures are best read as lower bounds: a campaign counted as not using a technique may simply be one whose vendor report omitted it, so true prevalence could be higher.","The selection rule of choosing one high-damage campaign per timeframe skews the sample toward well-resourced, well-documented actors, so the trend that fallback channels are increasing may be stronger for prominent campaigns than for the general APT population.","A direct test of the paper's generalization would be to repeat the same campaign-level feature extraction on a sample built from all reports released in a fixed later window, without impact-based preselection, and check whether the HTTPS and fallback-channel shares are reproduced.","The taxonomy could be extended to protocols the paper flags but does not fully count, such as DNS over HTTPS, cloud APIs, and P2P SMB, where the same evasion logic is likely migrating."],"forward_implications":["A network intrusion detection system that does not treat HTTP(S) payloads and DNS queries as primary detection surfaces will miss the two channels that carry most APT C&C traffic.","Any volume-based detector can be evaded by the 60.6% of campaigns that divide traffic across multiple C&C servers, so detector features should be per-flow or per-context rather than aggregate-volume.","Decrypting TLS at the perimeter is insufficient when 54.5% of campaigns wrap command traffic in additional encoding or cipher layers underneath the session.","Detection of malicious domains has to cover dynamic DNS, typosquatting, TLD squatting, FQDN hijacking, and DGA, because 45% of campaigns use DNS beyond simple resolution.","Because 24.2% of campaigns deliver malware through spearphishing links, domain and URL reputation belongs inside the network-detection loop, not only at the mail gateway."],"supporting_citations":[{"why":"Supplies the detailed account of dynamic DNS, hijacked FQDNs, HTTPS C&C, and large C&C server counts that anchors several of the paper's measurements.","marker":"[3]"},{"why":"Provides the HTTP(S) traffic dataset from which the paper validates traffic-based techniques such as protocol impersonation and stealthy bursts.","marker":"[4]"},{"why":"Establishes the method of investigating APTs through public technical reports and supplies the campaign list that inspired the paper's selection.","marker":"[6]"},{"why":"Provides the adversary technique catalog whose stages shape the paper's network-level TTP taxonomy.","marker":"[9]"},{"why":"Provides the collected APT, phishing, and legitimate domain dataset used for the DNS-based TTP analysis.","marker":"[25]"},{"why":"Supplies the seven-stage lifecycle model whose C&C stage the paper singles out for subdivision into DNS resolution, server location, and malicious traffic.","marker":"[26]"},{"why":"Defines domain fronting, the evasion technique the paper measures at 51.5% of campaigns.","marker":"[117]"}],"fun_headline_variants":["HTTPS masks 81% of APT command traffic","APT evasions rely on HTTPS and DNS, study finds","Split C&C servers help 60% of APTs evade detection","No APT survives without C&C; HTTPS is the favorite mask","22-year APT study: HTTP(S) and DNS are evasion cornerstones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline percentages rest on the assumption that the 33 selected campaigns and the 118 vendor reports used to code them are representative of all APTs, and that techniques missing from a report were not actually used.","fun_headline_variants_meta":{"raw":{"variants":["HTTPS masks 81% of APT command traffic","APT evasions rely on HTTPS and DNS, study finds","Split C&C servers help 60% of APTs evade detection","No APT survives without C&C; HTTPS is the favorite mask","22-year APT study: HTTP(S) and DNS are evasion cornerstones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000375,"raw_usage":{"total_tokens":2038,"prompt_tokens":1023,"completion_tokens":1015,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":925}},"tokens_in":639,"tokens_out":1015,"duration_ms":10401,"temperature":1.0,"reasoning_tokens":925,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:32:10.969231+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile a fresh sample of every APT campaign disclosed by several major vendors between 2023 and 2025, code the same TTP features without excluding any campaign, and compare the HTTPS share (81%), DNS share (45%), fallback-channel share (60.6%), and multi-hop proxy share (51.5%); a substantial shortfall would show that the paper's percentages are artifacts of its campaign-selection rule.","supporting_citations":[{"cited_title":"Blocking- resistant communication through domain fronting.,","cited_arxiv_id":null,"evidence_quote":"Defines domain fronting, the evasion technique the paper measures at 51.5% of campaigns."}],"review_version":1}