REVIEW 4 major objections 6 minor 58 references
AI-Powered Spearphishing Cyber Attacks: Fact or Fiction?
T0 review · 4 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Most people, even when told fakes may be present, misjudged 66% of AI-generated audio clips and 43% of AI-generated video clips, and the fakes were built on an ordinary desktop.
desk verdict The paper's headline 66%/43% detection-failure claims are not supported by the reported analysis — those are overall error rates over mixed real/fake sets, not missed-fake rates — but the study is a genuine attempt worth a careful revision rather than a desk reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a detection experiment built around deepfakes—synthetic audio and video that replace a person's likeness or voice—and by the toolchain that produced them. In the questionnaire, the video condition had 8 clips (3 fake, 5 genuine), the audio condition had 13 clips (3 fake, 10 genuine), and the combined condition had 8 lip-synced clips; participants were told that fakes were present and were asked to identify the generated items and give their reasons. The generated media themselves—a face swap, a cloned voice, and a lip-synced combination—are the objects whose realism the experiment measures.
What would settle it
Recompute the audio and video error rates separately for fake items and for genuine items from the raw responses; if most incorrect audio selections were participants labelling genuine clips as fake rather than missing fake clips, the headline claim that 66% of people failed to identify AI-created audio is not supported.
Extended reading notes
Core claim
The paper's central claim is that synthetic audio and video are already convincing enough to support spearphishing attacks, and that the production side is no longer a barrier. Using a public dataset of 43 speakers as source material, the authors generated swapped-face video with an open-source face-swapping tool, cloned speech with a commercial voice-synthesis service, and lip-synced combinations with a speech-driven animation model, all on an ordinary desktop computer. In a questionnaire with 44 respondents, participants were told that some clips were fakes and were asked to identify them; the paper reports 66.2% of audio selections were incorrect and 43.1% of video selections were incorrect, and in the combined condition the audio component stayed about as hard to detect (68.3% incorrect) while the video component became easier (30.2% incorrect). The paper reads these rates as evidence that a large share of such attacks would succeed, that synthetic audio is the most dangerous component, and that people without prior knowledge of deepfakes—along with older adults—are especially vulnerable.
Load-bearing premise
The central claim depends on treating an incorrect answer on a test where participants were told fakes exist as the same thing as being deceived by a fake in a real-world attack.
Editorial extensions
If this is right
- Because the fakes were produced on a desktop computer with accessible tools, the cost and skill barrier for running deepfake spearphishing attacks is low enough that this is a current threat, not a distant one.
- Audio is the riskier medium in this study: it was misidentified more often than video, and in the combined condition the audio portion stayed nearly as hard to detect while the video portion became easier to spot.
- Older respondents and respondents with no prior knowledge of deepfakes had lower detection rates, which points to awareness and training as targeted countermeasures.
- The poor visual quality of existing lip-syncing tools currently helps defenders, since adding lip-synced video made the video component easier to detect.
- The documented real-world case of a CEO's cloned voice gives the lab results a concrete anchor: the authors treat the audio error rate as the share of audio-based attacks that could plausibly succeed.
Reading between the lines
- Beyond the paper's aggregate numbers, the 66% and 43% figures are rates over selections, not over people; a participant who incorrectly flags several genuine clips as fake counts as an error without having been deceived by a fake. Recomputing the rates for fake-only and genuine-only items would separate true deception from an over-cautious response bias.
- Because participants were told that fakes might be present, the experiment may overstate how well people would perform when nothing has primed their suspicion; running the same test without any warning would give a cleaner measure of real-world vulnerability and could plausibly show even higher deception rates.
- The paper uses one speaker's voice and actors recorded in identical studio conditions, so the results may not transfer to attacks that impersonate a widely recognized public figure; repeating the protocol with a well-known voice would test whether familiarity helps detection or makes the clone more believable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether deepfake audio and video could enable spearphishing attacks. It reviews reported and hypothetical cases, creates deepfake media using DeepFaceLab, ResembleAI, and speech-driven animation on a consumer desktop, and then presents 44 participants with genuine and fake clips across separate audio, video, and combined conditions. The abstract's headline finding is that 66% of participants failed to identify AI-created audio as fake and 43% failed to identify such videos as fake, leading the authors to conclude that deepfake-enabled spearphishing is a serious and easily accessible threat.
Significance. If the headline numbers were valid, the study would be a useful empirical data point on human detection of deepfake media in a security context. The paper gives credit for several things: it actually creates deepfake media rather than only speculating, it uses a public audio-video dataset, it collects data across age groups and before/after threat perceptions, and it compares audio, video, and combined conditions. However, the central quantitative claim is not supported by the analysis as reported: the 66% and 43% figures are overall incorrect selection rates over a mixture of genuine and fake clips, not failure-to-detect rates for fake clips. Without a corrected analysis, the abstract's main claim and the conclusion's inference about attack success do not follow from the experiment.
major comments (4)
- [Section 5 and Section 4.4.1] The abstract's headline claim is not what Section 5 actually measures. Section 4.4.1 states that the video task contains 3 fake and 5 genuine clips and the audio task contains 3 fake and 10 genuine clips, while Section 5 reports 'of the audio selections made 66.2% were incorrect' and '43.1% incorrect when selecting video examples.' These are overall error rates over all clips, pooling missed fakes with false positives on genuine clips. Because genuine clips heavily outnumber fakes (10:3 in audio, 5:3 in video), a participant who is told fakes are present and therefore mislabels many genuine clips as fake can be counted as 'incorrect' even while detecting every actual fake. The abstract's phrase 'failed to identify AI created audio as fake' requires the miss rate restricted to fake items, which is never reported. Please report a per-condition confusion matrix, the fake-only miss rate, the false-alarm rate on genuine clips, and confidence intervals.
- [Section 6 (Conclusion)] The inference that 'if only 34% of assumptions were correct when attempting to identify fake audio samples then ... the remaining 66% of attacks would convince a victim into falling for a phishing attack' conflates an incorrect forced-choice classification in a laboratory task with falling victim to a spearphishing attack. The experiment did not measure whether participants would click a link, enter credentials, transfer funds, or otherwise comply. Moreover, Section 4.4.1 tells participants that fake examples are present, which is unlike most real-world encounters and can inflate suspicion and false positives. This conclusion is therefore not supported by the data.
- [Section 5, Table 2] The age-disaggregated numbers in Table 2 appear inconsistent with the overall correct-rate figures in Section 5. Using the participant counts in Table 1, a weighted average of the 'Video Detection Rate' column is roughly 36%, whereas Section 5 reports 56.9% correct when selecting video examples. This suggests that Table 2 and Section 5 are using different definitions of 'detection rate' (for example, per-fake detection versus per-selection accuracy), or that the table contains an error. The metric must be defined explicitly for each reported statistic, and raw counts should accompany percentages. In addition, several age cells contain only one or two participants (e.g., 'Below 18' has n=1), so the conclusion in Section 6 that 'the older an individual is, the less likely they are to correctly identify artificial media' is not supported by these data.
- [Section 5 generally] No inferential statistics, confidence intervals, or chance-level baselines are provided. For the audio task, a participant who simply classified every clip as genuine would be correct on 10 of 13 clips (76.9%), while a participant who classified every clip as fake would be correct on only 3 of 13 (23.1%); the reported 33.8% overall correct rate therefore cannot be interpreted without a chance baseline and a measure of variability. The absence of such analysis is particularly consequential for the small age subgroups where percentages such as 100% and 0% appear in Tables 3 and 4.
minor comments (6)
- [Section 4.1] The hardware described (Intel i7-8700K, GTX 1080 Ti, 32GB RAM) is perhaps better described as a 'consumer desktop' rather than a 'low-spec computing facility,' which is the phrase used in the introduction and conclusion.
- [Section 4.4.1] The phrase 'a varied combination of real and genuine audio and video' is confusing because 'real' and 'genuine' are synonyms; please clarify the intended contrast.
- [Tables 3 and 4] Entries of 100.00% and 0.00% in small age cells should be suppressed or accompanied by raw counts, and the 'N/A' cells should be explained (e.g., no participants from Form A in those age groups).
- [Abstract and Section 5] The abstract rounds 66.2% and 43.1% to 66% and 43%, respectively; this is acceptable, but the paper should ensure that all quoted percentages are traceable to the same denominator and metric.
- [References] References [38] and [39] both appear to cite the same arXiv paper by Li and Lyu; one duplicate should be removed, and the in-text citations should be checked.
- [General presentation] There are numerous typographical and formatting errors, including 'confidently,' 'Artificial,' and inconsistent spacing in the references; a careful proofread is needed.
Circularity Check
No circularity: the paper is an empirical measurement study with no fitted parameters, equations, or self-citation chain; the disputed 66%/43% figures are a construct-validity issue, not a circular derivation.
full rationale
The paper's central claim is an experimental result: 66% of participants failed to identify AI-created audio as fake and 43% failed for video. The underlying analysis is measurement, not derivation. There are no equations, fitted parameters, or models whose outputs are fed back as inputs. The authors use third-party datasets and tools (Sanderson's VidTIMIT, DeepFaceLab, Resemble AI, Speech-Driven Animation) and external references; no load-bearing conclusion rests on a citation to the authors' own prior work. The abstract's phrasing does conflate an overall incorrect-selection rate (66.2% of audio selections and 43.1% of video selections, Section 5) with the fake-only miss rate implied by 'failed to identify AI created audio/video.' That conflation is a threat to the validity of the headline inference, and Section 4.4.1's uneven mix of 3 fake vs 10 genuine audio clips and 3 fake vs 5 genuine video clips, together with Section 5's own acknowledgement of false positives, are relevant limitations. However, conflated metrics and unsupported inductive leaps are not circularity: the reported rates are not defined in terms of the conclusion, no parameter is fitted to a subset and then reported as a prediction of a closely related quantity, and the result is not equivalent to its inputs by construction. The paper is therefore best characterized as empirically weak in its interpretation but not circular.
Assumptions & free parameters
assumptions (4)
- domain assumption Failure to label a deepfake as fake in a survey is equivalent to susceptibility to a real spearphishing attack.
- domain assumption The 44 self-selected participants are representative of 'average individuals' and of likely spearphishing targets.
- domain assumption Deepfakes generated by DeepFaceLab, ResembleAI, and Speech-Driven Animation on the stated hardware are representative of real-world AI-powered spearphishing media.
- domain assumption The base-rate structure and the warning that some clips are fake do not distort the headline error rates.
Cite this review
Pith. "Pith review of AI-Powered Spearphishing Cyber Attacks: Fact or Fiction?." pith.science (2026). https://pith.science/paper/VEPNVGH2
@misc{pith2026250200961,
author = {Pith},
title = {Pith review of: AI-Powered Spearphishing Cyber Attacks: Fact or Fiction?},
year = {2026},
howpublished = {\url{https://pith.science/paper/VEPNVGH2}},
note = {Machine review of arXiv:2502.00961}
}
read the original abstract
Due to society's continuing technological advance, the capabilities of machine learning-based artificial intelligence systems continue to expand and influence a wider degree of topics. Alongside this expansion of technology, there is a growing number of individuals willing to misuse these systems to defraud and mislead others. Deepfake technology, a set of deep learning algorithms that are capable of replacing the likeness or voice of one individual with another with alarming accuracy, is one of these technologies. This paper investigates the threat posed by malicious use of this technology, particularly in the form of spearphishing attacks. It uses deepfake technology to create spearphishing-like attack scenarios and validate them against average individuals. Experimental results show that 66% of participants failed to identify AI created audio as fake while 43% failed to identify such videos as fake, confirming the growing fear of threats posed by the use of these technologies by cybercriminals.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
, 2008. Vidtimit audio-video dataset. URL:http://conradsanderson. id.au/vidtimit/
work page 2008
-
[2]
Anna skelton - analyzing the effects of deepfakes on market manipulation -def con 27 aivillage
, 2019. Anna skelton - analyzing the effects of deepfakes on market manipulation -def con 27 aivillage. URL:https://www.youtube.com/ watch?v=BlS-inTxnp8
work page 2019
-
[3]
Conmen made /uni20AC8m by impersonating french minister - is- raeli police
, 2019. Conmen made /uni20AC8m by impersonating french minister - is- raeli police. URL: https://www.theguardian.com/world/2019/mar/28/ conmen-made-8m-by-impersonating-french-minister-israeli-police
work page 2019
-
[4]
Phishing attacks: defending your organisation
, 2019. Phishing attacks: defending your organisation. URL:https: //www.ncsc.gov.uk/guidance/phishing
work page 2019
-
[5]
Why deepfakes pose an unprecedented threat to businesses
, 2019. Why deepfakes pose an unprecedented threat to businesses. URL: https://aibusiness.com/document.asp?doc%5C_id=760904
work page 2019
-
[6]
, 2020. URL: https://publish.twitter.com/?query=https%3A%2F% 2Ftwitter.com%2Felonmusk%2Fstatus%2F1256239815256797184&widget= Tweet
work page 2020
-
[7]
, 2020. Adobe. vfx amp; motion graphics software | adobe after ef- fects. URL:https://www.adobe.com/uk/products/aftereffects.html
work page 2020
-
[8]
,2020.Aigeneratedvoices resembleai.URL: https://www.resemble. ai/
work page 2020
Show all 58 references
-
[9]
Dinoman/speech-driven-animation
, 2020. Dinoman/speech-driven-animation. URL: https://github. com/DinoMan/speech-driven-animation
2020
-
[10]
Fraudsters used ai to mimic ceo’s voice in un- usual cybercrime case
, 2020. Fraudsters used ai to mimic ceo’s voice in un- usual cybercrime case. URL: https://www.wsj.com/articles/ fraudsters-use-ai-to-mimic-ceos-voice-in-unusual-cybercrime-case-% 11567157402
2020
-
[11]
, 2020. Maltego. URL:https://www.maltego.com/
2020
-
[12]
resemble-ai/resemblyzer
, 2020. resemble-ai/resemblyzer. URL: https://github.com/ resemble-ai/Resemblyzer
2020
-
[13]
Voice biometric authentication amp; anti-fraud for call cen- ters
, 2020. Voice biometric authentication amp; anti-fraud for call cen- ters. URL: https://www.pindrop.com/
2020
-
[14]
Security awareness training: A review, in: Proceedings of the World Congress on Engi- neering, pp
Al-Daeef, M.M., Basir, N., Saudi, M.M., 2017. Security awareness training: A review, in: Proceedings of the World Congress on Engi- neering, pp. 5–7
2017
-
[15]
Arik, S.O., Chrzanowski, M., Coates, A., Diamos, G., Gibian- sky, A., Kang, Y., Li, X., Miller, J., Ng, A., Raiman, J., et al.,
-
[16]
Artificialintelligenceand uk national security
Babuta,A.,Oswald,M.,Janjeva,A.,2020. Artificialintelligenceand uk national security
2020
-
[17]
Pindr0p: using single-ended audio features to de- terminecallprovenance,in: Proceedingsofthe17thACMconference on Computer and communications security, pp
Balasubramaniyan,V.A.,Poonawalla,A.,Ahamad,M.,Hunter,M.T., Traynor, P., 2010. Pindr0p: using single-ended audio features to de- terminecallprovenance,in: Proceedingsofthe17thACMconference on Computer and communications security, pp. 109–120
2010
-
[18]
Threat of deepfakes draws legislator and biometrics industry attention
Burt, C., 2019. Threat of deepfakes draws legislator and biometrics industry attention. URL: https://www.biometricupdate.com/201902/ threat-of-deepfakes-draws-legislator-and-biometrics-industry-attention
2019
-
[19]
Analysis | how misinformation helped spark an attempted coup in gabon
Cahlan, S., 2020. Analysis | how misinformation helped spark an attempted coup in gabon. URL: https://www.washingtonpost.com/politics/2020/02/13/ how-sick-president-suspect-video-helped-sparked-an-attempted-coup-% gabon/
2020
-
[20]
Carella,A.,Kotsoev,M.,Truta,T.M.,2017.Impactofsecurityaware- ness training on phishing click-through rates, in: 2017 IEEE Interna- tional Conference on Big Data (Big Data), IEEE. pp. 4458–4466
2017
-
[21]
URL: https://www.ncsc.gov.uk/guidance/ whaling-how-it-works-and-what-your-organisation-can-do-about-it
Centre, N.C.S., 2020. URL: https://www.ncsc.gov.uk/guidance/ whaling-how-it-works-and-what-your-organisation-can-do-about-it
2020
-
[22]
Snapshot Paper - Deepfakes and Audiovisual Disinformation
for Data Ethics, C., Innovation, 2019. Snapshot Paper - Deepfakes and Audiovisual Disinformation
2019
-
[23]
Detecting audio deep fakes with ai
Dessa, 2019. Detecting audio deep fakes with ai. URL: https:// medium.com/dessa-news/detecting-audio-deepfakes-f2edfd8e2b35
2019
-
[24]
dessa-oss/fake-voice-detection
Dessa-Oss, 2020. dessa-oss/fake-voice-detection. URL: https:// github.com/dessa-oss/fake-voice-detection
2020
-
[25]
Cyber Security Breaches Survey
Department for Digital, Culture, M., Sport, 2020. Cyber Security Breaches Survey. Ph.D. thesis
2020
-
[26]
Why fake video, audio may not be as powerful in spreading disinformation as feared
Ewing, P., 2020. Why fake video, audio may not be as powerful in spreading disinformation as feared. URL: https://www.npr.org/2020/05/07/851689645/ why-fake-video-audio-may-not-be-as-powerful-in-spreading-% disinformation-as-feare?t=1595494301688
2020
-
[27]
Why deepfakes are a net positive for human- ity
Forbes, 2020. Why deepfakes are a net positive for human- ity. URL: https://www.forbes.com/sites/simonchandler/2020/03/09/ why-deepfakes-are-a-net-positive-for-humanity/#e512a3b2f84f
2020
-
[28]
iperov/deepfacelab
Iperov, 2020. iperov/deepfacelab. URL:https://github.com/iperov/ DeepFaceLab
2020
-
[29]
Transfer learning fromspeakerverificationtomultispeakertext-to-speechsynthesis,in: Advances in neural information processing systems, pp
Jia, Y., Zhang, Y., Weiss, R., Wang, Q., Shen, J., Ren, F., Nguyen, P., Pang, R., Moreno, I.L., Wu, Y., et al., 2018. Transfer learning fromspeakerverificationtomultispeakertext-to-speechsynthesis,in: Advances in neural information processing systems, pp. 4480–4490
2018
-
[30]
Artificial intelligence in digital media: The era of deepfakes
Karnouskos, S., 2020. Artificial intelligence in digital media: The era of deepfakes. IEEE Transactions on Technology and Society 1, 138–147
2020
-
[31]
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., Aila, T.,
-
[32]
Cyber security breaches survey
Klahr, R., 2017. Cyber security breaches survey. Ph.D. thesis. Uni- versity of Portsmouth
2017
-
[33]
Deepfakes: a new threat to face recognition? assessment and detection
Korshunov, P., Marcel, S., 2018a. Deepfakes: a new threat to face recognition? assessment and detection. arXiv preprint arXiv:1812.08685
-
[34]
Speakerinconsistencydetectionin tampered video, in: 2018 26th European Signal Processing Confer- ence (EUSIPCO), IEEE
Korshunov,P.,Marcel,S.,2018b. Speakerinconsistencydetectionin tampered video, in: 2018 26th European Signal Processing Confer- ence (EUSIPCO), IEEE. pp. 2375–2379
2018
-
[35]
This day in history: Hacked ap tweet about white house explosions triggers panic
Langlois, S., 2018. This day in history: Hacked ap tweet about white house explosions triggers panic. URL: https://www.marketwatch.com/story/ this-day-in-history-hacked-ap-tweet-about-white-house-explosions-% triggers-panic-2018-04-23
2018
-
[36]
IEEE transactions on visualization and computer graphics 19, 1859–1871
Le,B.H.,Zhu,M.,Deng,Z.,2013.Markeroptimizationforfacialmo- tion acquisition and deformation. IEEE transactions on visualization and computer graphics 19, 1859–1871
2013
-
[37]
In ictu oculi: Exposing ai gen- erated fake face videos by detecting eye blinking
Li, Y., Chang, M.C., Lyu, S., 2018. In ictu oculi: Exposing ai gen- erated fake face videos by detecting eye blinking. arXiv preprint : Preprint submitted to Elsevier Page 10 of 11 AI-Powered Spearphishing arXiv:1806.02877
2018 arXiv
-
[39]
Exposing deepfake videos by detecting face warping artifacts
Li, Y., Lyu, S., 2018b. Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656
-
[40]
Face geometry and appearance modeling: concepts and applications
Liu, Z., Zhang, Z., 2011. Face geometry and appearance modeling: concepts and applications. Cambridge University Press
2011
-
[41]
Deep face recogni- tion
Parkhi, O.M., Vedaldi, A., Zisserman, A., 2015. Deep face recogni- tion
2015
-
[42]
Deepfacelab: A simple, flexible and extensible face swapping framework
Petrov, I., Gao, D., Chervoniy, N., Liu, K., Marangonda, S., Umé, C., Jiang, J., RP, L., Zhang, S., Wu, P., et al., 2020. Deepfacelab: A simple, flexible and extensible face swapping framework. arXiv preprint arXiv:2005.05535
2020 arXiv
-
[43]
Historyofphishing
Phishing.Org,. Historyofphishing. URL: https://www.phishing.org/ history-of-phishing
-
[44]
arXiv preprint arXiv:1710.07654
Ping,W.,Peng,K.,Gibiansky,A.,Arik,S.O.,Kannan,A.,Narang,S., Raiman,J.,Miller,J.,2017.Deepvoice3: Scalingtext-to-speechwith convolutional sequence learning. arXiv preprint arXiv:1710.07654
2017 arXiv
-
[45]
Exploring historical and emerg- ing phishing techniques and mitigating the associated security risks
Rader, M., Rahman, S., 2015. Exploring historical and emerg- ing phishing techniques and mitigating the associated security risks. arXiv preprint arXiv:1512.00082
2015 arXiv
- [46]
-
[47]
Recurrentconvolutionalstrategiesforfacemanipulation detection in videos
Sabir, E., Cheng, J., Jaiswal, A., AbdAlmageed, W., Masi, I., Natara- jan,P.,2019. Recurrentconvolutionalstrategiesforfacemanipulation detection in videos. Interfaces (GUI) 3
2019
-
[48]
Multi-region probabilistic his- tograms for robust and scalable identity inference, in: International conference on biometrics, Springer
Sanderson, C., Lovell, B.C., 2009. Multi-region probabilistic his- tograms for robust and scalable identity inference, in: International conference on biometrics, Springer. pp. 199–208
2009
-
[49]
Towards detection of morphed face images in electronic travel documents, in: 2018 13th IAPR International Workshop on Document Analysis Systems (DAS), IEEE
Scherhag, U., Rathgeb, C., Busch, C., 2018. Towards detection of morphed face images in electronic travel documents, in: 2018 13th IAPR International Workshop on Document Analysis Systems (DAS), IEEE. pp. 187–192
2018
-
[50]
Scherhag, U., Rathgeb, C., Merkle, J., Breithaupt, R., Busch, C.,
-
[51]
Facenet: A unified embedding for face recognition and clustering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp
Schroff, F., Kalenichenko, D., Philbin, J., 2015. Facenet: A unified embedding for face recognition and clustering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 815–823
2015
-
[52]
End-to-end speech-driven facial animation with temporal gans
Vougioukas, K., Petridis, S., Pantic, M., 2018. End-to-end speech-driven facial animation with temporal gans. arXiv preprint arXiv:1805.09313
2018 arXiv
-
[53]
Thedeepfakethreattofacebiometrics
Wojewidka,J.,2020. Thedeepfakethreattofacebiometrics. Biomet- ric Technology Today 2020, 5–7
2020
-
[54]
Exposing deep fakes using inconsis- tent head poses, in: ICASSP 2019-2019 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), IEEE
Yang, X., Li, Y., Lyu, S., 2019. Exposing deep fakes using inconsis- tent head poses, in: ICASSP 2019-2019 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), IEEE. pp. 8261–8265
2019
-
[55]
Two-stream neu- ral networks for tampered face detection, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE
Zhou, P., Han, X., Morariu, V.I., Davis, L.S., 2017. Two-stream neu- ral networks for tampered face detection, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE. pp. 1831–1839
2017
-
[56]
Zhu, B., Fang, H., Sui, Y., Li, L., 2020. Deepfakes for medical video de-identification: Privacy protection and diagnostic informa- tion preservation, in: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pp. 414–420. : Preprint submitted to Elsevier Page 11 of 11
2020
-
[2017]
arXiv preprint arXiv:1702.07825
Deep voice: Real-time neural text-to-speech. arXiv preprint arXiv:1702.07825
-
[2019]
IEEE Access 7, 23012–23026
Face recognition systems under morphing attacks: A survey. IEEE Access 7, 23012–23026
-
[2020]
8110–8119
Analyzing and improving the image quality of stylegan, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8110–8119
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.