REVIEW 1 major objections 1 cited by
A weighted partial similarity framework detects partial code reuse at the project level more reliably than prior methods.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 10:22 UTC pith:4Q5YCNSX
load-bearing objection The paper introduces sensible mechanisms for project-level birthmark comparison but the evaluation setup using same-project versions falls short of testing the partial cross-project reuse it aims to address. the 1 major comments →
Project-wise Comparison of Software Birthmarks Using Weighted Partial Similarity
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By establishing a symmetric aggregation framework for project-wise birthmark comparison, introducing a weighting scheme that prioritizes larger modules, and applying a partial similarity method focused on the top fraction of highly similar module pairs, the approach achieves robust and stable detection of partial code reuse at the project level and consistently outperforms existing methods when evaluated on resilience and credibility across 35 Java projects in ten categories.
What carries the argument
Symmetric aggregation of module-level similarities, combined with size-based weighting and top-fraction partial similarity selection.
Load-bearing premise
Treating different versions of the same project as valid examples of partial reuse accurately models how modules are selectively copied between unrelated projects.
What would settle it
If the method fails to identify known partial module reuses between distinct unrelated projects or produces many false positives among independent projects, the performance advantage would be falsified.
If this is right
- Detects reuse even when only a small fraction of modules is copied.
- Reduces false positives from incidental similarities in small modules.
- Maintains stable results across varied project categories.
- Improves the combined resilience-credibility score via harmonic mean.
Where Pith is reading between the lines
- The method could support automated checks for open-source license compliance when projects borrow code fragments.
- Project-level analysis may become necessary for any realistic reuse or plagiarism tool.
- Testing the same weighting and partial focus on non-Java codebases would check whether the gains generalize.
- Embedding the approach in build or version-control pipelines could enable ongoing reuse monitoring.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a framework for project-wise software birthmark comparison via symmetric aggregation of module-level similarities. It introduces a weighting scheme that prioritizes larger modules to mitigate noise from small ones, and a partial similarity method that considers only the top fraction of highly similar module pairs to detect partial reuse. The approach is evaluated on 35 open-source Java projects across ten categories by treating pairs of different versions of the same project as positive reuse cases. Performance is measured via resilience and credibility (combined by harmonic mean), with the claim that the method consistently outperforms existing approaches for robust project-level partial reuse detection. The dataset and artifacts are released publicly.
Significance. If the central claim holds after addressing the evaluation gap, the work would provide a practical advance in birthmark-based detection of partial code reuse at the project level, which is more representative of real-world scenarios than module-level comparisons. The public release of the dataset and experimental artifacts is a clear strength, enabling reproducibility and follow-on research.
major comments (1)
- [Evaluation section] Evaluation section: Positive cases are constructed exclusively from version pairs within the same 35 projects. These pairs share evolutionary history and large overlapping codebases, unlike the selective module copying across unrelated projects that motivates the weighting scheme and top-fraction partial similarity. Because the mechanisms target the cross-project partial case yet the reported outperformance is measured only on the within-project version task, it is unclear whether the performance advantage generalizes to the intended scenario.
Simulated Author's Rebuttal
We thank the referee for highlighting this important distinction in the evaluation design. We provide a point-by-point response below and indicate where revisions will be made to strengthen the manuscript.
read point-by-point responses
-
Referee: [Evaluation section] Evaluation section: Positive cases are constructed exclusively from version pairs within the same 35 projects. These pairs share evolutionary history and large overlapping codebases, unlike the selective module copying across unrelated projects that motivates the weighting scheme and top-fraction partial similarity. Because the mechanisms target the cross-project partial case yet the reported outperformance is measured only on the within-project version task, it is unclear whether the performance advantage generalizes to the intended scenario.
Authors: We agree that the positive cases are intra-project version pairs rather than selective cross-project module copying. Version pairs were chosen because they provide controlled, realistic partial reuse (modules added/removed/modified across releases) while allowing direct measurement of resilience under evolutionary change—the core property birthmarks must satisfy. This is a standard evaluation practice in the birthmark literature for the same reason. Nevertheless, the referee correctly notes that this does not fully replicate the cross-project selective-reuse scenario that motivates the weighting and partial-similarity mechanisms. We will revise the evaluation section to (a) explicitly acknowledge this limitation, (b) add a new subsection discussing how the proposed mechanisms are expected to behave under cross-project reuse, and (c) include at least one additional cross-project experiment using the released dataset (pairing modules from unrelated projects within the same category) to provide supporting evidence. These changes will be reflected in the revised manuscript. revision: yes
Circularity Check
No circularity; derivation and evaluation are self-contained against external benchmarks
full rationale
The paper defines a symmetric aggregation framework plus weighting and top-fraction partial-similarity mechanisms from first principles to address partial reuse and incidental similarity. These definitions stand independently of the evaluation data. The evaluation applies the method to a public dataset of 35 Java projects (version pairs treated as reuse cases) and reports resilience/credibility harmonic means; no parameters are fitted to the test outcomes, no self-citation chain supports the central claims, and no equation reduces to its own input by construction. The choice of positive examples is an external modeling decision, not a definitional loop.
Axiom & Free-Parameter Ledger
read the original abstract
Software birthmarks provide a robust approach to detecting code plagiarism even under substantial modifications, while distinguishing independently developed software. Existing similarity measures are typically applied at the module level (e.g., source or class files). However, in practice, software reuse often occurs at the project level, where only a subset of modules may be reused. This setting introduces two key challenges: (1) partial reuse, where reused modules constitute only a small fraction of the project, and (2) incidental similarity from small modules, which can lead to false positives. In this paper, we establish a framework for project-wise birthmark comparison based on a symmetric aggregation of module-level similarities. On top of this framework, we propose two complementary mechanisms to address the above challenges. First, we introduce a weighting scheme that assigns higher importance to larger modules, reducing the influence of noisy matches from small modules. Second, we propose a partial similarity method that focuses on the top fraction of highly similar module pairs, enabling robust detection of partial reuse. We evaluate the proposed approach on 35 open-source Java projects across ten categories, where different versions of the same project are treated as reuse cases. The dataset and experimental artifacts are made publicly available to support reproducibility. Performance is assessed using two complementary properties of software birthmarks, resilience and credibility, combined via their harmonic mean. The results show that the proposed method consistently outperforms existing approaches, achieving robust and stable detection of partial code reuse at the project level.
Figures
Forward citations
Cited by 1 Pith paper
-
Detection of LLM-assisted Code Plagiarism Using k-gram Software Birthmarks
Opcode k-gram software birthmarks, especially 2-grams with Dice similarity, detect LLM code paraphrasing with high Hmean, though ChatGPT-style models are hardest to catch.
Reference graph
Works this paper leans on
-
[1]
Open Source Software Detection using Function-level Static Software Birthmark
D. Kim, S. Cho, S. Han, M. Park, and I. You, “Open Source Software Detection using Function-level Static Software Birthmark”,Journal of Internet Services and Information Security (JISIS), vol.4, no.4, pp. 25– 37, 2014. doi: 10.22667/JISIS.2014.11.31.025
-
[2]
D´ ej` aVu: A Map of Code Duplicates on GitHub
C. V. Lopes, P. Maj, P. Martins, V. Saini, D. Yang, J. Zitny, H. Sajnani, and J. Vitek, “D´ ej` aVu: A Map of Code Duplicates on GitHub”, inProc. ACM Program. Lang., vol. 1, 28 pages, 2017. doi: 10.1145/3133908
-
[3]
Identifying Open-Source License Violation and 1-day Security Risk at Large Scale
R. Duan, A. Bijlani, M. Xu, T. Kim, and W. Lee, “Identifying Open-Source License Violation and 1-day Security Risk at Large Scale”,CCS ’17, USA, 2017. doi: 10.1145/3133956.3134048
-
[4]
Stack Overflow: A Code Laundering Platform?
L. An, o. Mlouki, F. Khomh, and G. Antoniol, “Stack Overflow: A Code Laundering Platform?”. arXiv:1703.03897 [cs.SE], 2017
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[5]
Sourcerer's Apprentice and the study of code snippet migration
S. Romansky, C. Chen, B. Malhotra, and A. Hindle, “Sourcerer’s Apprentice and the study of code snippet migration”. arXiv:1808.00106 [cs.SE], 2018
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[6]
A Study of Potential Code Borrowing and License Violations in Java Projects on GitHub
Y. Golubev, M. Eliseeva, N. Povarov, and T. Bryksin, “A Study of Potential Code Borrowing and License Violations in Java Projects on GitHub”. arXiv:2002.05237 [cs.SE], 2020
-
[7]
Detecting the Theft of Programs Us- ing Birthmarks
H. Tamada, M. Nakamura, A. Monden, and K. Matsumoto, “Detecting the Theft of Programs Us- ing Birthmarks”,Information Science Technical Re- port, NAIST-IS-TR2003014, ISSN 0919-9527, Grad- uate School of Information Science, Nara Institute of Science and Technology, Japan, 2003
work page 2003
-
[8]
k-gram based software birthmarks
G. Myles and C. Collberg, “k-gram based software birthmarks”,Proc. 2005 ACM symposium on Applied Computing, pp.314—318, ACM, 2005
work page 2005
-
[9]
Design and evaluation of birthmarks for de- tecting theft of Java programs
H. Tamada, M. Nakamura, A. Monden, and K. Mat- sumoto, “Design and evaluation of birthmarks for de- tecting theft of Java programs”, inProc. IASTED SE 2004, pp.569—575, Innsbruck, Austria, 2004
work page 2004
-
[10]
Java birthmarks - detecting the software theft
H. Tamada, M. Nakamura, A. Monden, and K. Mat- sumoto, “Java birthmarks - detecting the software theft”,IEICE Trans. Inf. Syst., E88-D(9): pp.2148— 2158, 2005. doi: 10.1093/ietisy/e88-d.9.2148
-
[11]
API-based software birthmarking method using fuzzy hashing,
D. Lee, D. Kang, Y. Choi, J. Kim, and D. Won, “API-based software birthmarking method using fuzzy hashing,”IEICE Trans. Inf. Syst., E99.D(7): pp.1836-–1851, 2016. doi: 10.1587/transinf.2015EDP7379
-
[12]
A static API birthmark for windows binary executa- bles
S. Choi, H. Park, H. I. Lim, and T. Han, “A static API birthmark for windows binary executa- bles”,Journal of Systems and Software, 82(5):862— 873, 2009
work page 2009
-
[13]
Detecting Software Theft with API Call Sequence Sets
D. Schuler and V. Dallmeier, “Detecting Software Theft with API Call Sequence Sets”, inProc. of the 8th Workshop Software Reengineering (WSR’06), Germany, 2006
work page 2006
-
[14]
Comparison of Similarity Functions for n-gram Software Birthmarks,
N. Fedorov, H. Tamada, H. Inayoshi, and A. Mon- den, “Comparison of Similarity Functions for n-gram Software Birthmarks,” inProc. of the 2024 The 6th World Symposium on Software Engineering (WSSE), pp.169—176, 2024. doi: 10.1145/3698062.3698087
-
[15]
Detecting software theft via whole program path birthmarks
G. Myles and C. Collberg, “ Detecting software theft via whole program path birthmarks”,Information se- curity, pp.404—415, Springer, 2004
work page 2004
-
[16]
Using software birth- marks to identify similar classes and major function- alities
T. Kakimoto, A. Monden, Y. Kamei, H. Tamada, M. Tsunoda, and K. Matsumoto, “Using software birth- marks to identify similar classes and major function- alities”, inProc. of the 2006 International Workshop on Mining Software Repositories (MSR ’06), Associa- tion for Computing Machinery, pp.171–172, 2006. doi: 10.1145/1137983.1138026. 16 Table 5: Open-sourc...
-
[17]
IBM Corporation, “Effects of nested classes.” [On- line]. Available:www.ibm.com/docs/en/clearcase/ 11.0.0?topic=omake-effects-nested-classes, Accessed on: Mar. 10, 2025
work page 2025
-
[18]
Oracle Corporation, “Nested Classes.” [Online]. Available:https://docs.oracle.com/javase/ tutorial/java/javaOO/nested.html, Accessed on: Mar. 10, 2025
work page 2025
-
[19]
Software plagiarism detection: a graph- based approach
D.-K. Chae, J. Ha, S.-W. Kim, B. Kang, and E. G. Im, “Software plagiarism detection: a graph- based approach”,Proc. 22nd ACM Intern. Conf. on Inf. and Knowl. Manag., pp.1577–1580, 2013. doi: 10.1145/2505515.2507848
-
[20]
A sur- vey of software watermarking
W. Zhu, C. Thomborson, and F.-Y. Wang, “A sur- vey of software watermarking”,Intel. and Sec. Inf., pp.454-–458, Springer, 2005
work page 2005
-
[21]
A practical method for watermarking Java programs
A. Monden, H. Iida, K. Matsumoto, K. Inoue, and K. Torii, “A practical method for watermarking Java programs”, inProc. 24th IEEE compsac2000, pp.191– 197, Taipei, Taiwan, 2000
work page 2000
-
[22]
A method for detecting the theft of java programs through anal- ysis of the control flow information
H.-I. Lim, H. Park, S. Choi, and T. Han, “A method for detecting the theft of java programs through anal- ysis of the control flow information”,Information and Software Technology, vol.51, no.9, pp.1338-–1350, 2009
work page 2009
-
[23]
Software birthmark design and estimation: A systematic litera- ture review
S. Nazir, S. Shahzad, and N. Mukhtar, “Software birthmark design and estimation: A systematic litera- ture review”,Arab. Journal for Science and Engineer- ing, vol.44, pp.3905–3927, 2019. doi: 10.1007/s13369- 019-03718-9
-
[24]
Design and Evaluation of Dynamic Soft- ware Birthmarks Based on API Calls
H. Tamada, K. Okamoto, M. Nakamura, and A. Monden, “Design and Evaluation of Dynamic Soft- ware Birthmarks Based on API Calls”,Information Science Technical Report, NAIST-IS-TR2007011, ISSN 0919-9527, Graduate School of Information Science, Nara Institute of Science and Technology, Japan, 2007
work page 2007
-
[25]
B. Yuan, J. Wang, Z. Fang, and L. Qi, “A New Soft- ware Birthmark based on Weight Sequences of Dy- namic Control Flow Graph for Plagiarism Detection”, The Computer Journal, vol.61, no.8, pp.1202—1215,
-
[26]
doi: 10.1093/comjnl/bxy055
-
[27]
Dynamic Software Birthmarks to Detect the Theft of Windows Applications
H. Tamada, K. Okamoto, M. Nakamura, and A. Monden, “Dynamic Software Birthmarks to Detect the Theft of Windows Applications”,International Symposium on Future Software Technology, vol. 20, no. 22, 2004
work page 2004
-
[28]
Software Plagiarism Detection with Birthmarks Based on Dynamic Key Instruc- tion Sequences
Z. Tian, Q. Zheng, T. Liu, M. Fan, E. Zhuang, and Z. Yang, “Software Plagiarism Detection with Birthmarks Based on Dynamic Key Instruc- tion Sequences”,IEEE Trans. on Software Engi- neering, vol.41, no.12, pp.1217–1235, 2015. doi: 10.1109/TSE.2015.2454508
-
[29]
Malware Variant Detec- tion Using Similarity Search over Sets of Control Flow Graphs
S. Cesare and Y. Xiang, “Malware Variant Detec- tion Using Similarity Search over Sets of Control Flow Graphs”,2011IEEE 10th International Conference on 17 Trust, Security and Privacy in Computing and Com- munications, Changsha, China, 2011, pp. 181–189. doi: 10.1109/TrustCom.2011.26
-
[30]
A summary of the international standard date and time notation
Markus Kuhn, “A summary of the international standard date and time notation”, [Online]. Avail- able:https://www.cl.cam.ac.uk/ ~mgk25/iso- time.html. Accessed on: Mar. 13, 2025
work page 2025
-
[31]
ReDeBug: finding unpatched code clones in entire os distribu- tions
J. Jang, A. Agrawal, and D. Brumley, “ReDeBug: finding unpatched code clones in entire os distribu- tions”, inProc. of the 33rd IEEE Symposium on Se- curity and Privacy (Oakland), USA, 2012
work page 2012
-
[32]
C. Alejandro, H. Tamada, and Y. Kanzaki, “Towards the auto extraction for the dynamic software birth- marks with the inputs from the plaintiff software”, Proc. of the Forum on Information Technology (FIT), vol.22, no.1, pp.165–168, 2023
work page 2023
-
[33]
Clone search for malicious code correlation
P. Charland, B. C. Fung, and M. R. Farhadi, “Clone search for malicious code correlation”,In ATO RTO Symposium on Information Assurance and Cyber De- fense (IST-111), 2012
work page 2012
-
[34]
Public Git Archive: a Big Code dataset for all
V. Markovtsev, W. Long, “Public Git Archive: a Big Code dataset for all”,Proc. of the 15th International Conference on Mining Software Repositories (ICSE ’18), pp.34–37, 2018. doi: 10.1145/3196398.3196464
-
[35]
A Software Birthmark Based on Dynamic Op- code n-gram
B. Lu, F. Liu, X. Ge, B. Liu, and X. Luo, “A Software Birthmark Based on Dynamic Op- code n-gram”,International Conference on Seman- tic Computing (ICSC 2007), pp.37–44, 2007. doi: 10.1109/ICSC.2007.15
-
[36]
Program Characterization Using Runtime Values and Its Application to Software Plagiarism Detection
Y.-C. Jhi, X. Jia, X. Wang, S. Zhu, P. Liu, and D. Wu, “Program Characterization Using Runtime Values and Its Application to Software Plagiarism Detection”, inProc. of the ACM/IEEE 33rd Inter- national Conference on Software Engineering (ICSE 2011), Software Engineering in Practice Track, USA, 2011
work page 2011
-
[37]
Behav- ior Based Software Theft Detection
X. Wang, Y.-C. Jhi, S. Zhu, and P. Liu, “Behav- ior Based Software Theft Detection”,Proc. of the 16th ACM Conference on Computer and Communi- cations Security (CCS’09), pp.280–290, USA, 2009. doi: 10.1145/1653662.1653696
-
[38]
DKISB: Dynamic key instruction sequence birth- mark for software plagiarism detection
Z. Tian, Q. Zheng, T. Liu, and M. Fan, “DKISB: Dynamic key instruction sequence birth- mark for software plagiarism detection”,in Proc. 2013 IEEE 10th International Conference on High Per- formance Computing and Communications & 2013 IEEE International Conference on Embedded and Ubiquitous Computing, pp. 619-–627, 2013. doi: 10.1109/HPCC.and.EUC.2013.93
-
[39]
D. Schuler, V. Dallmeier, and C. Lindig, “A dynamic birthmark for Java,” inProc. of the twenty-second IEEE/ACM international conference on Automated software engineering (ASE ’07), pp. 274-–283, 2007
work page 2007
-
[40]
Dy- namic k-gram based software birthmark
Y. Bai, X. Sun, G. Sun, X. Deng, and X. Zhou, “Dy- namic k-gram based software birthmark”, inProc. 19th Australian Softw. Eng. Conf., pp. 644-–649, 2008
work page 2008
-
[41]
GPLAG: detec- tion of software plagiarism by program dependence graph analysis
C. Liu, C. Chen, J. Han, P. S. Yu, “GPLAG: detec- tion of software plagiarism by program dependence graph analysis”,KDD ’06, pp. 872–881, 2006
work page 2006
-
[42]
Program Logic Based Software Plagiarism Detection
F. Zhang, D. Wu, P. Liu, and S. Zhu, “Program Logic Based Software Plagiarism Detection”,2014 IEEE 25th International Symposium on Software Reliability Engineering, pp. 66–77, 2014
work page 2014
-
[43]
De- tecting software theft via system call based birth- marks
X. Wang, Y.-C. Jhi, S. Zhu, and P. Liu, “De- tecting software theft via system call based birth- marks”,Computer Security Applications Conference ACSAC’09., 2009
work page 2009
-
[44]
Current Sta- tus of Vizio Case
Software Freedom Conservancy, Inc., “Current Sta- tus of Vizio Case”, Software Freedom Conservancy, [Online]. Available:https://sfconservancy.org/ copyleft-compliance/vizio.html. Accessed on: May 24, 2025
work page 2025
-
[45]
R. G. Sanders, S. V. Vakili, J. A. Schlaff, D. N. Schultz, and S. P. Hoffman, “CASE NO.: 30-2021- 01226723-CU-BC-CJC. COMPLAINT FOR: (1) BREACH OF CONTRACT; and (2) DECLARA- TORY RELIEF”, SUPERIOR COURT OF THE STATE OF CALIFORNIA COUNTY OF OR- ANGE - CENTRAL JUSTICE CENTER. [Online]. Available:https://sfconservancy.org/static/ docs/software-freedom-conser...
work page 2021
-
[46]
IBM Cor- poration v. Teraproc Inc. (7:16-cv-07989)
District Court S.D. New York, “IBM Cor- poration v. Teraproc Inc. (7:16-cv-07989)”, CourtListener. [Online]. Available:https: //www.courtlistener.com/docket/4524777/ibm- corporation-v-teraproc-inc/. Accessed on: May 24, 2025
-
[47]
Plagiarism in Programming Assignments
M. Joy and M. Luck, “Plagiarism in Programming Assignments”, University of Warwick. Department of Computer Science. (Department of Computer Science Research Report). (Unpublished), 1998
work page 1998
-
[48]
Tri Le, A. Carbone, J. Sheard, M. Schuhmacher, M. de Raath and C. Johnson, “Educating computer pro- gramming students about plagiarism through use of a code similarity detection tool”,2013 Learning and Teaching in Computing and Engineering (LATICE ’13), pp. 98–105, 2013. doi: 10.1109/LaTiCE.2013.37
-
[49]
A Generic Approach to Automatic Deobfus- cation of Executable Code
B. Yadegari, B. Johannesmeyer, B. Whitely, and S. Debray, “A Generic Approach to Automatic Deobfus- cation of Executable Code”.In 2015 IEEE Sympo- sium on Security and Privacy. IEEE, USA, pp. 674- –691, 2015. doi: 10.1109/SP.2015.47. 18
-
[50]
A Taxon- omy of Obfuscating Transformations
C. Collberg, C. Thomborson, and D. Low, “A Taxon- omy of Obfuscating Transformations”.Department of Computer Science, The University of Auckland. New Zealand, 1997. URL:http://www.cs.auckland.ac. nz/staff-cgi-bin/mjd/csTRcgi.pl?serial
work page 1997
-
[51]
A survey on software clone detection research,
C. Roy and J. Cordy, “A survey on software clone detection research,” School of Computing, TR 2007- 541, 2007
work page 2007
-
[52]
A Static Java Birthmark Based on Control Flow Edges,
H. -i. Lim, H. Park, S. Choi and T. Han, “A Static Java Birthmark Based on Control Flow Edges,” 2009 33rd Annual IEEE International Computer Soft- ware and Applications Conference, Seattle, WA, USA, 2009, pp. 413–420, doi: 10.1109/COMPSAC.2009.62
-
[53]
Plagiarism detection for multithreaded soft- ware based on thread-aware software birthmarks,
Z. Tian, Q. Zheng, T. Liu, M. Fan, X. Zhang and Z. Yang, “Plagiarism detection for multithreaded soft- ware based on thread-aware software birthmarks, ” In Proceedings of the 22nd International Conference on Program Comprehension (ICPC 2014). Associa- tion for Computing Machinery, New York, USA, 2014, pp. 304–313, doi: 10.1145/2597008.2597143. B Biography...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.