Pith. sign in

REVIEW 2 major objections 5 minor 64 references

SEDCoT raises COBOL-to-C correctness by at least 12% over LLM baselines while keeping translations as readable as human-written C.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 21:44 UTC pith:2B7LQQ3E

load-bearing objection Solid engineering pipeline for COBOL-to-C that delivers a real multi-model lift and readable output; the GnuCOBOL dual-use is a real but bounded limitation, not a collapse of the claim. the 2 major comments →

arxiv 2607.04092 v1 pith:2B7LQQ3E submitted 2026-07-05 cs.SE

SEDCoT: Enhancing LLM-Based COBOL Code Translation via Symbolic Execution and Delta Debugging

classification cs.SE
keywords code translationcode repairlarge language modelsymbolic executiondelta debuggingCOBOLlegacy modernization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

COBOL still runs critical banking and government systems, but rule-based translators produce unreadable C and plain large language models produce incorrect C because COBOL is low-resource and has unusual data and control semantics. This paper claims that a three-phase pipeline—LLM draft, hybrid test generation that mixes symbolic execution on GnuCOBOL’s intermediate C with LLM-generated cases, and iterative LLM repair that finally feeds delta-debugged minimal failing inputs—closes most of that accuracy gap without sacrificing readability. On 319 CodeNet COBOL programs scored against a 500-input golden suite whose oracle is GnuCOBOL itself, the method lifts average pass rates by 12–63% over the strongest prior LLM translation pipelines while human and automated judges rate the resulting C near human-written quality. The practical payoff is a modernization path that yields maintainable code rather than opaque compiler dumps or brittle LLM drafts.

Core claim

Combining LLM initial translation with symbolic-execution and LLM test suites, then repairing residual failures with delta-debugged counterexamples, produces COBOL-to-C translations whose behavioral pass rates exceed those of UniTrans and HRJR by at least 12% while remaining far more readable than GnuCOBOL’s rule-based output.

What carries the argument

SEDCoT’s three-phase loop: (1) LLM draft of C, (2) hybrid test suite obtained by symbolically executing GnuCOBOL’s intermediate C and by prompting an LLM on the original COBOL, (3) iterative LLM repair that, after a few ordinary rounds, supplies only the minimal delta-debugged failing inputs.

Load-bearing premise

That matching GnuCOBOL’s behavior on roughly five hundred synthetically mutated inputs is a faithful proxy for the true semantics of the original COBOL program.

What would settle it

Take a held-out set of COBOL programs whose correct C translations are independently known; if SEDCoT’s golden-suite pass rates no longer exceed UniTrans/HRJR by roughly 12% or higher, or if the readability advantage over GnuCOBOL disappears, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. SEDCoT is a three-phase COBOL-to-C translation framework: (1) LLM initial translation with compilation-error repair, (2) hybrid test generation that runs KLEE symbolic execution on GnuCOBOL intermediate C and also solicits LLM tests from the original COBOL, and (3) iterative LLM repair driven by failing tests, with delta debugging applied only in the final round to produce minimal counterexamples. On 319 CodeNet COBOL programs the method raises average golden-test pass ratio by at least 12.2% (and up to 63%) over UniTrans and HRJR across four LLMs, while human and LLM readability ratings place SEDCoT near human-written C and far above GnuCOBOL output. Ablations attribute gains mainly to symbolic tests and late-stage delta debugging; coverage and unique-bug analyses argue that symbolic inputs expose more latent faults than LLM tests despite comparable line/branch coverage.

Significance. Legacy COBOL modernization is a high-stakes industrial problem; a pipeline that simultaneously improves functional fidelity and readability over pure LLM and pure rule-based baselines is practically valuable. The work supplies a clear experimental protocol (temperature 0, fixed three repair rounds matching baselines, multi-LLM evaluation, component ablations, human readability study) and a public dataset subset, making the central claim falsifiable and the engineering contribution reusable. The explicit integration of symbolic execution and delta debugging into an LLM repair loop for a low-resource language is a concrete methodological advance even if absolute correctness remains below industrial compilers.

major comments (2)
  1. [Section 4.2 / Table 5] Section 4.2 and Table 5: the evaluation oracle and the symbolic-test generator both rest on GnuCOBOL (oracle set to 1.000; KLEE runs on GnuCOBOL intermediate C, Eq. 5). The paper itself records that GnuCOBOL fails a non-zero fraction of the NIST COBOL 85 suite and that CodeNet dialect metadata is incomplete (Section 4.1). Without an independent COBOL 85 reference (or even a small manual audit of the 319 programs) it is impossible to distinguish genuine recovery of original COBOL semantics from improved matching of GnuCOBOL’s own model. This dual-use is load-bearing for the ≥12% claim and should be quantified or bounded.
  2. [Section 4.1 / Section 8] Section 4.1 and Threats (Section 8): the study is restricted to 319 function-level CodeNet programs and to C as the sole target language because open-source COBOL-to-Java translation yielded only 11 successes. The abstract and introduction frame the contribution as a general COBOL modernization technique; the manuscript should either demonstrate transfer to at least one additional target or substantially qualify the scope claim so that readers do not over-generalize the 12% lift.
minor comments (5)
  1. [Header / ACM Reference Format] Throughout: author placeholder “Trovato et al.” and ACM copyright year 2018 remain; replace with the actual author list and current year.
  2. [Section 3.3.1] Section 3.3.1 title and body: “GunCOBOL” / “promgramming” are typos; correct to GnuCOBOL / programming.
  3. [Table 6 / Section 6] Table 6 “Failed” rows for SEDCoT remain high (73–82%); a short discussion of residual failure modes (beyond the two case studies in Section 6) would help readers gauge remaining risk.
  4. [Figure 2] Figure 2 pipeline diagram is helpful but the numbering of steps (μ, κ, λ …) is hard to map onto the textual Phase I–III description; align labels.
  5. [Section 5.5] Section 5.5 readability study uses only 10 programs for human raters; report inter-rater agreement (e.g., Krippendorff’s α) and confidence intervals.

Circularity Check

1 steps flagged

No load-bearing circular derivation; dual use of GnuCOBOL as IR source and evaluation oracle is a validity choice, not a by-construction result.

specific steps
  1. other [Section 4.2 (oracle) + Section 3.3.1 Eq. 5 (test generation)]
    "GnuCOBOL therefore acts as the authoritative execution standard. ... the rule-based compiler translates COBOL programs into functionally equivalent intermediate code ... The symbolic execution engine S is then applied to I_C to generate the test cases: T = S(I_C, X)"

    GnuCOBOL supplies both the intermediate C that KLEE explores for repair tests and the sole runtime oracle for the held-out golden suite. This is a shared semantic model, not a mathematical identity that forces the ≥12% lift; baselines face the same oracle and the golden suite is independently mutated. Mild dual-use only.

full rationale

SEDCoT is an empirical systems paper whose central claim (Table 5: ≥12.2% average golden-test pass-ratio lift over UniTrans/HRJR) is a comparative measurement under a single, openly declared oracle, not a first-principles derivation. Correctness is defined as matching GnuCOBOL outputs on a held-out suite of ~500 mutated inputs (Section 4.2); GnuCOBOL is set to 1.000 by that definition. Repair-time tests are generated by running KLEE on GnuCOBOL’s intermediate C (Eq. 5, Section 3.3.1) plus LLM tests, with ground-truth outputs also taken from GnuCOBOL. That dual role creates a mild methodological dependence, but it does not force the reported lift: the golden suite is held-out and generated by independent mutation strategies, baselines are scored under the identical oracle, and no free parameter is fitted and then re-labeled a prediction. There are no self-citation uniqueness theorems, no ansatz smuggled via prior author work, and no equation that reduces to its own inputs by construction. The dual-use concern is properly a threat to external validity (already noted in Section 8 and by the paper’s own NIST-pass-rate caveat), not circularity of the derivation chain. Score 1 reflects only that minor dual-use observation; the comparative claim remains independently measured.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The central claim rests on a small set of engineering choices (repair budget, KLEE resource caps, synthetic mutation strategies) and on the domain assumption that GnuCOBOL is an adequate semantic oracle. No new physical or mathematical entities are postulated; the free parameters are ordinary hyper-parameters of the experimental protocol.

free parameters (3)
  • max_repair_iterations = 3
    Fixed at 3 (2 ordinary + 1 delta-debug round) to match prior baselines; the performance curves in Table 9 show that the final gain is sensitive to this budget.
  • KLEE_resource_limits = 80M expr / 20 min
    80 M expressions, 20 min total, 2 min/state, 5 min coverage; chosen to keep generation tractable on a laptop and directly affect which programs receive symbolic tests.
  • golden_suite_size_and_mutation_mix = 500 (5×100)
    500 inputs per program, 100 from each of five mutation strategies; the pass-rate metric is defined with respect to this suite.
axioms (3)
  • domain assumption GnuCOBOL’s runtime behavior on the filtered CodeNet programs is a faithful ground-truth oracle for COBOL semantics.
    Stated in §4.2; every correctness number is computed against this oracle.
  • domain assumption Temperature-0 decoding plus a fixed three-round repair budget yields stable, comparable results across models.
    Implementation details §4.4; used to eliminate sampling variance.
  • ad hoc to paper Symbolic execution on the GnuCOBOL intermediate C plus LLM-generated tests together provide sufficient coverage to expose the majority of semantic bugs.
    Justified by the coverage and unique-bug counts in RQ3, but not independently verified outside the paper’s suite.
invented entities (1)
  • SEDCoT three-phase pipeline no independent evidence
    purpose: Orchestrates LLM translation, hybrid test generation, and delayed delta-debug repair for COBOL-to-C.
    The pipeline itself is the methodological contribution; it has no existence independent of the paper’s experiments.

pith-pipeline@v1.1.0-grok45 · 27383 in / 2761 out tokens · 29645 ms · 2026-07-11T21:44:48.111070+00:00 · methodology

0 comments
read the original abstract

COBOL remains critical across banking, insurance, and government infrastructure. However, maintenance is increasingly challenging due to outdated technologies, sparse documentation, and developer retirement, necessitating code translation into modern languages like C. Traditional rule-based transcompilers yield outputs that are difficult to read and maintain, while general-purpose large language models (LLMs) achieve suboptimal correctness because COBOL is a low-resource language with distinct logic patterns. To bridge this gap, we propose SEDCoT, a novel COBOL-to-C translation framework. SEDCoT first leverages LLMs for initial translation, then combines symbolic execution with LLM guidance to generate test suites and iteratively repair semantic discrepancies. Finally, it integrates delta debugging to minimize failing tests into succinct counterexamples, accelerating automated code repair. Evaluating SEDCoT on a public COBOL-to-C dataset demonstrates that it outperforms state-of-the-art baselines by at least 12% while producing translations with substantially higher readability than rule-based alternatives.

Figures

Figures reproduced from arXiv: 2607.04092 by Alexander Knapp, Chunyang Chen, Phillip Entin, Wenchao Gu.

Figure 1
Figure 1. Figure 1: Example of delta debugging: a string containing [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of our test-driven COBOL-to-C translation refinement pipeline. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of Bug-Triggering Effectiveness: Symbolic Execution vs. LLM-Generated Test Cases. [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Correlation between translated code correctness and readability (The X-axis denotes the pass rate of [PITH_FULL_IMAGE:figures/full_fig_p017_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

64 extracted references · 5 canonical work pages

  1. [1]

    Juan Altmayer Pizzorno and Emery D Berger. 2025. CoverUp: Effective High Coverage Test Generation for Python. Proceedings of the ACM on Software Engineering2, FSE (2025), 2897–2919

  2. [2]

    Joshua Bailey and Charles Nicholas. 2025. Symbolic Execution in Practice: A Survey of Applications in Vulnerability, Malware, Firmware, and Protocol Analysis.arXiv preprint arXiv:2508.06643(2025)

  3. [3]

    David M. Beazley. 1996. SWIG: an easy to use tool for integrating scripting languages with C and C++. InProceedings of the 4th Conference on USENIX Tcl/Tk Workshop, 1996 - Volume 4(Monterey, California)(TCLTK’96). USENIX Association, USA, 15

  4. [4]

    Seshia, and Alvin Cheung

    Sahil Bhatia, Jie Qiu, Niranjan Hasabnis, Sanjit A. Seshia, and Alvin Cheung. 2024. Verified Code Transpilation with LLMs. InAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, , Vol. 1, No. 1, Article . Publication date: July 2018. SEDCoT: Enhancing LLM-Based COBOL Code Translation via Symbolic Exec...

  5. [5]

    Veliswa Boya and Didier Durand. 2021. Serverless COBOL: Rejuvenating Legacy Code with Open Source Soft- ware. https://aws.amazon.com/de/blogs/opensource/serverless-cobol-rejuvenating-legacy-code-with-open-source- software/ Accessed: 2026-05-21

  6. [6]

    Raymond P. L. Buse and Westley Weimer. 2010. Learning a Metric for Code Readability.IEEE Trans. Software Eng.36, 4 (2010), 546–558. doi:10.1109/TSE.2009.70

  7. [7]

    Cristian Cadar, Daniel Dunbar, Dawson R Engler, et al. 2008. Klee: unassisted and automatic generation of high-coverage tests for complex systems programs.. InOSDI, Vol. 8. 209–224

  8. [8]

    2017.COBOL Is Everywhere

    David Cassel. 2017.COBOL Is Everywhere. Who Will Maintain It?Retrieved July 8, 2025 from https://thenewstack.io/ cobol-everywhere-will-maintain/, archived at [https://web.archive.org/web/20250612083600/https://thenewstack.io/ cobol-everywhere-will-maintain/]

  9. [9]

    Le Chen, Bin Lei, Dunzhi Zhou, Pei-Hung Lin, Chunhua Liao, Caiwen Ding, and Ali Jannesari. 2024. Fortran2CPP: Automating Fortran-to-C++ Translation using LLMs via Multi-Turn Dialogue and Dual-Agent Integration.arXiv preprint arXiv:2412.19770(2024)

  10. [10]

    Lowden, and Simon Sobisch

    Gary Cutler, Vincent Coen, Brian Tiffin, Bill Klein, László Erdős, Arnold Trembley, Edward Hart, Ron Norman, James K. Lowden, and Simon Sobisch. 2020.GnuCOBOL Programmer’s Guide. Retrieved July 9, 2025 from https: //gnucobol.sourceforge.io/HTML/gnucobpg.html, archived at [https://web.archive.org/web/20250626160559/https: //gnucobol.sourceforge.io/HTML/gnu...

  11. [11]

    Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, and Michel C Desmarais. 2024. Effective test generation using pre-trained large language models and mutation testing.Information and Software Technology 171 (2024), 107468

  12. [12]

    Colin Diggs, Michael Doyle, Amit Madan, Eric O Scott, Emily Escamilla, Jacob Zimmer, Naveed Nekoo, Paul Ursino, Michael Bartholf, Zachary Robin, et al. 2025. Leveraging LLMs for Legacy Code Modernization: Evaluation of LLM- Generated Documentation. In2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code). IEEE, 177–184

  13. [13]

    Min, Gail Kaiser, Junfeng Yang, and Baishakhi Ray

    Yangruibo Ding, Jinjun Peng, Marcus J. Min, Gail Kaiser, Junfeng Yang, and Baishakhi Ray. 2024. SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning. InAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Associates, Inc., 6027...

  14. [14]

    Sarah Fakhoury, Devjeet Roy, Adnan Hassan, and Vernera Arnaoudova. 2019. Improving Source Code Readability: Theory and Practice. In2019 IEEE/ACM 27th International Conference on Program Comprehension (ICPC). 2–12. doi:10. 1109/ICPC.2019.00014

  15. [15]

    Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. 2023. Automated repair of programs from large language models. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1469–1481

  16. [16]

    Shmulik Froimovich, Raviv Gal, Wesam Ibraheem, and Avi Ziv. 2025. Quality Evaluation of COBOL to Java Code Transformation. arXiv:2507.23356 [cs.SE] https://arxiv.org/abs/2507.23356

  17. [17]

    Shubham Gandhi, Manasi Patwardhan, Jyotsana Khatri, Lovekesh Vig, and Raveendra Kumar Medicherla. 2024. Translation of low-resource COBOL to logically correct and readable Java leveraging high-resource Java refinement. In Proceedings of the 1st International Workshop on Large Language Models for Code. 46–53. doi:10.1145/3643795.3648388

  18. [18]

    Samat Gaynutdinov, Saveliy Grigoryev, Pavel Iatchenii, Elena Ilina, Dmitry Ivanov, Vladislav Kalugin, Aleksei Pleshakov, Pavel Ponomarev, Konstantin Rybkin, Svetlana Shmidt, Vadim Volodin, and Alexey Utkin. 2022. Pre- sentation: UTBot Simplifies Auto Test Generation. https://www.utbot.org/static/KLEE_workshop2022_abstract- 9591232a9941df34577a134609dbbe29...

  19. [19]

    2025.GnuCOBOL - A free COBOL compiler

    Bernard Giroud, Brian Tiffin, Keisuke Nishida, Simon Sobisch, and Roger While. 2025.GnuCOBOL - A free COBOL compiler. Retrieved August 25, 2025 from https://sourceforge.net/projects/gnucobol/, archived at [https://web.archive. org/web/20250825030006/https://sourceforge.net/projects/gnucobol/]

  20. [20]

    Jingxuan He, Gishor Sivanrupan, Petar Tsankov, and Martin Vechev. 2021. Learning to Explore Paths for Symbolic Execution. InProceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security(Virtual Event, Republic of Korea)(CCS ’21). Association for Computing Machinery, New York, NY, USA, 2526–2540. doi:10. 1145/3460120.3484813

  21. [21]

    Ali Reza Ibrahimzada, Kaiyao Ke, Mrigank Pawagi, Muhammad Salman Abid, Rangeet Pan, Saurabh Sinha, and Reyhaneh Jabbarvand. 2025. AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation.Proceedings of the ACM on Software Engineering2, FSE (2025), 2454–2476. , Vol. 1, No. 1, Article . Publication date: July ...

  22. [22]

    International Organization for Standardization. 2023. ISO/IEC 1989:2023 - Information technology - Programming languages, their environments and system software interfaces - Programming language COBOL. https://www.iso.org/ standard/74527.html Accessed: 2026-05-21

  23. [23]

    Prithwish Jana, Piyush Jha, Haoyang Ju, Gautham Kishore, Aryan Mahajan, and Vijay Ganesh. 2023. Cotran: An llm-based code translator using reinforcement learning with feedback from compiler and symbolic execution.arXiv preprint arXiv:2306.06755(2023)

  24. [24]

    Mingsheng Jiao, Tingrui Yu, Xuan Li, Guanjie Qiu, Xiaodong Gu, and Beijun Shen. 2023. On the Evaluation of Neural Code Translation: Taxonomy and Benchmark. arXiv:2308.08961 [cs.SE] https://arxiv.org/abs/2308.08961

  25. [25]

    John Johnson, Sergio Lubo, Nishitha Yedla, Jairo Aponte, and Bonita Sharif. 2019. An Empirical Study Assessing Source Code Readability in Comprehension. In2019 IEEE International Conference on Software Maintenance and Evolution (ICSME). 513–523. doi:10.1109/ICSME.2019.00085

  26. [26]

    Svetoslav Karaivanov, Veselin Raychev, and Martin Vechev. 2014. Phrase-Based Statistical Translation of Programming Languages. InProceedings of the 2014 ACM International Symposium on New Ideas, New Paradigms, and Reflections on Programming & Software(Portland, Oregon, USA)(Onward! 2014). Association for Computing Machinery, New York, NY, USA, 173–184. do...

  27. [28]

    Atul Kumar, Diptikalyan Saha, Toshiaki Yasue, Kohichi Ono, Saravanan Krishnan, Sandeep Hans, Fumiko Satoh, Gerald Mitchell, and Sachin Kumar. 2024. Automated Validation of COBOL to Java Transformation. InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 2415–2418. doi:10.1145/3691620.3695365

  28. [29]

    K. Lano, P. T. Breuer, and H. Haughton. 1993. Reverse-engineering cobol via formal methods.Journal of Software Main- tenance: Research and Practice5, 1 (1993), 13–35. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/smr.4360050103 doi:10.1002/smr.4360050103

  29. [30]

    Kevin Lano and Hanan Siala. 2024. Using model-driven engineering to automate software language translation. Automated Software Engineering31, 1 (Feb. 2024). doi:10.1007/s10515-024-00419-y

  30. [31]

    Fangjian Lei, Jiawen Liu, Shayan Noei, Ying Zou, Derek Truong, and William Alexander. 2025. Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models.arXiv preprint arXiv:2507.02182(2025)

  31. [32]

    Caroline Lemieux, Jeevana Priya Inala, Shuvendu K Lahiri, and Siddhartha Sen. 2023. Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models. In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 919–931

  32. [33]

    Stephan Lukasczyk and Gordon Fraser. 2022. Pynguin: Automated unit test generation for python. InProceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings. 168–172

  33. [34]

    Wenqiang Luo, Jacky Wai Keung, Boyang Yang, Jacques Klein, Tegawende F Bissyande, Haoye Tian, and Bach Le. 2025. Unlocking LLM Repair Capabilities in Low-Resource Programming Languages Through Cross-Language Translation and Multi-Agent Refinement.arXiv preprint arXiv:2503.22512(2025)

  34. [35]

    Marcos Macedo, Yuan Tian, Pengyu Nie, Filipe R Cogo, and Bram Adams. 2024. InterTrans: Leveraging transitive intermediate translations to enhance LLM-based code translation.arXiv preprint arXiv:2411.01063(2024)

  35. [36]

    1999.COBOL Test Suites

    Carmelo Montanez-Rivera. 1999.COBOL Test Suites. Retrieved August 25, 2025 from https://www.itl.nist.gov/div897/ ctg/cobol_form.htm, archived at [https://web.archive.org/web/20230917224831/https://www.itl.nist.gov/div897/ctg/ cobol_form.htm]

  36. [37]

    Vikram Nitin. 2024. Using AI to Automate the Modernization of Legacy Software Applications. InProceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering(Sacramento, CA, USA)(ASE ’24). Association for Computing Machinery, New York, NY, USA, 2514–2517. doi:10.1145/3691620.3695610

  37. [39]

    OpenRouter. 2026. OpenRouter. https://openrouter.ai/ Accessed: 2026-05-21

  38. [41]

    2024.Qwen2.5 Coder 32B Instruct

    Inc OpenRouter. 2024.Qwen2.5 Coder 32B Instruct. Retrieved August 5, 2025 from https://openrouter.ai/qwen/qwen- 2.5-coder-32b-instruct, archived at [https://web.archive.org/web/20250709025049/https://openrouter.ai/qwen/qwen- 2.5-coder-32b-instruct]

  39. [42]

    2025.Google: Gemma 3 27B

    Inc OpenRouter. 2025.Google: Gemma 3 27B. Retrieved August 5, 2025 from https://openrouter.ai/google/gemma-3- 27b-it, archived at [https://web.archive.org/web/20250719132359/https://openrouter.ai/google/gemma-3-27b-it]

  40. [43]

    Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and Reyhaneh Jabbarvand. 2024. Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code. InProceedings of the IEEE/ACM 46th International Conference on , Vol. 1, No. 1, A...

  41. [44]

    Annibale Panichella, Fitsum Meshesha Kifetew, and Paolo Tonella. 2015. Reformulating branch coverage as a many- objective optimization problem. In2015 IEEE 8th international conference on software testing, verification and validation (ICST). IEEE, 1–10

  42. [45]

    Annibale Panichella, Fitsum Meshesha Kifetew, and Paolo Tonella. 2017. Automated test case generation as a many- objective optimisation problem with dynamic selection of the targets.IEEE Transactions on Software Engineering44, 2 (2017), 122–158

  43. [46]

    Sebastian Poeplau and Aurélien Francillon. 2020. Symbolic execution with {SymCC}: Don’t interpret, compile!. In 29th USENIX Security Symposium (USENIX Security 20). 181–198

  44. [47]

    Daryl Posnett, Abram Hindle, and Premkumar T. Devanbu. 2011. A simpler model of software readability. InProceedings of the 8th International Working Conference on Mining Software Repositories, MSR 2011 (Co-located with ICSE), Waikiki, Honolulu, HI, USA, May 21-28, 2011, Proceedings, Arie van Deursen, Tao Xie, and Thomas Zimmermann (Eds.). ACM, 73–82. doi:...

  45. [48]

    Asha Rajbhoj, Akanksha Somase, Tanay Sant, Ajim Pathan, Purvesh Doud, and Vinay Kulkarni. 2025. Leveraging LLM for software modernization: COBOL Functionality Extraction Case study. In2025 40th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW). 14–21. doi:10.1109/ASEW67777.2025.00012

  46. [49]

    Nishath Rajiv Ranasinghe, Shawn M Jones, Michal Kucer, Ayan Biswas, Daniel O’Malley, Alexander Buschmann Most, Selma Liliane Wanna, and Ajay Sreekumar. 2025. LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study.arXiv preprint arXiv:2504.15424(2025)

  47. [50]

    Matthew Renze. 2024. The effect of sampling temperature on problem solving in large language models. InFindings of the association for computational linguistics: EMNLP 2024. 7346–7356

  48. [51]

    Reuters. [n. d.].COBOL blues. Retrieved August 5, 2025 from https://www.reuters.com/graphics/USA-BANKS-COBOL/ 010040KH18J/, archived at [https://web.archive.org/web/20250726190454/https://www.reuters.com/graphics/USA- BANKS-COBOL/010040KH18J/]

  49. [52]

    Baptiste Roziere, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020. Unsupervised translation of programming languages.Advances in neural information processing systems33 (2020), 20601–20611

  50. [53]

    2025.opensource COBOL 4J

    Yutaro Sakamoto. 2025.opensource COBOL 4J. https://github.com/opensourcecobol/opensourcecobol4j

  51. [54]

    Marina Sakharova, Abhinav Anand, and Mira Mezini. 2025. Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs.arXiv preprint arXiv:2504.15210(2025)

  52. [55]

    Agnia Sergeyuk, Olga Lvova, Sergey Titov, Anastasiia Serova, Farid Bagirov, and Timofey Bryksin. 2024. Assessing Consensus of Developers’ Views on Code Readability.CoRRabs/2407.03790 (2024). arXiv:2407.03790 doi:10.48550/ ARXIV.2407.03790

  53. [56]

    2020.USA suchen dringend Programmierer für uralte Behördensysteme

    Kathrin Stoll. 2020.USA suchen dringend Programmierer für uralte Behördensysteme. Retrieved August 5, 2025 from https://www.welt.de/wirtschaft/webwelt/article207536129/Cobol-USA-suchen-wegen-Coronakrise-Programmierer- fuer-alte-Systeme.html, archived at [https://web.archive.org/web/20240702212854/https://www.welt.de/wirtschaft/ webwelt/article207536129/Co...

  54. [57]

    Ekaterina Tochilina, Vyacheslav Tamarin, Dmitry Mordvinov, Valentyn Sobol, Sergey Pospelov, Alexey Menshutin, Yury Kamenev, and Dmitry Ivanov. 2024. UTBot Python at the SBFT Tool Competition 2024. InProceedings of the 17th ACM/IEEE International Workshop on Search-Based and Fuzz Testing (SBFT ’24). ACM, 41–42. doi:10.1145/3643659. 3643934

  55. [58]

    Ekaterina Tochilina, Vyacheslav Tamarin, Dmitry Mordvinov, Valentyn Sobol, Sergey Pospelov, Alexey Menshutin, Yury Kamenev, and Dmitry Ivanov. 2024. UTBot Python at the SBFT Tool Competition 2024. InProceedings of the 17th ACM/IEEE International Workshop on Search-Based and Fuzz Testing(Lisbon, Portugal)(SBFT ’24). Association for Computing Machinery, New...

  56. [59]

    Weixi Tong and Tianyi Zhang. 2024. Codejudge: Evaluating code generation with large language models.arXiv preprint arXiv:2410.02184(2024)

  57. [60]

    Fengcai Wen, Emad Aghajani, Csaba Nagy, Michele Lanza, and Gabriele Bavota. 2021. Siri, Write the Next Method. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). 138–149. doi:10.1109/ICSE43902.2021.00025

  58. [61]

    Chen Yang, Junjie Chen, Bin Lin, Jianyi Zhou, and Ziqi Wang. 2024. Enhancing llm-based test generation for hard-to- cover branches via program analysis.arXiv preprint arXiv:2404.04966(2024)

  59. [62]

    Zhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung, Jia Li, Shuo Liu, Yifan Hong, Xiaoxue Ma, Zhi Jin, and Ge Li

  60. [63]

    ACM Softw

    Exploring and Unleashing the Power of Large Language Models in Automated Code Translation.Proc. ACM Softw. Eng.1, FSE, Article 71 (July 2024), 24 pages. doi:10.1145/3660778

  61. [64]

    Zhiqiang Yuan, Weitong Chen, Hanlin Wang, Kai Yu, Xin Peng, and Yiling Lou. 2024. Transagent: An llm-based multi-agent system for code translation.arXiv preprint arXiv:2409.19894(2024)

  62. [65]

    Alon Zakai. 2011. Emscripten: an LLVM-to-JavaScript compiler. InProceedings of the ACM International Conference Companion on Object Oriented Programming Systems Languages and Applications Companion(Portland, Oregon, USA) , Vol. 1, No. 1, Article . Publication date: July 2018. 24 Trovato et al. (OOPSLA ’11). Association for Computing Machinery, New York, N...

  63. [66]

    Andreas Zeller. 1999. Yesterday, my program worked. Today, it does not. Why?ACM SIGSOFT Software engineering notes24, 6 (1999), 253–267

  64. [67]

    Terry Yue Zhuo. 2024. ICE-Score: Instructing Large Language Models to Evaluate Code. InFindings of the Association for Computational Linguistics: EACL 2024, Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian’s, Malta, 2232–2242. https://aclanthology.org/2024.findings-eacl.148/ Received 20 February 2007; revised ...