Pith. sign in

Paper Citation Record · LEDGER

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 4 inbound Pith citation observations for arXiv:2412.18989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18989 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:03.644079Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:35.374595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:21:10.908503Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact4
  • verified fuzzy32
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c8e02bf-26b7-4949-a787-60feb4091439 · outbound

This paper cites An empirical study on the usage of transformer models for code completion,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study An empirical study on the usage of transformer models for code completion,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.663359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.403178Z digest=sha256:5a2630573854e3ebcbb650601b30e9544445cb1ee79e9cfa8067a1b2ed978f6a

Observation 8e557a6a-917a-4907-9809-e78844a645d6 · outbound

This paper cites SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.409292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.409292Z digest=sha256:4107ccadbdda1dd689a2a50e70b786f490652c474d15cf3695a26c3f0b66ee24

Observation 3708045c-766f-4a85-9052-c70a950cc1c6 · outbound

This paper cites Toward deep learning software reposi- tories,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Toward deep learning software reposi- tories,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.646725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.415672Z digest=sha256:a2c9075f4cc92eb4579823a3480a09717babae4b0ce26996138b42f1cad34005

Observation 3a94e3f5-60d1-4571-97a2-eb2b061cd0d6 · outbound

This paper cites Few-shot training llms for project-specific code-summarization,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Few-shot training llms for project-specific code-summarization,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.631605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.420970Z digest=sha256:05f45ce15ddce1cbd0f283d1fa11603e44618e9b5e01d7efd02ed4811a0ccc32

Observation 49064042-8472-40b9-98c2-5b66006880b9 · outbound

This paper cites Inferfix: End-to-end program repair with llms,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Inferfix: End-to-end program repair with llms,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.610995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.425961Z digest=sha256:32bc2e8449ac8146be1410b106e34fbf01cc830fec560b54eed09afcaa5803b6

Observation 1feb2ec2-6ea3-41c4-99e9-b77bd00cb82c · outbound

This paper cites Deep learning code fragments for code clone detection,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Deep learning code fragments for code clone detection,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.591439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.430951Z digest=sha256:655e8f141f5b2eced2cefceaa8f0ec021d7a4bbddf4556653ae6ef191ce799b0

Observation 25084cca-18a8-4ac4-81ed-8bd501399e17 · outbound

This paper cites On learning meaningful assert statements for unit test cases,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study On learning meaningful assert statements for unit test cases,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.572962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.436960Z digest=sha256:19a46edd8ade3505e5292f678bf976611ba0eb0b485f8c160f7bd5f6bb2233f3

Observation bf8022e6-97f5-4fd6-b529-bb16756d5eb4 · outbound

This paper cites A systematic literature review on the use of deep learning in software engineering research,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study A systematic literature review on the use of deep learning in software engineering research,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.557003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.441520Z digest=sha256:fa819d2f870c0bc7dd8132fbc8673304d954ad4023295e5862d518ea0e7aca42

Observation d9daa52d-26e9-4ff6-969f-92aa962e5e51 · outbound

This paper cites When and Why Your Code Starts to Smell Bad (and Whether the Smells Go Away),.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study When and Why Your Code Starts to Smell Bad (and Whether the Smells Go Away),

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.538147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.445796Z digest=sha256:1bf71056f654ccbe94a8ec25dbf124b09f3f3ac76d51ea49f88a83be803d8983

Observation 1d358ee4-d47a-49a8-8f7e-611103c185d6 · outbound

This paper cites On the diffuseness and the impact on maintainability of code smells: a large scale empirical investigation,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study On the diffuseness and the impact on maintainability of code smells: a large scale empirical investigation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.518728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.450187Z digest=sha256:c6a5eaca002f5d097a25ffcc9fa3c88bd0db57a08b4e3f1c1c62fa77292f5145

Observation 08cafb2d-6691-48e3-8983-fca985485e3f · outbound

This paper cites Vulnerability Handling of AI-Generated Code -- Existing Solutions and Open Challenges.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Vulnerability Handling of AI-Generated Code -- Existing Solutions and Open Challenges

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-11T01:04:04.024541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.454292Z digest=sha256:1567fcede5d40d40d0fa37233ca8c05139f272a6f321c1fd5d8a8fc1bae2a3f9

Observation 9beb1c46-93eb-45af-ba9b-eb21d3080a0c · outbound

This paper cites CodeLMSec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code lan- guage models,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study CodeLMSec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code lan- guage models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.499930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.458850Z digest=sha256:1eb9e9584970158d5f77b76f1dd207348c71cfa4a7cfe112457d6d1445a853b5

Observation 604cae6f-c46c-4e2e-ab6e-d00da5a12fac · outbound

This paper cites SALLM: Security Assessment of Generated Code.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study SALLM: Security Assessment of Generated Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.463934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.463934Z digest=sha256:5d611fac4f95f398c3f266b7c5b5c612e05ea38504c5cfb9d84c0a11ba7a8777

Observation d75ffbc4-0343-4067-b2aa-57b852050647 · outbound

This paper cites DLAP: A Deep Learning Augmented Large Language Model Prompting Framework for Software Vulnerability Detection.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study DLAP: A Deep Learning Augmented Large Language Model Prompting Framework for Software Vulnerability Detection

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T01:04:03.985672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.468913Z digest=sha256:fdbeaeea7d45f237a1d23fec427d6983735288bc963f578011ab59c15f8fd704

Observation 616c364a-45f6-4e45-a6e0-8660d05b530d · outbound

This paper cites Evaluating Large Language Models in Detecting Test Smells.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Evaluating Large Language Models in Detecting Test Smells

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T01:04:03.960809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.474105Z digest=sha256:3d13cc4defa8f7ea2de837e40086dffbc577cb8d371cac152f164dc4e4b06354

Observation 88b79e18-58ca-4138-9e1e-ea1cbd336ac1 · outbound

This paper cites Code smell detection using hy- brid machine learning algorithms,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Code smell detection using hy- brid machine learning algorithms,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.478649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.479283Z digest=sha256:026cf8b96c00ecb307ce7b2d294971338e1790c55262ae2443badea71119b221

Observation fcebad28-2922-4413-a8d3-c80377a9a3ee · outbound

This paper cites Multi-Label Code Smell Detection with Hybrid Model based on Deep Learning,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Multi-Label Code Smell Detection with Hybrid Model based on Deep Learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.459092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.483872Z digest=sha256:e75e943c3289dd55d4ebe4283c2b50cac3a7389998b0bc55dd86682f9ce7f2d1

Observation afc4a3f8-d920-4833-a499-935fc97b243e · outbound

This paper cites iSMELL: Assembling LLMs with Expert Toolsets for Code Smell Detection and Refactoring,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study iSMELL: Assembling LLMs with Expert Toolsets for Code Smell Detection and Refactoring,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.438375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.489197Z digest=sha256:7ceef524bc762918a0a054043de25e5594c6a0a20449cca3ee6a589bd180649b

Observation 872550b3-bded-4f1f-8a79-fee044d9b43f · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study BLEU: a method for automatic evaluation of machine translation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.417579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.493760Z digest=sha256:1dadf08266ccde32a37fd323d16f9b6eac7926c9f57dda244e7780349bc1cc9a

Observation 31fd5b27-340c-49fd-87e2-4660d7410cbd · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.498305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.498305Z digest=sha256:44228c197c4c2f412b120f641087eb22f107c4fcae6dc6179c51175ed7d745c7

Observation 535866fd-a9b0-41fc-807f-5ff433ee9696 · outbound

This paper cites ROUGE: A Package for Automatic Evaluation of Sum- maries,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study ROUGE: A Package for Automatic Evaluation of Sum- maries,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.400047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.503459Z digest=sha256:fa5efce86b059a9a6957d8819fcb68b553209eda4783e5ba0506814f4c2188c0

Observation e8b9eec1-90af-4a1c-b751-d4012a64dd18 · outbound

This paper cites Meteor: an automatic metric for mt evaluation with high levels of correlation with human judgments,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Meteor: an automatic metric for mt evaluation with high levels of correlation with human judgments,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.380061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.508403Z digest=sha256:fc0927dee2ac722262039ffbfbca5758ab74ad07e07e0169a27da7324f5f53a6

Observation d58e88c7-0c18-41e5-9ac3-51fec2e21e34 · outbound

This paper cites A systematic literature review on the use of deep learning in software engineering research,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study A systematic literature review on the use of deep learning in software engineering research,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.360462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.513445Z digest=sha256:c9e038c0703e49d8c91d3b65990eb31d07e5b358047e267b1ed134c1940f2f0f

Observation f91b7c1d-0f8a-463d-b121-c915bd08c090 · outbound

This paper cites Towards More Trust- worthy and Interpretable LLMs for Code through Syntax-Grounded Explanations,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Towards More Trust- worthy and Interpretable LLMs for Code through Syntax-Grounded Explanations,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.518465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.518465Z digest=sha256:0cf1e5aecbe16fc1a083063942d3156f81bfc45fb6b089d5c25a6a7901c0f06a

Observation c6271d76-e3a4-4208-b32b-886d0b184242 · outbound

This paper cites Which syntactic capabilities are statistically learned by masked language models for code?.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Which syntactic capabilities are statistically learned by masked language models for code?

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.341291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.523362Z digest=sha256:a4854be946d9924d0f49ff2fbeee612ab8a232d95e2514361781272538aaae6e

Observation d843fc25-c085-40c5-a4bb-2e7e0b3b0dae · outbound

This paper cites Codesmells,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Codesmells,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.322889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.528834Z digest=sha256:ebca29ba5dcef92d4a6a024bfcdd61ce0b701889de770fa75124a2541aeb0c7c

Observation 8dcf6c74-f5d1-493f-9e93-defc121d301c · outbound

This paper cites Visualizing and Understanding Recurrent Networks.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Visualizing and Understanding Recurrent Networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.533227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.533227Z digest=sha256:1bc1abda558a6daa477a4ef374e747d39794267508819de29024c531e04e1333

Observation 53a9e426-9a90-445d-872d-308cadaa0d49 · outbound

This paper cites an unresolved cited work.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T01:04:04.291678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.537843Z digest=sha256:81dfcae3d1e605271d08bdc24fbbdb7aed1073d2472e37f63d5badfd7057b562

Observation 59a871b3-bfc3-46ff-82a7-213bc33082dd · outbound

This paper cites Benchmarking causal study to interpret large language models for source code,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Benchmarking causal study to interpret large language models for source code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.274862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.544720Z digest=sha256:9f2a691bb0c53745587c46fd2567707e2819f201a7b0c0036e67e26dfbed9682

Observation 75a4ba42-b2fc-46b5-b3e8-8a138d00d3b3 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Code Llama: Open Foundation Models for Code

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.549799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.549799Z digest=sha256:72018096788b169e26188714893fe9c2daca1166bc29bc33beccecafdf6b7e0f

Observation f8e76506-26e5-45db-bf6d-e992454b5e6d · outbound

This paper cites Mistral 7B.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Mistral 7B

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.556840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.556840Z digest=sha256:374bfef9e87dd6476e60c8f7de958209c667d9c311131732af84bd16f99d748f

Observation 7b08d064-358f-46ec-92d7-ce1252d348c2 · outbound

This paper cites Code smell detection using hy- brid machine learning algorithms,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Code smell detection using hy- brid machine learning algorithms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.256271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.562787Z digest=sha256:4c86d028ab4271fe0b3c6088efc65109ccc73f138322869462beeb8ef84d7306

Observation 6c9b0ac3-f9b6-4aa7-8564-e1f76c1ae49b · outbound

This paper cites Machine learning powered code smell de- tection as a business improvement tool,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Machine learning powered code smell de- tection as a business improvement tool,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.239837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.568264Z digest=sha256:a73d118dc865b33c2fb73a42057a0e818b9cf143f1a394dc6fceabd2d617fda9

Observation c1e00833-c93b-4aa2-b548-6e61aa40d5f3 · outbound

This paper cites Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMs.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:03.574583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:03.574583Z digest=sha256:059e6eb531e03d7994b7324a01baef73bb4c49481192fff3ce36e456aa91ce9b

Observation b746481d-dcd9-4cf4-8980-9f7b8922378e · outbound

This paper cites A Systematic Literature Review on the Code Smells Datasets and Validation Mechanisms,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study A Systematic Literature Review on the Code Smells Datasets and Validation Mechanisms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.222541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.580778Z digest=sha256:609434f9b412139394fc4c28531c51f237ec7386a6bb7b69f5ade5f727133c46

Observation 66a55f9b-506f-41dd-a668-36d8ef0acdc9 · outbound

This paper cites ml-Codesmell: A code smell prediction dataset for machine learning approaches,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study ml-Codesmell: A code smell prediction dataset for machine learning approaches,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.204730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.589571Z digest=sha256:031ae66d341ed3b477d7d2c5221bc43cf86f42e9b5ed5df3e3c2092fc990ea1b

Observation 912c9ca7-7a9b-4cf9-b82d-db08d4c33db1 · outbound

This paper cites Mlcq: Industry-relevant code smell data set,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Mlcq: Industry-relevant code smell data set,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.187375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.596066Z digest=sha256:339c4b7aed7f9c1b48d0b89ac86f4c6ce905c31e64e14fc3fde9c0aa1a09f6e0

Observation 2538c857-0072-43b6-91b9-45b9acff69f0 · outbound

This paper cites Evaluating the accuracy of machine learning algorithms on detecting code smells for different developers,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Evaluating the accuracy of machine learning algorithms on detecting code smells for different developers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.171338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.601761Z digest=sha256:b4f5613461b6d5eb3cade80a57c199ef5a9c922121d7209f9d6c075758c082a4

Observation fcb823cf-1ffc-4338-a32d-292805d193a8 · outbound

This paper cites DACOS—a manually annotated dataset of code smells,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study DACOS—a manually annotated dataset of code smells,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.153186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.606679Z digest=sha256:7a6e332289405b6c9a28e0ec64bc606d1fb30df28dd2a52dc98d58e593bff032

Observation 912d416b-9b36-40d1-a474-3e8579c501c2 · outbound

This paper cites The Technical Debt Dataset,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study The Technical Debt Dataset,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.129712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.612873Z digest=sha256:093f58351f330eb91a1e94afe3fcb0bdd55a8c59d369e004e836dc611f63df47

Observation 69b60e96-dbfd-4169-8031-fa9fb4039542 · outbound

This paper cites Using code evolution information to improve the quality of labels in code smell datasets,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Using code evolution information to improve the quality of labels in code smell datasets,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.110103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.618462Z digest=sha256:e7aff3d6bffae3586b5210e7745436c2a6b08a6c185b20a6633481647b752650

Observation 8b65d380-44ba-4d29-8582-1a4e1e7ec806 · outbound

This paper cites Prompt Learning for Multi-Label Code Smell Detection: A Promising Approach.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Prompt Learning for Multi-Label Code Smell Detection: A Promising Approach

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-11T01:04:03.704852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.625710Z digest=sha256:2588cccdb204a00983e5291f7da4c09fd9186d4708aa505de8691b6adc13b15e

Observation 944f827f-5e2a-4dfb-bb44-64a276c1f7d0 · outbound

This paper cites Toward a theory of causation for interpreting neural code models,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study Toward a theory of causation for interpreting neural code models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.093037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.632997Z digest=sha256:60693e1272b4ac3fce4d0f161f27d00208b8f373144cab697c42f30a6c6e76af

Observation 1ae8623c-9850-465a-ad51-b3071431b412 · outbound

This paper cites ”why should i trust you?.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study ”why should i trust you?

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.076285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.638043Z digest=sha256:78398f10c37525b642dcb280ed20b9f54c08655852496340de28f4ff43c87b40

Observation 0ca3de16-dcf5-4eba-9bfa-cba00155576b · outbound

This paper cites A unified approach to interpreting model predictions,.

How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study A unified approach to interpreting model predictions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:04.059394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T01:04:03.644079Z digest=sha256:75a818c8b9b32de4e3ce0a074128191ac1d52210ef3db73553b46fd0a12c1285

Pith citing papers

Observation 9e665133-d85c-4869-ae5b-c4ded1d91d73 · inbound

Optimizing Token Consumption in LLMs: A Nano Surge Approach for Code Reasoning Efficiency cites this paper.

Optimizing Token Consumption in LLMs: A Nano Surge Approach for Code Reasoning Efficiency How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:35.374595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:17:35.374595Z digest=sha256:c9961380b31189fe2faadd64980cdbde03c4050769551a94cd17027625a42a34

Observation c8d06b7e-9688-49bb-bc04-b9eb0b08fdf5 · inbound

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 cites this paper.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.402491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.402491Z digest=sha256:982f185f0eaa05164c182607d2407c330e8bf51853785196ee774c7f63102b1d

Observation 75937992-3c33-4279-abbe-2724dc63252a · inbound

A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code cites this paper.

A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T06:49:10.487496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:49:10.487496Z digest=sha256:bbc3c7fc1eca75544fc88d29c3814f6e7b1d9daeda619f68d99c31f9444807f9

Observation c4be3898-66d8-4c11-9a73-d9fa90977d3d · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:10.911009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:f56237c75824d6b762b04c5ddb597d908b99124bbb62276feecb6197ff9fcbe4