Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 3 inbound Pith citation observations for arXiv:2506.07594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07594 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:41.214905Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T16:24:25.357338Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 23247134-ec7a-462b-b92b-b0da15036b50 · outbound

This paper cites On the relation of test smells to software code quality,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study On the relation of test smells to software code quality,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.756482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.943703Z digest=sha256:473fe39e8fdbe35594be1a851f12feb98589ff025dd395c08ba65b5f90db1939

Observation 8a415b7c-ea53-4703-807f-b2a336d20b4f · outbound

This paper cites On the diffusion of test smells in automatically generated test code: An empirical study,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study On the diffusion of test smells in automatically generated test code: An empirical study,

Reference 2

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:35:42.283784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.949345Z digest=sha256:ea04737390a79882ae23e932ab2440a1c6a88d779b0699c22457e80396d765df

Observation db152661-eaa6-4851-905d-3fa62b9f4660 · outbound

This paper cites When and why your code starts to smell bad,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study When and why your code starts to smell bad,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.741730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.954889Z digest=sha256:f67039f1e07284e7ef7621a66f792650d4178a8ef337134c3f5d2460297ad247

Observation 688e61d6-4cdb-49a5-a2de-f687eabf87dc · outbound

This paper cites An empirical analysis of the distribution of unit test smells and their impact on software maintenance,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical analysis of the distribution of unit test smells and their impact on software maintenance,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.726187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.960193Z digest=sha256:937c940107ce224791a26637726326f556b56f4ad9472df36502edfbb0f4e024

Observation 568dc405-c62c-401c-b339-7d00346a4871 · outbound

This paper cites Just-in-time test smell detection and refactoring: The darts project,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Just-in-time test smell detection and refactoring: The darts project,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.966381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.966381Z digest=sha256:542c5c60197dfe33456dd8d91ec7d34c33887fa535eb76f1ccb6eaddb905dfe5

Observation 1d6ec493-2c87-4666-9895-45d04d52dc64 · outbound

This paper cites GPT-4 Technical Report.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.971440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.971440Z digest=sha256:160926956ec2fa651ebc9c3f43a1c6a297bb83c17da8357cc06a894087dbdf0a

Observation 748b762e-2e35-4b29-a32f-bc2578dbe139 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.977343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.977343Z digest=sha256:908510c6635c37389542ec6eb4a72a49633c9a4959cf609109c4d8cc7652b614

Observation aea9bf46-fe6a-4274-9c69-d0eb6c0509fe · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.982398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.982398Z digest=sha256:d77e7dd636f296f62086a0e1d2d6f70047658e9b357f837b4fd6ebc0e8fdd4e2

Observation 08a1ec8d-69e3-4d06-b149-1248e6a9c634 · outbound

This paper cites Codebert: A pre-trained model for programming and natural languages,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Codebert: A pre-trained model for programming and natural languages,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.710348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.987859Z digest=sha256:36f22e92d87f0b68f1a605c1613a7f4ac9f40f65b59477c5f814233ff56fe1e4

Observation 37b2e45a-2e62-4589-a29e-30195993cc57 · outbound

This paper cites Codexglue: A machine learning benchmark dataset for code understanding and generation,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Codexglue: A machine learning benchmark dataset for code understanding and generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.692702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.992839Z digest=sha256:f5ec3f4602c17be86f44c4eaf7111a745b6f16fd72602e455c4d8ddcb7200e3a

Observation 84c2c9f9-5d92-4b0f-858b-61b0e668b247 · outbound

This paper cites Top programming languages - the state of the octoverse 2022,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Top programming languages - the state of the octoverse 2022,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.677548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:40.998380Z digest=sha256:86cc43db5252a04fad9dee33df4cb46889a5852ea3f5a03a72d22b2d333f6116

Observation fe91d1d0-56e7-4e29-8824-7615b7467435 · outbound

This paper cites Utilization of pre-trained language model for adapter-based knowledge transfer in software engineering,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Utilization of pre-trained language model for adapter-based knowledge transfer in software engineering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.660795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.003516Z digest=sha256:d17c12456829c3621408f2228a69c068385cfdf37a17dd4e8c05634cce51cca3

Observation 11ba22fa-d781-4eb2-b61f-2334d7482892 · outbound

This paper cites To- wards efficient fine-tuning of pre-trained code models: An experimental study and beyond,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study To- wards efficient fine-tuning of pre-trained code models: An experimental study and beyond,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.645541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.008583Z digest=sha256:c151ba6b85172d6c9fd0071eae704cf736d91e7651b2dbc02b23ecc7ce105934

Observation d2584707-6a9a-40f0-b8f8-2338e8f740ac · outbound

This paper cites An empirical comparison of pre-trained models of source code,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical comparison of pre-trained models of source code,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.013423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.013423Z digest=sha256:5732dbddd560eba49c1f50ef7f395a550e8ce24b5213a084dd0e2305f9b6cced

Observation ab33be99-7a39-4325-9b52-4b05844df9d9 · outbound

This paper cites (2024) Testsmellsrefactoringbyllms.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study (2024) Testsmellsrefactoringbyllms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.629396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.018130Z digest=sha256:313934bc9f0446ebe0fd952f8a89413a81a337334efe775aeca72eaad19da07c

Observation abc5f614-3046-456b-a24d-e7157091aecb · outbound

This paper cites Large language models for software engineering: A systematic literature review,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Large language models for software engineering: A systematic literature review,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.023442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.023442Z digest=sha256:9ad1a07a7cb6ca4893c6a4571d9a7ae8884014986b14a69e627614e8b0b86ba2

Observation 9007f747-60b9-4a80-9dc6-1896b7c41e46 · outbound

This paper cites Software testing with large language models: Survey, landscape, and vision,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Software testing with large language models: Survey, landscape, and vision,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.028343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.028343Z digest=sha256:27a3195ce1c88bedfb39e0ba972da1bad361eb462547aea86cea3dcfc497d815

Observation 604f551a-1ac6-4e77-876c-0e1f253f2d38 · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical evaluation of using large language models for automated unit test generation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.034055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.034055Z digest=sha256:c7158dfa6f26ecf92c560a14631ab34c2f61a91ef0e11832e4f5fd1e618a238a

Observation 16a0c54f-72b7-452f-ad91-f24e0a7b3b19 · outbound

This paper cites Automated test case repair using language models,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Automated test case repair using language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.592740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.038870Z digest=sha256:42f3b827db1c952d1c54f1b7d81a1d401fa80532571b91dbdb5c3ec7cb7ee75c

Observation a607be27-a5ec-42e0-9f50-e90ee04bff8a · outbound

This paper cites Chatunitest: a chatgpt- based automated unit test generation tool,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Chatunitest: a chatgpt- based automated unit test generation tool,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.577039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.044486Z digest=sha256:4813db3b472db088a4485302aaa093c5f2ce9d56b673734306ab40e98cee41e8

Observation 7501baa6-8233-41ba-922e-bc3d59d8bb78 · outbound

This paper cites An empirical study of using large language models for unit test generation,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical study of using large language models for unit test generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.561428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.049310Z digest=sha256:6d4ca2a73a10d8d32eaf68a97308a975de60f40ab039a4515635cd3a18e1534f

Observation d4444700-dfc5-48e5-8936-21cab8d709c6 · outbound

This paper cites Towards an understanding of large language models in software engineering tasks,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Towards an understanding of large language models in software engineering tasks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.545326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.055046Z digest=sha256:21e61482523b2b5fa69de5ac0ed96c90360b9b60a9c98f7f25442c46c22caa7f

Observation 01643684-ce9e-417e-b4ce-ee1b3ecd2056 · outbound

This paper cites Pynose: a test smell detector for python,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Pynose: a test smell detector for python,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.060347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.060347Z digest=sha256:3a47808c2ec7e55e20eea252d4de84bd6e9cb40a5bfd7dcbce0eba4a3f0659bf

Observation 37e1047a-9964-4933-ad72-31fc5fd62b2c · outbound

This paper cites Tempy: Test smell detector for python,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Tempy: Test smell detector for python,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.065639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.065639Z digest=sha256:1c356406d361f91bab24c8cf73219cca73e83f62634901cc9f7c5bfadc759a2f

Observation 1808a497-898a-4853-89d7-46f6bb0b93c4 · outbound

This paper cites Handling test smells in python: Results from a mixed-method study,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Handling test smells in python: Results from a mixed-method study,

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:35:41.834355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.070373Z digest=sha256:79840c77f8417ca0a86cba9c11e7a2f81dcb2aabf6833dcbd443ae8c8fd14d5c

Observation 0b467529-eab7-4e27-910f-cd8f9cd3ffd2 · outbound

This paper cites A trend analysis of test smells in python test code over commit history,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study A trend analysis of test smells in python test code over commit history,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.525234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.075592Z digest=sha256:61106a6fb6f18afed94b2c14004177ca7da4e1bc7dd0635ec39f26ad0e8f7d5b

Observation a7830eb5-2748-452a-9657-0adedbb77461 · outbound

This paper cites Pytest-smell: A smell detection tool for python unit tests,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Pytest-smell: A smell detection tool for python unit tests,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.080279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.080279Z digest=sha256:30e9c5744a4eba74a162d6f49d361396b903792c3018c0d351d7497fe26ab060

Observation 15ad84c6-b314-495e-a34f-c35c84b732d8 · outbound

This paper cites A trend analysis of test smells in python test code over commit history,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study A trend analysis of test smells in python test code over commit history,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.085461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.085461Z digest=sha256:365c27816eae0c9e3f0484a25e9257b9eacceeb9cdf59518ffe55cd9c52dc6fc

Observation 6e9b218f-0a68-451f-b6e1-53b2cb1d204a · outbound

This paper cites Detecting test smells in python test code generated by LLM: an empirical study with github copilot,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Detecting test smells in python test code generated by LLM: an empirical study with github copilot,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.090496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.090496Z digest=sha256:68aaebdd93f4ca8c11351840b41a8db89165d16b3705b957a81f36c7f1f28fc7

Observation 27036ac3-7f44-4dc0-90e4-a0c565ff228e · outbound

This paper cites Tsdetect: An open source test smells detection tool,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Tsdetect: An open source test smells detection tool,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.095249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.095249Z digest=sha256:8e62f23e6cf4f27df8ea928b49eb7c8988a78ca7917a641c4fc2dcd84ccd4636

Observation f37741b2-b5b4-4428-a81e-8f656088678b · outbound

This paper cites The secret life of test smells - an empirical study on test smell evolution and maintenance,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study The secret life of test smells - an empirical study on test smell evolution and maintenance,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.507134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.100530Z digest=sha256:9ceab6e42d112f77ac1978c7861a9a75fa752de5966f7677cf4cab992b8530ea

Observation c5a472aa-b24b-487a-893f-dded75dc24ff · outbound

This paper cites An empirical investigation into the nature of test smells,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical investigation into the nature of test smells,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.490301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.105536Z digest=sha256:1ac670da47231aadef08597caa752ef33f7612495595d180adf9956b3ae80f62

Observation 55a94a4e-6461-4fba-9507-0c1234c75a4d · outbound

This paper cites An empirical evaluation of raide: A semi-automated approach for test smells detection and refactoring,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical evaluation of raide: A semi-automated approach for test smells detection and refactoring,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.473239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.110958Z digest=sha256:47457e8442e8794207bf07d984556404b72ce2ae6d9ac4f4bcdc68f61c489ae3

Observation 4c6232e1-13ef-464f-9272-7c5832301035 · outbound

This paper cites Machine learning-based test smell detection,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Machine learning-based test smell detection,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.457599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.116048Z digest=sha256:91fbc01d55a6ee076bc1cce77cb6fc9e0456cf65ea5bfebea72366178663bcb3

Observation 17333b22-3370-4470-ada3-de15e04eaf6e · outbound

This paper cites Ml test smell detection - online appendix,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Ml test smell detection - online appendix,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.441334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.121328Z digest=sha256:aa8a810ad98c2c00f6b1544e05b8d0c1d68452eb106c5f8cdfc2a4a7c4775833

Observation 979a075e-875f-4c4e-817a-651e9631ee66 · outbound

This paper cites The Prompt Report: A Systematic Survey of Prompt Engineering Techniques.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study The Prompt Report: A Systematic Survey of Prompt Engineering Techniques

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.127695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.127695Z digest=sha256:8f9fa85851463253104efaf83281a38bf382d5a2e0607a7712af549e04789c0c

Observation 9e1f8995-92c8-409a-9f5e-50e8760143c9 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.133174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.133174Z digest=sha256:42c13aee707964d6f06bd782e60341962cac322559089729ed7bfed71ddd8293

Observation e92c0d03-5ce9-4438-8ffd-bad42eb089bf · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Finetuned Language Models Are Zero-Shot Learners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.138681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.138681Z digest=sha256:96cf4ead4ed0a31cd0aef1e5ba558f5322356414491e2cf5fc941e79ce7c706d

Observation bcf2c254-a14a-4662-9088-e3d66ed00209 · outbound

This paper cites Language models are few-shot learners,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Language models are few-shot learners,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.143500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.143500Z digest=sha256:68f1455c54c38e72a1cc0f126bbd8a0a5af2c84565a7e88a75cbfabbc65d146d

Observation e8bcc4d9-42d0-42c7-bd7c-674bfb4baf9a · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.154116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.154116Z digest=sha256:ab190d5d57fdcef9e769102da99080098675bfdc1c4d18360fb2e6949a850137

Observation f6690eac-0686-4d46-8b51-481bc0203440 · outbound

This paper cites Enhancing zero-shot chain-of-thought reasoning in large language models through logic,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Enhancing zero-shot chain-of-thought reasoning in large language models through logic,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.414785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.159476Z digest=sha256:3a70595f3570a29fb19f5dce2d4bc452ed08e3e22af8ad4ef95aa7fadcc14867

Observation c01ba966-08cf-47cb-af39-6179098f4152 · outbound

This paper cites Wilcoxon, Individual Comparisons by Ranking Methods.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Wilcoxon, Individual Comparisons by Ranking Methods

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.164376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.164376Z digest=sha256:1f4ac7857ea22b6e035e99055ff3aa6aee2297cdba93136e4e24f0c3807add3c

Observation 4b784f3d-ec10-4cef-ab7d-6c9848bebb76 · outbound

This paper cites Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.169640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.169640Z digest=sha256:50eab0187efe50669c1f14ab49ae4926f4b11183e415563684e49e43e5983972

Observation 203e52a2-3eb4-4dd5-93e2-9ec74aabd054 · outbound

This paper cites Towards effective validation and integration of llm-generated code,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Towards effective validation and integration of llm-generated code,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.399120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.174703Z digest=sha256:a1c858f36587c5b6a41c476f66d364da39b754559e27318347d352141ad07a57

Observation 4d3c0078-843f-4abe-a90e-af4cd3796809 · outbound

This paper cites Challenges and opportunities in integrating llms into con- tinuous integration/continuous deployment (ci/cd) pipelines,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Challenges and opportunities in integrating llms into con- tinuous integration/continuous deployment (ci/cd) pipelines,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.383731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.179595Z digest=sha256:0366f40b9c4ffb1bbae61d9fcafe16b8338ca686d1835a4db15b80dc71e28f2d

Observation 3866c2cb-de0c-4740-9325-25b67ecdb778 · outbound

This paper cites Next-generation refactoring: Combining llm insights and ide capabilities for extract method,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Next-generation refactoring: Combining llm insights and ide capabilities for extract method,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.367755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.184895Z digest=sha256:e15fdbae07c790bcc2b2c2147f35bb228bc94b3a808459c95b132cfcf88ebabf

Observation bcf1d462-afeb-49f6-8b97-81e8497ba694 · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.350942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.189548Z digest=sha256:4ce5039deb0cfa53b3fae40dd9632a4db4ef8246b5730d5589e485f7fa44502a

Observation ba98147c-6e57-4c7a-86bd-39625fad115e · outbound

This paper cites Autorefactoring: A platform to build refactoring agents,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Autorefactoring: A platform to build refactoring agents,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.335039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.194572Z digest=sha256:722180eeb0d5c907a8e6e51d36128b73abcbc2b9dc79a731940e11f7bcdcc970

Observation de96ca94-7c8f-43c9-8533-e95f24ffaad4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study DeepSeek-V3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.200114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.200114Z digest=sha256:c46ec9de51924e38262072f484c3c27e75f736f9f707c72484b86b40bcaa9947

Observation d33ee6e0-9751-4ba5-9b5b-b0d791ab5ddc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.204978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.204978Z digest=sha256:4af734ad0d41668232aaaa7e03e0bafa7f4607eb9879de19bacb772590f51483

Observation d9736758-fc93-45ee-819c-908bea04db40 · outbound

This paper cites Runeson, M.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Runeson, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.318400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.210159Z digest=sha256:91e8fcc375160117978cbc6cfab259fc4dd48c20edbbfb34e791ee7a19e617b6

Observation 9a477ee6-450c-4e12-8886-ead2dd72cacc · outbound

This paper cites Qualitative methods in empirical studies of software engineering,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Qualitative methods in empirical studies of software engineering,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.302511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:35:41.214905Z digest=sha256:3b7be6247685a270d3841bc9ce026c4198a4aa89db7e27193513e3a5398cfce3

Observation a9133213-2008-441f-ab73-53fa2c8f35fb · outbound

This paper cites Language Models are Few-Shot Learners.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.149159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.149159Z digest=sha256:572cf340e83d27a56a52b0ec27135d5feb0859ea44aa0ceff7ae450608334c01

Pith citing papers

Observation a8bbc461-08d7-43e7-a15d-eac1263a4d91 · inbound

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code cites this paper.

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:14.700137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T07:59:25.198736Z digest=sha256:d65e43e01b43ac4d9a49f9da806af61ac170677c8ade114b3f43719aacfc3ab0

Observation 8902ec67-7c8b-4674-894b-3fbf0d010e3c · inbound

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing cites this paper.

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.517733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:19:18.667967Z digest=sha256:da74ac7ac5eebbf42822a358355b2c6cdd6905620b30747cff1ba81ea549efcf

Observation 1fdc2dcf-d38d-430e-9dd7-f7c350880795 · inbound

Qiskit Code Migration with LLMs cites this paper.

Qiskit Code Migration with LLMs Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-06-26T16:29:35.581303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T16:24:25.357338Z digest=sha256:bb00de6f5d8c404143cf42292f4d291fabd3b2975d8a8a2304e43c981232b11f