Pith. sign in

Paper Citation Record · LEDGER

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2505.19502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19502 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:49.895278Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:33.789468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:32:34.591005Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8c3daba-9935-4ca6-9220-153383d74bfc · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.793868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.793868Z digest=sha256:a5c8530504469c9139c64db4a1f64afe6478e7d52424b3ec48669f55de449b13

Observation 37ac8ab1-e125-492e-a656-74e769683d7f · outbound

This paper cites Efficient and green large language models for software engineering: Vision and the road ahead,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient and green large language models for software engineering: Vision and the road ahead,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.852067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.852067Z digest=sha256:22e4801d72554bb73a68c5fa9e46f472d5c8ec18075ca0507622942bb4d6c114

Observation 7059d352-e049-4c97-b5ff-c4b3d8c384a3 · outbound

This paper cites Large language models for software engi- neering: A systematic literature review,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Large language models for software engi- neering: A systematic literature review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.977776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.977776Z digest=sha256:fa6f7f567fd4b3999ec841b4f9260fa21c60873b694b2eb1520d6eb218198b12

Observation d5db457a-4915-4066-a1b2-91a769ee34f4 · outbound

This paper cites Exploring the capabilities of llms for code change related tasks,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploring the capabilities of llms for code change related tasks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:56.142825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:44.075982Z digest=sha256:8405a3650f7ca7ae6e10ed0d61e69be67e6e31e3b77a793e56ca798366cb8f8e

Observation 436ddac7-8fc6-453a-bd73-f5cb6ff0165b · outbound

This paper cites Chain- of-thought in neural code generation: From and for lightweight language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain- of-thought in neural code generation: From and for lightweight language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.190067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.190067Z digest=sha256:e5a950b113d231daba24dcc7f425cda4ece510e66ab212d4c05d54e208484283

Observation edc0ac39-d979-410b-899e-af50a5ee8016 · outbound

This paper cites An empirical study of retrieval-augmented code generation: Challenges and opportunities,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation An empirical study of retrieval-augmented code generation: Challenges and opportunities,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.323662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.323662Z digest=sha256:e2d13077eb8804580488e903621740a7113c7527e8ea98e7a5de631e180bb3bb

Observation b8e62aa2-67fc-4b4c-8ff7-ee8e4befcdde · outbound

This paper cites A review on code generation with llms: Application and evaluation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A review on code generation with llms: Application and evaluation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.419952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.419952Z digest=sha256:0b05879fd3afc2d578f1670a501628f9829bb5a244e129b648eaf80281b0fabc

Observation a016940c-4a33-4e26-af3d-37407ca86c7d · outbound

This paper cites HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.535532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.535532Z digest=sha256:a72c5cd31f7d4b7f16f9c5ec642c4b2894251f59b23636a8509afdc073c5289d

Observation da753d1c-7523-4255-8b2f-db694ff4245c · outbound

This paper cites Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.722335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:44.695755Z digest=sha256:c21807bb4fa118f9120e902f29b51c070aa4b1568aac1f922b41bb753efad40c

Observation 6f840e5e-74ea-4099-b2ff-15a0f23c8e5e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Bleu: a method for automatic evaluation of machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.800778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.800778Z digest=sha256:1f5458accd31cb8950af8be67fabf0d0cf7ee9beb44fa3d411470e1b99d349da

Observation be427ef5-801d-43d2-b29f-3dd2615fcca8 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Rouge: A package for automatic evaluation of summaries,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.882525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.882525Z digest=sha256:3e7bfd55b4776a017185bd9dae389a363a4f434bc545e6eaad40e4953a101483

Observation 61f6370f-1789-4085-b2ff-43d2915ce993 · outbound

This paper cites chrf: character n-gram f-score for automatic mt evalu- ation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation chrf: character n-gram f-score for automatic mt evalu- ation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.399227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:45.020391Z digest=sha256:8ececcca08242707c5ffb90c8ed0c335f9cf1609c5dce0578f1e3ed8f4119ff0

Observation 788e8390-42a0-43fb-bbc9-d3f73c7c27f7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.176339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.176339Z digest=sha256:228519ffd3ef66f8dcab9419dbb291cf3ec5ee960b6cc72b04a3c511266261bb

Observation 7bd5d4a7-0764-482d-93dc-49dec8202507 · outbound

This paper cites Exploit- gen: Template-augmented exploit code generation based on codebert,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploit- gen: Template-augmented exploit code generation based on codebert,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.097618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:45.322765Z digest=sha256:31b44d7ae4ce329454279367a2456c3ce772ef51dac9322533f175fde266a928

Observation 328742c4-9d6d-4754-9765-2a5d317b0ba1 · outbound

This paper cites Are nlp metrics suitable for evaluating gen- erated code?.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Are nlp metrics suitable for evaluating gen- erated code?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.811838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:45.485559Z digest=sha256:8816fde21c35de3c051e4f3ecddd18a58976b9cdd270d5a9723f653e2a76af13

Observation af2248f9-f5c9-4c45-84c7-7ea2822b126d · outbound

This paper cites On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:51.337928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:45.617534Z digest=sha256:28771a5209a7919140b3811e1191d4aea2ed765e5fdcfcb2b75ac78847ef3a57

Observation 7f780868-1084-45f3-a25b-b9518ed9b730 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.728953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.728953Z digest=sha256:a93185f87bfda403988daa3ae4c4116afaff2db31df1c26871ce4a52bdaa3ae0

Observation 606da85a-ab5e-4433-87a0-1a1720486b5f · outbound

This paper cites A Survey on LLM-as-a-Judge.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A Survey on LLM-as-a-Judge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.866412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.866412Z digest=sha256:dfab6467e014caf85cf854cb92cbd1ffb85d5f4689bf2a8b126a9443675e6054

Observation 6591cce8-2b4c-45b4-92b5-db04020bcf7d · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.986897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.986897Z digest=sha256:85302086ab2fe3bd1d61f434d739d59fb062d0b69af97502cb8e7373ed62ef5f

Observation 954b232b-51e2-48e3-9a83-80db1bc437ab · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.158301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.158301Z digest=sha256:33ba0e173e22157abb8d4696c93625bb9501d3fb335adbc84c26e3f8f3fc5ab9

Observation 2dc63c72-91a8-42d9-ac61-f8f830f53b5b · outbound

This paper cites Fight fire with fire: How much can we trust chatgpt on source code-related tasks?.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Fight fire with fire: How much can we trust chatgpt on source code-related tasks?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.479297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:46.267717Z digest=sha256:4869be49db8fdf7e455b8ef04e2c66325a52772c3b2be1b1f05cc4e16276f425

Observation df1fc234-6b03-4450-a8f2-74ceb248c922 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.412755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.412755Z digest=sha256:07c7e1eddec008b3eb6a3b68b4f1d49ff81eaebf57c4ccf447aa98f01c1be691

Observation 5eb22316-f96c-4534-9ccc-4ab52109222b · outbound

This paper cites Pissa: Principal singular values and singular vectors adaptation of large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Pissa: Principal singular values and singular vectors adaptation of large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.106941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:46.534245Z digest=sha256:81a11a18750afefb9e08f41c05c8c3d2f6abbfe1a5fd554c4f622cc4cbed98e5

Observation fb82432f-3dcd-4202-a6f1-cedd7489ad02 · outbound

This paper cites DeepSeek-V3 Technical Report.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.672703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.672703Z digest=sha256:edf2e39c71200bee2901345bd0b14081bc709b38f3caf935f20f51c0efc6381c

Observation 19d5527f-75f7-4917-b9f3-278883ebead8 · outbound

This paper cites Preference leakage: A contamination problem in llm-as-a-judge,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Preference leakage: A contamination problem in llm-as-a-judge,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.831139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.831139Z digest=sha256:4529e15bfe1bf6af3958fe87de0baa7de3c17e9589c710d19e5d50dbac87e374

Observation 44ad428f-1c5c-4197-ba0d-3a659e98494a · outbound

This paper cites Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.843328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:46.961223Z digest=sha256:d985f830a5bd83693308f3b2eb590959603a1a1ea712b430fb64aa420378660c

Observation 97dc3270-991f-4aca-8444-5ffd1dab3267 · outbound

This paper cites Crystalbleu: precisely and efficiently mea- suring the similarity of code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Crystalbleu: precisely and efficiently mea- suring the similarity of code,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.498388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:47.156377Z digest=sha256:e023cf332060f59db034223f1980ee0e79409eae65ccdc1ee8d83edfe85e7aa8

Observation a48577a0-d159-428b-897e-195126bdf97b · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:47.284456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:47.284456Z digest=sha256:9a6bf3be1c9230e7e54699c1c92f60404149933357e6bb525f1db90114472dec

Observation a87f8fb3-0a76-475d-95e9-3ec920ebc06c · outbound

This paper cites Codebertscore: Eval- uating code generation with pretrained models of code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codebertscore: Eval- uating code generation with pretrained models of code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.273794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:47.431493Z digest=sha256:d04aeceb1f5e6c62b94fb03d6fc6930c2f8aa2aa1a26f928b2902183ba41d285

Observation a1f8e276-a258-4a87-8e3a-99cd99b7399d · outbound

This paper cites Codescore: Eval- uating code generation by learning code execution,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codescore: Eval- uating code generation by learning code execution,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.904879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:47.555464Z digest=sha256:912b7e31658952d2243ccdc67a27d24abc5ee256cd5d53ed929e44b6f04c0093

Observation b97d9c13-1d3a-4464-9b9d-2d0527eba567 · outbound

This paper cites CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:50.455945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:47.689046Z digest=sha256:0987799f994c0fb60caa1218022be85c97dd57e44ff3d097f7979509256703f3

Observation 2aaf3ad2-026c-4451-b6ed-2ba896733148 · outbound

This paper cites Ice-score: Instructing large language models to evaluate code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Ice-score: Instructing large language models to evaluate code,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.590439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:47.812514Z digest=sha256:924a3257a28f4f9b687a0491375090bb9da04f33627a5452357b1aaec0255e50

Observation 831fc490-587b-43e4-9ad2-14fa7a542494 · outbound

This paper cites Codejudge: Evaluating code generation with large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codejudge: Evaluating code generation with large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.336829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:47.980501Z digest=sha256:d528332fb92ce0c3b9294db32e8332b61012ddac07fd46fead9c18a9283f300b

Observation d1b05140-ba45-4990-8fa0-9b199a50692d · outbound

This paper cites Benchmarks and metrics for evalua- tions of code generation: A critical review,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Benchmarks and metrics for evalua- tions of code generation: A critical review,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.066356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:48.119267Z digest=sha256:deedfac479a8896b15d0dacb9b5f3ac34256bd952d0d66c2b5ce737098be9e4d

Observation 20add06a-1154-416b-8b05-f3a185bb4f3b · outbound

This paper cites From Code to Courtroom: LLMs as the New Software Judges.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From Code to Courtroom: LLMs as the New Software Judges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.242132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.242132Z digest=sha256:61c33c41198994c0b0649c265a46975c071b353ead76ccfba8ebccb315cbc41c

Observation 42377c96-b83f-44c4-8a1d-359778743253 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:51.761783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:18:48.442996Z digest=sha256:727a4afde018cb6838ae74e69e28763105429571aa6efe6862610124226329ef

Observation c35ee74f-45a5-416f-9a9e-a9027fa1014a · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.551215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.551215Z digest=sha256:f6304f16bac44a50523f0e2dde970c62bd5425c232003bbece4be6b229fcd27e

Observation d5e8fd5b-c57c-445c-88d3-ed7d2cb88b75 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Qwen2.5-Coder Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.666706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.666706Z digest=sha256:79fc9738efa780bbb11bea1c0abc0bea4a6a3a54b6793b4b4bb2237f83843718

Observation 5b65b982-8887-435a-ac28-b5f595e03134 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.828289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.828289Z digest=sha256:e1f84c7942060daeb9a7a2d1780529a7a56644fae036c2dbbb97596e5789cde1

Observation d316dd0b-02f7-4b68-99aa-12cdd7cda6d6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain-of-thought prompting elicits reasoning in large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.977432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.977432Z digest=sha256:3f7e655adc950d8af348b80ea2e4a793ed5ad551e746bc3db58164d5dd9cb970

Observation 8a620128-32a8-4efc-9f36-9fb7a9a7ba48 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient memory management for large language model serving with pagedattention,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.169400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.169400Z digest=sha256:4fd7285ca6cc164910535ba8b8153213bc7d5bb2505826e882c259273a33f099

Observation 3a2a8f9d-ef7c-44b5-8dbe-1e0898697da3 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.297876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.297876Z digest=sha256:3e2a08ac47d0ebe873f802bd48ef9f8b4e2b3286e51758bd2edb15ac5a538e52

Observation d4780c4b-4d30-4d97-8fbc-306e2706f1f9 · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.422123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.422123Z digest=sha256:b181830788bf33d1ca3e05987efc260907b0835ce1b05989135a473b7b83c303

Observation f38a23b2-d917-42f4-92ce-7d3fdfb8d31c · outbound

This paper cites Less is More: Towards Green Code Large Language Models via Unified Structural Pruning.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Less is More: Towards Green Code Large Language Models via Unified Structural Pruning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.570812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.570812Z digest=sha256:8fb37ea4ac716fb2050795e3f784e4ead6c9011532cffd537cc2f0d39e15b574

Observation 4877cb16-649c-4715-ab71-f88e482f3690 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.765210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.765210Z digest=sha256:96e58a87ade5a9d1b7142dbd1d15b75cf3203ed7df42d5c65e8b4972c7a1ad17

Observation 6d7b34fb-a8b3-4a53-b612-d2274275d5a2 · outbound

This paper cites A coefficient of agreement for nominal scales,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A coefficient of agreement for nominal scales,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.895278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.895278Z digest=sha256:2f9917e0f8211b7d7c26d8302336b0fb794de290580bd7b89436d7facb9ca256

Pith citing papers

Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:32:34.645100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.789468Z digest=sha256:f4e9582f1a15049b5ebcfde4a85624a2015e1d02a107e15127f092732c8acad0