Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:49.895278Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2505.19502.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:49.895278Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:33.789468Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T17:32:34.591005Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c8c3daba-9935-4ca6-9220-153383d74bfc · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ac8ab1-e125-492e-a656-74e769683d7f · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient and green large language models for software engineering: Vision and the road ahead,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7059d352-e049-4c97-b5ff-c4b3d8c384a3 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Large language models for software engi- neering: A systematic literature review,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5db457a-4915-4066-a1b2-91a769ee34f4 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploring the capabilities of llms for code change related tasks,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 436ddac7-8fc6-453a-bd73-f5cb6ff0165b · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain- of-thought in neural code generation: From and for lightweight language models,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc0ac39-d979-410b-899e-af50a5ee8016 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation An empirical study of retrieval-augmented code generation: Challenges and opportunities,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e62aa2-67fc-4b4c-8ff7-ee8e4befcdde · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A review on code generation with llms: Application and evaluation,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a016940c-4a33-4e26-af3d-37407ca86c7d · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da753d1c-7523-4255-8b2f-db694ff4245c · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f840e5e-74ea-4099-b2ff-15a0f23c8e5e · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Bleu: a method for automatic evaluation of machine translation,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be427ef5-801d-43d2-b29f-3dd2615fcca8 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Rouge: A package for automatic evaluation of summaries,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f6370f-1789-4085-b2ff-43d2915ce993 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation chrf: character n-gram f-score for automatic mt evalu- ation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 788e8390-42a0-43fb-bbc9-d3f73c7c27f7 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Evaluating Large Language Models Trained on Code
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd5d4a7-0764-482d-93dc-49dec8202507 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploit- gen: Template-augmented exploit code generation based on codebert,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 328742c4-9d6d-4754-9765-2a5d317b0ba1 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Are nlp metrics suitable for evaluating gen- erated code?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af2248f9-f5c9-4c45-84c7-7ea2822b126d · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f780868-1084-45f3-a25b-b9518ed9b730 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606da85a-ab5e-4433-87a0-1a1720486b5f · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A Survey on LLM-as-a-Judge
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6591cce8-2b4c-45b4-92b5-db04020bcf7d · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954b232b-51e2-48e3-9a83-80db1bc437ab · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc63c72-91a8-42d9-ac61-f8f830f53b5b · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Fight fire with fire: How much can we trust chatgpt on source code-related tasks?
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df1fc234-6b03-4450-a8f2-74ceb248c922 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb22316-f96c-4534-9ccc-4ab52109222b · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Pissa: Principal singular values and singular vectors adaptation of large language models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb82432f-3dcd-4202-a6f1-cedd7489ad02 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-V3 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d5527f-75f7-4917-b9f3-278883ebead8 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Preference leakage: A contamination problem in llm-as-a-judge,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ad428f-1c5c-4197-ba0d-3a659e98494a · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97dc3270-991f-4aca-8444-5ffd1dab3267 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Crystalbleu: precisely and efficiently mea- suring the similarity of code,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a48577a0-d159-428b-897e-195126bdf97b · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87f8fb3-0a76-475d-95e9-3ec920ebc06c · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codebertscore: Eval- uating code generation with pretrained models of code,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1f8e276-a258-4a87-8e3a-99cd99b7399d · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codescore: Eval- uating code generation by learning code execution,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b97d9c13-1d3a-4464-9b9d-2d0527eba567 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2aaf3ad2-026c-4451-b6ed-2ba896733148 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Ice-score: Instructing large language models to evaluate code,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 831fc490-587b-43e4-9ad2-14fa7a542494 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codejudge: Evaluating code generation with large language models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1b05140-ba45-4990-8fa0-9b199a50692d · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Benchmarks and metrics for evalua- tions of code generation: A critical review,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20add06a-1154-416b-8b05-f3a185bb4f3b · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From Code to Courtroom: LLMs as the New Software Judges
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42377c96-b83f-44c4-8a1d-359778743253 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c35ee74f-45a5-416f-9a9e-a9027fa1014a · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e8fd5b-c57c-445c-88d3-ed7d2cb88b75 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Qwen2.5-Coder Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b65b982-8887-435a-ac28-b5f595e03134 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d316dd0b-02f7-4b68-99aa-12cdd7cda6d6 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain-of-thought prompting elicits reasoning in large language models,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a620128-32a8-4efc-9f36-9fb7a9a7ba48 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient memory management for large language model serving with pagedattention,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2a8f9d-ef7c-44b5-8dbe-1e0898697da3 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4780c4b-4d30-4d97-8fbc-306e2706f1f9 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38a23b2-d917-42f4-92ce-7d3fdfb8d31c · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Less is More: Towards Green Code Large Language Models via Unified Structural Pruning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4877cb16-649c-4715-ab71-f88e482f3690 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d7b34fb-a8b3-4a53-b612-d2274275d5a2 · outbound
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A coefficient of agreement for nominal scales,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · inbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.