Pith. sign in

Paper Citation Record · LEDGER

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness

As of 5 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2604.09564.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.09564 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T23:01:55.096929Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5661c9f6-5b41-4bc8-a50d-faa7efb1c8a6 · outbound

This paper cites an unresolved cited work.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-15T23:06:51.304111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:6f19cc20fce2f1f53ffcd6d1fcf3d3e98da7188eb235b638f9fee3534e0118da

Observation 5c2ff6e6-2358-44e5-9ca9-355a42a156af · outbound

This paper cites A Survey on Code Generation with LLM-based Agents.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness A Survey on Code Generation with LLM-based Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:02:21.366092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:cb355a26f48bd2e660b2e9120c3fbc1309c8df99648ef0b65ce985bd0e304f33

Observation 774770c3-1269-4ff3-8412-b8c11c9ae061 · outbound

This paper cites an unresolved cited work.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-15T23:06:51.291526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:e1130aa49dc7e95122150f022df24e759915ecb530895b6a36c6cfa27f0d5fa9

Observation 6b8049a5-3c1d-4a66-8960-41fb4b839f1b · outbound

This paper cites A Survey on LLM-as-a-Judge.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness A Survey on LLM-as-a-Judge

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T23:06:50.930309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:43e107a2ca70d15499b55acf9d365fabf13274a91d7fb58c50d26eba3374b82c

Observation 6677676a-13d8-4c2e-bde9-38e2c787193d · outbound

This paper cites SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:06:50.939021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:549ac76cc6f2e9c11183bf204a58069fad0783aeb445fd8592d29e2c0e16531c

Observation 9ef12b3d-cf9c-4e36-988f-93e8aa83d798 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness A Survey on Large Language Models for Code Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:06:51.179233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:c0dbdf358c02b8bfdb2456d154cd33ed8a982c72fa33cc1df6c929b4c37684d7

Observation 6718c6bf-869c-4b07-a343-76d0d4db2d49 · outbound

This paper cites Large language models effectively leverage document-level context for literary translation, but critical errors persist.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Large language models effectively leverage document-level context for literary translation, but critical errors persist

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:06:50.977008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:21b9d5103c02d1f5467cb972143e3a48b5beaea9f64da358af3f27f72e1b3207

Observation 459af4fe-e283-4b02-91e8-eca64ff582bb · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 10

Resolution
malformed identifier
local_arxiv, observed 2026-05-15T23:06:51.192504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:226726eaf5c2bebc8f3d3af22ffcdc86a3737e1c70450aeacdb25676c482a614

Observation a3e03cd9-7796-48c6-8607-3e5bd9aa29e4 · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando De Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando De Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 11

Resolution
verified exact
doi, observed 2026-05-15T23:06:50.969415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:2d0ab3d702a5b13cac74d54f29bb18e32129001e6f70ab552b67ac6af81d2da5

Observation a70708e0-41b8-4b7c-8ade-7638642c3cbb · outbound

This paper cites Symeonidis.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Symeonidis

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:06:50.993796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:b01e365015a0940d3a750a941695247011964861810991e22efc2cce0108c0fb

Observation ccbf649f-fca7-4cd9-84dd-478f70a8b900 · outbound

This paper cites Proceedings of the 2023.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Proceedings of the 2023

Reference 14

Resolution
verified exact
doi, observed 2026-05-15T23:06:50.963107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:9382293f63e6afdebfb4d590fd48849745211909722f48429f44e47432a23b4f

Observation 7f4a62c9-0fea-46ca-b75f-e87a190f902c · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 16

Resolution
malformed identifier
doi_truncated, observed 2026-05-15T23:06:51.000957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:445e38261a212614a58b803fb3579dfc34a0ba7acb4e6ec9abc9d78c9b8b025d

Observation e8caa12c-d629-416f-914f-e8d7668338a3 · outbound

This paper cites OpinSummEval: Revisiting Automated Evaluation for Opinion Summarization.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness OpinSummEval: Revisiting Automated Evaluation for Opinion Summarization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:06:51.185522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:e4f3944dcb87ccb9a4e1feba965fbbbf8170d813e132785590cd3b855e576bba

Observation 98206fa7-5edc-4d7b-98ab-e93110cd61e7 · outbound

This paper cites Human-In-the-Loop Software Development Agents.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Human-In-the-Loop Software Development Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:06:50.955770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:9d85641490de834a1ec8615e231784294846cfa019f8ced1f860289ac11035c1

Observation 3c54921f-bb19-49f1-bfd8-3cfed707a095 · outbound

This paper cites an unresolved cited work.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-15T23:06:51.299797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:f7582e922b08b30ba91422899bc34aa964115b0aad581604f3f7d7dd0e983faf

Observation 376fdf86-bd84-44fe-b48f-fab0bd439017 · outbound

This paper cites doi: 10.18653/v1/2021.emnlp-main.685.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness doi: 10.18653/v1/2021.emnlp-main.685

Reference 20

Resolution
verified exact
doi, observed 2026-05-15T23:06:50.898160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:be224f6860596ca262bc3669ad1ff61d4bf835cacba6cfdf0387e6f730ecb57c

Observation 28087e12-89ab-4fef-b6ba-21889fa226ea · outbound

This paper cites an unresolved cited work.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T23:06:51.295681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:d02727647182384341b5f7860057eb2bb94358a65d8a320b4ec0807d0a3d7eb3

Observation 61061a34-b92a-49cb-85ae-b64b7e677833 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T23:06:50.947153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:c5f31320d9da965505f9556ef83b553e5365cd94658fade59a1e413e44eae013

Observation 7a607383-f8ca-40f9-9b7f-ea7399da330f · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:06:50.914179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:6f26b1a8e1d67b1ac2779e8a75a9d362423810d45760ca170dde4bfa4465c01e

Observation 69abd6e7-b9ac-4746-9fe1-aac1e2116b8f · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:06:51.010834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:01:55.096929Z digest=sha256:2b176e6545927a8973c084eec39df5d70ce1f6469bb52caf1e7a8cb81421ac92

Pith citing papers

No inbound Pith citation observations are available.