Pith. sign in

Paper Citation Record · LEDGER

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?

As of 23 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2411.17927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17927 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:46:20.720359Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 416a0a09-8e82-45ef-8f66-1d9a829c47e1 · outbound

This paper cites Emergent abilities of large language models,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Emergent abilities of large language models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.235060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.629809Z digest=sha256:eb5c2e123ed16c1030a30b432e64539d80b2f241eae10d5198fa4bce5529f1dd

Observation 5bc978b2-c081-4760-a448-9185af02ba53 · outbound

This paper cites Measuring Data.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Measuring Data

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:46:20.635199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:46:20.635199Z digest=sha256:1798aab2bc3adc4cf6c8f30888fb419a86d867b647e116519ba1ca732f2daddb

Observation 61a3bf9a-5e2c-4cbc-a950-ebcc70b0facc · outbound

This paper cites Codegen: An open large language model for code with multi-turn program synthesis,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Codegen: An open large language model for code with multi-turn program synthesis,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.219840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.640309Z digest=sha256:39f52ef535b277acd640349728cc84cb18c777ecf0745bd575a55b2c99890dfe

Observation 711e452f-e0a4-4044-a102-69c20d0a9b10 · outbound

This paper cites Codegen2: Lessons for training llms on programming and natural languages,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Codegen2: Lessons for training llms on programming and natural languages,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.204675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.644399Z digest=sha256:ffbc51b56f4efe4844d9a69576300710b3f3ec62c888f447d4a6837037d5c072

Observation 668ce3b1-fe11-4e53-8a0a-79854a11e0d2 · outbound

This paper cites Are emergent abilities of large language models a mirage?.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Are emergent abilities of large language models a mirage?

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.188892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.648696Z digest=sha256:4ba66043b2d9178d85331f2e49631322ba8333ba888868f55ace5c62087432fc

Observation 0c6cc93c-b5de-42ef-96f5-49d776dd93db · outbound

This paper cites Are emergent abilities in large language models just in-context learning?.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Are emergent abilities in large language models just in-context learning?

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.173138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.652847Z digest=sha256:ae26d0c38c546637d6670387f55272cd7f331a3695e54161058c3458946c3f14

Observation 60f56d2f-ec20-49d0-9401-6fff90858473 · outbound

This paper cites Beyond the imitation game: Quantify- ing and extrapolating the capabilities of language models,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Beyond the imitation game: Quantify- ing and extrapolating the capabilities of language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.157271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.657440Z digest=sha256:de3f68191126a444cc1ea1242202301095118221e7ddab218038268fd53f2603

Observation b83ed9ba-fb9c-4d6d-ba85-8e463f8b74a7 · outbound

This paper cites Beyond accuracy: Behavioral testing of nlp models with checklist,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Beyond accuracy: Behavioral testing of nlp models with checklist,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.142687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.662185Z digest=sha256:0c3e6bc9cd054525d2e9817c335b79884909fb4c0c5f569bf0e8c0da12f02367

Observation f5abe65a-cb60-4d74-a6e0-bb1042d0500e · outbound

This paper cites A Structured Review of the Validity of BLEU,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? A Structured Review of the Validity of BLEU,

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-12T11:46:20.666642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:46:20.666642Z digest=sha256:d8fbfd90d717d8b70131ab181e0ee89c0829da0e5a7261a0113ddd7a87fdb25d

Observation 59fd066c-88d0-4a62-adbb-23a430fbfd9c · outbound

This paper cites ORANGE: a method for evaluating automatic evaluation metrics for machine translation,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? ORANGE: a method for evaluating automatic evaluation metrics for machine translation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.127455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.671240Z digest=sha256:9366ca7f7940c0a47f0681db6aac2a1ebf4a5f3606e842b5a8fc9e020ddbadc9

Observation e9d362b6-e235-47ec-9ee4-4f28c95c61ec · outbound

This paper cites Codebleu: a method for automatic evaluation of code synthesis,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Codebleu: a method for automatic evaluation of code synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.111813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.675622Z digest=sha256:19e6dd924a5ed94394cda15c2eb28985624e895d8075a898a5c3a4f2a2eb30b7

Observation b9992ad7-473d-4397-92b5-725acb4cd55b · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Bleu: a method for automatic evaluation of machine translation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.094623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.679900Z digest=sha256:dea2fad5df993258c11dbccad4d9bb6fc3fb6a6d0a4ba71ad518e19fb6f6c298

Observation 5b8c493f-07b1-42b4-83ba-283f5fffc07c · outbound

This paper cites CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:46:20.684520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:46:20.684520Z digest=sha256:e5f46c15a5cd5a5ee6e36e653b3884021fa60f39c2636f19a76fb25cddd09fd3

Observation 4674ae1c-6d29-474e-84dc-8053821ecbf2 · outbound

This paper cites Pypi/k4black/codebleu,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Pypi/k4black/codebleu,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.078945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.689247Z digest=sha256:a25ff23cf2d64d4203ce8b6cb762bee0a7189c8a9ae29fd8b30947eeac4161e8

Observation b18a490e-abf3-40d8-95a8-6e2dbbca4656 · outbound

This paper cites Commit message generation for source code changes,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Commit message generation for source code changes,

Reference 15

Resolution
verified exact
doi, observed 2026-08-12T11:46:20.766944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.693722Z digest=sha256:8c7eeb3af991bf978704f38c1f9eaf5889174dbafb3b64d22c16036ed36dad68

Observation 5577396c-4775-498a-b4ba-2dc8edd8f87b · outbound

This paper cites Using large language models for commit message generation: A preliminary study,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Using large language models for commit message generation: A preliminary study,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.062819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.698369Z digest=sha256:7c10a4b21a90a06900d7d04a8b0b76a065443702800edad5230668fa701b8b3c

Observation 8529a88b-818d-46be-844e-261a1e4fd7fb · outbound

This paper cites Emergent capabilities of LLMs for software engineering,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Emergent capabilities of LLMs for software engineering,

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-12T11:46:20.982681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.702787Z digest=sha256:6f3a063442f38f603760ed7f9b0ff981d007f41f781ff739f7e0558f11e28020

Observation aa2e705e-ee08-4f5c-b3d3-af45a455abbc · outbound

This paper cites On the evaluation of commit message generation models: An experimental study,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? On the evaluation of commit message generation models: An experimental study,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.046873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.707149Z digest=sha256:f7ff966bf25e1e16eca4a25ff3f5b8ed8f441f2c80cfaad3d1045c91294b749e

Observation c5ed9eb3-ab0f-4eeb-b045-b066387e313b · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big?.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? On the dangers of stochastic parrots: Can language models be too big?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:46:20.711526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:46:20.711526Z digest=sha256:0e4feabb06c85bb71562e4c24f45ed9a6d41dd2b0168ab09ee34f374673e4864

Observation 5a61fe14-d7c4-4141-bbcf-41c14aea55bc · outbound

This paper cites Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:46:20.715940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:46:20.715940Z digest=sha256:d1f89ac6fc5a28bfb2d6ec30ac93e3ae0ce2ec379b0761e8b9e794eec7196f19

Observation a9fe2d3b-8c2d-4803-b330-362a36cf7ccc · outbound

This paper cites Social biases in NLP models as barriers for persons with disabilities,.

Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We? Social biases in NLP models as barriers for persons with disabilities,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:46:21.031613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T11:46:20.720359Z digest=sha256:b6402dfafe21c2376df2cf8fcb1c4fee72a40b2ec52f3a1b5fa85f8a2b87bb48

Pith citing papers

No inbound Pith citation observations are available.