Pith. sign in

Paper Citation Record · LEDGER

On the Robustness of LLMs' Internal Representation of Code Correctness

As of 15 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08266 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:15:27.811027Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e93653f-e357-4836-9f0d-3e8cb519fa4c · outbound

This paper cites In- tellicode compose: code generation using transformer,.

On the Robustness of LLMs' Internal Representation of Code Correctness In- tellicode compose: code generation using transformer,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.889232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.477382Z digest=sha256:668a42c1553e77d20957970524e5da54fba17a916ea5bfa084dcff7aed87d571

Observation 877a3a7d-f6f4-4373-9075-e6d02936af85 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On the Robustness of LLMs' Internal Representation of Code Correctness Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.483965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.483965Z digest=sha256:367835af860e5b512cab80128fd52ad32ff79109feb61f3eae5b3bb0d42603dd

Observation 17d18d29-37c6-4de7-a433-60a30775feb4 · outbound

This paper cites SWE-bench: Can language models resolve real-world GitHub issues?.

On the Robustness of LLMs' Internal Representation of Code Correctness SWE-bench: Can language models resolve real-world GitHub issues?

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.872981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.490111Z digest=sha256:536cbdfc9201383e4ed16eee90f9898f8e8859b9e8af8e47c4dc9d14e985ab69

Observation 4a874b0f-f250-4d20-8977-143a61dee2ab · outbound

This paper cites Grounded copi- lot: How programmers interact with code-generating models,.

On the Robustness of LLMs' Internal Representation of Code Correctness Grounded copi- lot: How programmers interact with code-generating models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.851862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.496038Z digest=sha256:143c364aeb91eb463743788df99fc735806be88ff7e15729f884ed27cb0dd3f6

Observation 06bddc3f-c8b3-4d2f-a2bd-bc7376315fdf · outbound

This paper cites Security weaknesses of copilot-generated code in GitHub projects: An empirical study,.

On the Robustness of LLMs' Internal Representation of Code Correctness Security weaknesses of copilot-generated code in GitHub projects: An empirical study,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.829988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.501758Z digest=sha256:5a3c921a119e0b04ab3238b2abb99c08634ed395d7956840bc469fdfda80a693

Observation 951b6ab5-a4a3-43f2-aba8-f5472d1c9c20 · outbound

This paper cites Towards understanding the characteristics of code generation errors made by large language models,.

On the Robustness of LLMs' Internal Representation of Code Correctness Towards understanding the characteristics of code generation errors made by large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.812477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.512294Z digest=sha256:cce675fa376006a40c1e602160107c2d867bd3a8f8c5435e3f51941dcded1020

Observation 3577482a-3108-49f5-9e20-97a6e1c724f1 · outbound

This paper cites The counterfeit conundrum: Can code lan- guage models grasp the nuances of their incorrect generations?.

On the Robustness of LLMs' Internal Representation of Code Correctness The counterfeit conundrum: Can code lan- guage models grasp the nuances of their incorrect generations?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.794242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.519104Z digest=sha256:c49d34b69edc28836a3f925fee0de57799b8e680c3eb89da8bd6728acfca0018

Observation 04ac78fc-cb0a-42f1-bf13-fc8f9ecebb41 · outbound

This paper cites Theoracleprobleminsoftwaretesting:Asurvey,.

On the Robustness of LLMs' Internal Representation of Code Correctness Theoracleprobleminsoftwaretesting:Asurvey,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.773546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.524564Z digest=sha256:7b8a5b55a081ac547d10936a0e995ea0712e0e72fd5000ba914db4f732048dfa

Observation 2e8662ee-86b1-44b1-b7c6-078b1641ed17 · outbound

This paper cites Justaskforcalibration:Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback,.

On the Robustness of LLMs' Internal Representation of Code Correctness Justaskforcalibration:Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.757064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.533572Z digest=sha256:4ca2b96a4bf1df837b8c45b69818360d61975a4d1b178e654b458e9fbe4e45c6

Observation ec70991c-df18-41cc-86e9-ae935561f171 · outbound

This paper cites Calibration and correctness of language models for code,.

On the Robustness of LLMs' Internal Representation of Code Correctness Calibration and correctness of language models for code,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.734956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.539073Z digest=sha256:4443f521c0bff5aa932b0f1501ef78abd94d3ebe3aa3f4c38a0b6420e8202c18

Observation e2f7a0e6-0154-4df4-bacf-c6fb1540aa19 · outbound

This paper cites On LLMs’ internal representation of code correctness,.

On the Robustness of LLMs' Internal Representation of Code Correctness On LLMs’ internal representation of code correctness,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.714600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.544629Z digest=sha256:085ae4de0a9b6a07467f8bc7bd8713b3acd9b8fc9c010c3a028368dbda40aa9a

Observation 10eeb383-e3d8-46ac-bf3b-12988748ab97 · outbound

This paper cites The internal state of an LLM knows when it’s lying,.

On the Robustness of LLMs' Internal Representation of Code Correctness The internal state of an LLM knows when it’s lying,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.693672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.550016Z digest=sha256:9fba9f0a5e8860d8e6825c4d886b393eafce6fa1e76c6fee49cce214e62e1805

Observation fbbc40d2-f6c0-426e-8212-1fc1624ccc49 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

On the Robustness of LLMs' Internal Representation of Code Correctness The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.555203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.555203Z digest=sha256:4168076dbb35d89676b91a49eb5ee6f4bfb0415ac32a5eceb93ca474b897fd52

Observation 6bbba6f8-3c22-4e6c-8a1a-5eef5803e0fb · outbound

This paper cites Discovering latent knowledge in language models without supervision,.

On the Robustness of LLMs' Internal Representation of Code Correctness Discovering latent knowledge in language models without supervision,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.677517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.561731Z digest=sha256:ee5fee7b1481b654d20f03e203d778bcf98fb833dff6c8885992625bb6343a02

Observation cf3d792a-fb7c-4c8a-bd53-e2aae564c61e · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

On the Robustness of LLMs' Internal Representation of Code Correctness Representation Engineering: A Top-Down Approach to AI Transparency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.567144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.567144Z digest=sha256:4780f07b64f20e0fcdcaa363812686461b0bad4f831a3f6c48eaa857e43a796d

Observation 5a80088c-9d98-417a-8a8c-48449a052a0a · outbound

This paper cites A unified understanding and evaluation of steering methods,.

On the Robustness of LLMs' Internal Representation of Code Correctness A unified understanding and evaluation of steering methods,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.573341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.573341Z digest=sha256:50eee3de3f26bb7679fd2f07fd07a2f355b3cd093ce0e05620f031a7a4a35b41

Observation 105803a4-9f5d-43df-80e7-d6e281737d44 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

On the Robustness of LLMs' Internal Representation of Code Correctness BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.578999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.578999Z digest=sha256:d03f7c83015c990e8614d9ebad1da2cd747d6df6ac04eae2f0ad4164dcf0b033

Observation 9025859a-db23-4db2-b2f3-eafc183a55a7 · outbound

This paper cites Program Synthesis with Large Language Models.

On the Robustness of LLMs' Internal Representation of Code Correctness Program Synthesis with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.585505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.585505Z digest=sha256:3fcc1e22010c262c5d14878bd5cd2a208af625c758cfb116dfc928591e3c1fcf

Observation 1e907841-bae2-4bc3-a7e2-4ea34aacff98 · outbound

This paper cites Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation,.

On the Robustness of LLMs' Internal Representation of Code Correctness Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.658163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.592054Z digest=sha256:6139ff5414784e2b8c69b5c2c844719e13fa6d0b8da3759adbe69db9ad421b92

Observation 233147c0-30ef-416d-a4ed-8c18fcbb2844 · outbound

This paper cites Attention is all you need,.

On the Robustness of LLMs' Internal Representation of Code Correctness Attention is all you need,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.598961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.598961Z digest=sha256:8d8adbbdb359f09fe1bf7f71dd13f69ba7356cc89313629327c829f99d5d3ee8

Observation ca461b9f-42b8-47a6-9f2b-f94c33d1eeb1 · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model,.

On the Robustness of LLMs' Internal Representation of Code Correctness Inference-time intervention: Eliciting truthful answers from a language model,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.619242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.603645Z digest=sha256:141d205b4d18e77a1bc27bd438576796373ef949bc61bc5a77dc63f269ff5463

Observation 3f69ca45-5c69-40c2-9c7b-aaee36175acc · outbound

This paper cites Principal component analysis,.

On the Robustness of LLMs' Internal Representation of Code Correctness Principal component analysis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.599552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.609121Z digest=sha256:a1d78b371254d8865409947cbc45c1226dd86372165c07cf9959595f3dd0ea85

Observation 43971965-fdbb-4614-9411-a674dbb661da · outbound

This paper cites Language Models (Mostly) Know What They Know.

On the Robustness of LLMs' Internal Representation of Code Correctness Language Models (Mostly) Know What They Know

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.613962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.613962Z digest=sha256:8c99256c5ac428fbf26f4bca8b624539e5782adb9dba61c50d4debeae65ba5ae

Observation 55d5e476-13b2-402e-8cf5-c843dae8592e · outbound

This paper cites Sifting through the chaff: On utilizing execution feed- back for ranking the generated code candidates,.

On the Robustness of LLMs' Internal Representation of Code Correctness Sifting through the chaff: On utilizing execution feed- back for ranking the generated code candidates,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.580268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.619749Z digest=sha256:dffa85438816a7c5add33700aebf967b17eb0e311e2e9310a96ce67c5df4a6e0

Observation 20abe70c-221c-4dbd-9983-ea0e058b2921 · outbound

This paper cites astroid: A common base representation of python source code,.

On the Robustness of LLMs' Internal Representation of Code Correctness astroid: A common base representation of python source code,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.556225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.625004Z digest=sha256:6a62936a7f077e11e71b1b4b38020e65315b5c63b87c81d0a8a9dd4e134dcf0d

Observation 2f981508-16eb-4fd7-8584-a88651a0981b · outbound

This paper cites Are mutants a valid substitute for real faults in software testing?.

On the Robustness of LLMs' Internal Representation of Code Correctness Are mutants a valid substitute for real faults in software testing?

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.534463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.629429Z digest=sha256:7d046fd53c8242deacfc4bd902c7e498d54fea4a0bac5ffa3fff9c8102009231

Observation f528ce79-a430-4fb0-ac23-0ad4d0db8031 · outbound

This paper cites Qwen2.5-Coder Technical Report.

On the Robustness of LLMs' Internal Representation of Code Correctness Qwen2.5-Coder Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.634934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.634934Z digest=sha256:0e6d3a990d7f304509f1d2f4d5779cf2ed446cbb771208edeb7ad77b28257e4f

Observation 984e61b7-771f-46f0-bc32-0b30099aa377 · outbound

This paper cites Model-agnostic quality assessment for LLM-generated code via dynamic internal representation selection,.

On the Robustness of LLMs' Internal Representation of Code Correctness Model-agnostic quality assessment for LLM-generated code via dynamic internal representation selection,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.508624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.640409Z digest=sha256:40ab3ceddc047c5f46a578fe2a4ba1b9a5f514e7b70d05e09b4a379894a7c988

Observation 3a0aa7bc-35bf-49a6-9bdc-9a6ab97dd86a · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances,.

On the Robustness of LLMs' Internal Representation of Code Correctness Probing classifiers: Promises, shortcomings, and advances,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.483479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.646478Z digest=sha256:8b7b319294c00a5f6648e63b735c44138572c17387e8df5677fa58705839eda0

Observation 1b5ee281-1079-4a65-9025-3e0d51655520 · outbound

This paper cites Designing and interpreting probes with control tasks,.

On the Robustness of LLMs' Internal Representation of Code Correctness Designing and interpreting probes with control tasks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.457140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.651644Z digest=sha256:6220f98cd3104d2123d2eb0939e311657f2c8273157298c0e0535f135d107de8

Observation 866dd0fd-62fe-44eb-9d9d-72a05f3bdee6 · outbound

This paper cites Probingtheprobing paradigm: Does probing accuracy entail task relevance?.

On the Robustness of LLMs' Internal Representation of Code Correctness Probingtheprobing paradigm: Does probing accuracy entail task relevance?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.437575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.658166Z digest=sha256:064cc67bbbf2948c87a65b0db1dd8bca4ca8c5572572cea3af6d19e7d05c161c

Observation 956918f9-deee-45ce-9b85-6fce21efb410 · outbound

This paper cites Amnesic probing: Behavioral explanation with amnesic counterfactuals,.

On the Robustness of LLMs' Internal Representation of Code Correctness Amnesic probing: Behavioral explanation with amnesic counterfactuals,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.420608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.663983Z digest=sha256:f554706eb90579eda6b42a70bd4c16308b7b6203743deab695f5efadc84a7827

Observation a3fc7e49-cf0b-4259-9e35-fa3a61a3e823 · outbound

This paper cites Probing classifiers are unreliable for concept removal and detection,.

On the Robustness of LLMs' Internal Representation of Code Correctness Probing classifiers are unreliable for concept removal and detection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.393262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.671328Z digest=sha256:f0297ba3cb11d34f0db68ef6836554424971abce13300b2fd251fcae97600ad8

Observation 49a90a7e-46d7-432e-ae7a-57799b7e2ee0 · outbound

This paper cites Semantic bug seeding: A learning- based approach for creating realistic bugs,.

On the Robustness of LLMs' Internal Representation of Code Correctness Semantic bug seeding: A learning- based approach for creating realistic bugs,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.368419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.677111Z digest=sha256:a6facea23a138e146aaae264da91f55fbf1098aaf4ba50e9a4a23cdddc58211c

Observation c2a10e7b-199b-453c-844e-3db071438fb9 · outbound

This paper cites Learningrealisticmutations:Bug creation for neural bug detectors,.

On the Robustness of LLMs' Internal Representation of Code Correctness Learningrealisticmutations:Bug creation for neural bug detectors,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.346430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.682540Z digest=sha256:ebe22537a60f06da3b6a871605bf1d68d7dfe1f42d172f82863c49625e4a495c

Observation e1cf6848-5e56-42e2-8d63-b78a11d1df34 · outbound

This paper cites On distribution shift in learning-based bug detectors,.

On the Robustness of LLMs' Internal Representation of Code Correctness On distribution shift in learning-based bug detectors,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.309513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.688620Z digest=sha256:077bb96cf20587efc7a31aff783ef87e54426ff244f783e602f6cc8bbe1bd515

Observation f4d00e17-e670-4c54-b3c7-4069bbd6b41d · outbound

This paper cites Large language models of code fail at completing code with potential bugs,.

On the Robustness of LLMs' Internal Representation of Code Correctness Large language models of code fail at completing code with potential bugs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.267087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.694570Z digest=sha256:7dd508f783ad5b944cc5d19a205be4e555eeccb5857f2d5a7c9b07e0eb0afe95

Observation c170a50a-94a1-4359-8d7d-2097cd0429b2 · outbound

This paper cites On the universal truthfulness hyperplane inside LLMs,.

On the Robustness of LLMs' Internal Representation of Code Correctness On the universal truthfulness hyperplane inside LLMs,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.209872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.700031Z digest=sha256:a3e4c85c75d390f9a8e099e34291f3dd11504e629e85ca929ec5dec554a59f95

Observation 27d87201-f698-42ec-b0e1-b566be2bdec8 · outbound

This paper cites Probing the geometry of truth: Consistency and generalization of truth directions in LLMs across logical transformations and question answering tasks,.

On the Robustness of LLMs' Internal Representation of Code Correctness Probing the geometry of truth: Consistency and generalization of truth directions in LLMs across logical transformations and question answering tasks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.181399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.705099Z digest=sha256:efc145f0eea5f1181484bef1e1a7be967e03ad55f838f153ae0224ec10d16643

Observation 7e286f9b-bd6c-4a1e-b83c-affe119ff2b7 · outbound

This paper cites The truth- fulness spectrum hypothesis,.

On the Robustness of LLMs' Internal Representation of Code Correctness The truth- fulness spectrum hypothesis,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.710684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.710684Z digest=sha256:0e5f3b9c7cb5b7c6c7856312fe1edb3d3d7d91256d05900a8efaf371ce48cbaf

Observation 02e1ab20-b429-4726-9e1d-2b8936af9daa · outbound

This paper cites Correctness assessment of code generated by large language models using internal representations,.

On the Robustness of LLMs' Internal Representation of Code Correctness Correctness assessment of code generated by large language models using internal representations,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.158848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.715508Z digest=sha256:4271b80907a3c855617b3821f870e580aa1317c75f7a1434a278e6480966b174

Observation 04b114ee-6da1-4788-8b3f-b491e1c509b9 · outbound

This paper cites LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations.

On the Robustness of LLMs' Internal Representation of Code Correctness LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.720442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.720442Z digest=sha256:2b2f94a381708c71e544c5fef2910bb05d27dc8e63c8cf45ed0e24b280edf60b

Observation fa053f60-75e7-4c9b-bf44-6b01ead5317f · outbound

This paper cites Risk assess- ment framework for code LLMs via leveraging internal states,.

On the Robustness of LLMs' Internal Representation of Code Correctness Risk assess- ment framework for code LLMs via leveraging internal states,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.134448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.725558Z digest=sha256:a83a0ca083f0f0755482136ba8f476ba8de5a0711ee926ed902a681397c5ef56

Observation e87dedce-ebf2-449e-947e-1bf07837bfb8 · outbound

This paper cites Localized calibrated uncertainty in code language models,.

On the Robustness of LLMs' Internal Representation of Code Correctness Localized calibrated uncertainty in code language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.731169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.731169Z digest=sha256:9ba0c74788eaf0c268fbfacd2fa2c6186e3260495872f0bf6eb7fd06770c3eb0

Observation 20a283e3-e804-485b-b22f-0e0d1d3b3185 · outbound

This paper cites Mechanistic interpretability of code correctness in llms via sparse autoencoders,.

On the Robustness of LLMs' Internal Representation of Code Correctness Mechanistic interpretability of code correctness in llms via sparse autoencoders,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.736628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.736628Z digest=sha256:fe998ba654bf5ff5c601c98666f5f1ed273a5c60ad176f4790baca9adf7eb3df

Observation 78d22ae2-ec23-4f3c-b2d8-e9ad0c32b590 · outbound

This paper cites Code- Circuit: Toward inferring LLM-generated code correctness via attribution graphs,.

On the Robustness of LLMs' Internal Representation of Code Correctness Code- Circuit: Toward inferring LLM-generated code correctness via attribution graphs,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.741684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.741684Z digest=sha256:e53a91278b85c7b119f0c5bf0a5ec0e512f0ce4d073970494f051bcc7b4bbcc0

Observation 9ca37055-02f3-42b8-8c18-d758295ca45e · outbound

This paper cites Emergentrepresentationsofprogramse- manticsinlanguagemodelstrainedonprograms,.

On the Robustness of LLMs' Internal Representation of Code Correctness Emergentrepresentationsofprogramse- manticsinlanguagemodelstrainedonprograms,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.108509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.748171Z digest=sha256:8e8c2955241def887e84ba6e66a23bf8c72c0118a34d32eb7a2bb6c138b2d1dc

Observation 6595ae70-6712-4d79-9955-54747d398513 · outbound

This paper cites Large language models for code: Security hardening and adversarial testing,.

On the Robustness of LLMs' Internal Representation of Code Correctness Large language models for code: Security hardening and adversarial testing,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.082713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.754549Z digest=sha256:ee2c001dcf79c2bd72d381302ae6712de11055c201553938c91f5eb795a858e7

Observation 4787b928-da74-40c0-8b65-ba750ea9d413 · outbound

This paper cites A Mixture of Linear Corrections Generates Secure Code.

On the Robustness of LLMs' Internal Representation of Code Correctness A Mixture of Linear Corrections Generates Secure Code

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:15:28.059625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.759902Z digest=sha256:df0ee7ea648c4c0c468b16837b77bb969bd38426a3c630fdaa3ca41e38b4fb00

Observation 91bdffa1-6118-4ea0-be7b-7aab2a813383 · outbound

This paper cites Steering large language models for vulnerability detection,.

On the Robustness of LLMs' Internal Representation of Code Correctness Steering large language models for vulnerability detection,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.063678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.765980Z digest=sha256:0ccc89687af89d9507633773b151da5316f1e0b4ca4c877a53d5f00b7d881008

Observation edf809b3-65eb-45f6-8e16-3ebf8a04af48 · outbound

This paper cites Are Sparse Autoencoders Useful for Java Function Bug Detection?.

On the Robustness of LLMs' Internal Representation of Code Correctness Are Sparse Autoencoders Useful for Java Function Bug Detection?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.770578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.770578Z digest=sha256:d49c10be18ac62342a59433d16b6cff4be312c4ae34fc7c02cdb491a5ea52c81

Observation 344778ce-3ac8-441f-9c03-a599cab63584 · outbound

This paper cites Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification.

On the Robustness of LLMs' Internal Representation of Code Correctness Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.776270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.776270Z digest=sha256:3370c60cc6dc04dfbe99242c7de7b655f771786498faf6f4b0cdc18862a44f43

Observation 313d14a4-4f59-4cc0-ae99-8c5a7c4ceba2 · outbound

This paper cites No answer needed: Predicting llm answer accuracy from question-only linear probes,.

On the Robustness of LLMs' Internal Representation of Code Correctness No answer needed: Predicting llm answer accuracy from question-only linear probes,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.783096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.783096Z digest=sha256:cda8e12a17c580133f98714a5da3b8605dc3340489bb031d70afeddd8357cfaa

Observation f40d70a3-dee2-49ce-982f-779c15591e09 · outbound

This paper cites The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models.

On the Robustness of LLMs' Internal Representation of Code Correctness The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.787692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.787692Z digest=sha256:e5811e5db54231b34c8bd6e27d3fa7f9af8b8126d053a9861845c0b387a1b4d1

Observation e204c350-10ab-4e78-952a-f3bdaafc945f · outbound

This paper cites ReCode: Robustness evaluation of code generation models,.

On the Robustness of LLMs' Internal Representation of Code Correctness ReCode: Robustness evaluation of code generation models,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.046534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.793486Z digest=sha256:e606e5430fee445369deb7507162dbda1fc12eba8b4ec4741c7f0b8d54a3fad7

Observation 4fca3363-7114-4e61-9223-9b1fec745fc3 · outbound

This paper cites Semantic robustness of models of source code,.

On the Robustness of LLMs' Internal Representation of Code Correctness Semantic robustness of models of source code,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:27.798858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:27.798858Z digest=sha256:f69953c229f1e0611bedc4052e388789a6d7beb749496164116797b3cacdc47a

Observation 685770ba-b533-4edc-8bc9-14f6fac4d64e · outbound

This paper cites Contrastive code representation learning,.

On the Robustness of LLMs' Internal Representation of Code Correctness Contrastive code representation learning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:29.019004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.803943Z digest=sha256:fd83e23da5b4a99aed195dc581668ef5e538a6df1b4a0be0cfe12c0e914aa3da

Observation 3768ae7c-8ebf-4c5a-bc02-f04928fba6f4 · outbound

This paper cites An automated methodol- ogy for generating labeled datasets of semantic errors in code,.

On the Robustness of LLMs' Internal Representation of Code Correctness An automated methodol- ogy for generating labeled datasets of semantic errors in code,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:28.993340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:15:27.811027Z digest=sha256:1bc897611407c39aa4579f0cd0acc6bc98c99edb93c543fc91397bdb35e9e57e

Pith citing papers

No inbound Pith citation observations are available.