Pith. sign in

Paper Citation Record · LEDGER

Evaluating Large Language Models Trained on Code

As of 7 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 100 inbound Pith citation observations for arXiv:2107.03374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.03374 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 2282 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch2

External citation measurements

1427
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6775eca7-a7e3-4cf9-b02b-b213ab340f21 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

Evaluating Large Language Models Trained on Code wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:24:27.750093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:498bd7b3a1487e3d890a18310cc98ee58f13fc8ba1eae577c72a223adedbc940

Observation 9c79fa98-a5a9-4777-9499-944352cfa36b · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Evaluating Large Language Models Trained on Code Generating Long Sequences with Sparse Transformers

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T12:24:27.756592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:20840a88c72776014802df7d2e3eaaaf1c587ffd14b2c688228ba58613e4b3fd

Observation 091eb877-cb95-4ef3-9bab-7174aa259756 · outbound

This paper cites Clarkson, M.

Evaluating Large Language Models Trained on Code Clarkson, M

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.748305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:973f3884c8b1a6569d9184cbc385ebc9db0f32d21dc3126f0aeb41fca3587f43

Observation 62706afc-d190-49d0-a0d6-d3a3bb9ae4d1 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Evaluating Large Language Models Trained on Code BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T12:24:27.258661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:8ce83b28f64de89a46eee5e4b896571f407a539c8e20c56f32eea42be6c4ad58

Observation db5c3b46-2d3f-4b2f-ac9d-447c9c050d76 · outbound

This paper cites an unresolved cited work.

Evaluating Large Language Models Trained on Code Unresolved cited work

Reference 5

Resolution
parse uncertain
raw_fallback, observed 2026-05-24T12:26:12.742901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:c44d942d01a0f3fc043c10bd881d06a718c0ec77dc34f7894491ac706421b074

Observation 1fcff2c7-3515-48fb-8c92-f5a2c60e3e84 · outbound

This paper cites Number of elements are less than k.

Evaluating Large Language Models Trained on Code Number of elements are less than k

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.738653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:f6040dc0911d97669bd7bba5c6c0c7074e0e61a6e959fce714e8a0e7c5a31863

Observation 34ff45a2-277d-4626-996f-3099c9723d85 · outbound

This paper cites "" ### COMPLETION 1 (WRONG): ### return x if n % x == 0 else y ### COMPLETION 2 (WRONG): ### if n > 1: return x if n%2 != 0 else y else: return.

Evaluating Large Language Models Trained on Code "" ### COMPLETION 1 (WRONG): ### return x if n % x == 0 else y ### COMPLETION 2 (WRONG): ### if n > 1: return x if n%2 != 0 else y else: return

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.733868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:6dbf855e40345fc2cb4228e6a649a69f249a47f3ff4dcd90cb5e99aac00f2701

Observation 848189d2-094e-41f1-824f-8b42eec8b8b7 · outbound

This paper cites remove all instances of the letter e from the string.

Evaluating Large Language Models Trained on Code remove all instances of the letter e from the string

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.729768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:62ed5b746f2fbc1b5a9f8471cb1305839720396a4e7e4cdd7bad08d233ba9837

Observation ada604cf-f260-47ec-8b10-c2e892e17738 · outbound

This paper cites replace all spaces with exclamation points in the string.

Evaluating Large Language Models Trained on Code replace all spaces with exclamation points in the string

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.725434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:05904e15aa2e2668e236f5b3ca3fc74dc1b9080af462796000fb102fd00e41f6

Observation 666f331c-d1e6-40ff-a969-95deaf213e16 · outbound

This paper cites convert the string s to lowercase.

Evaluating Large Language Models Trained on Code convert the string s to lowercase

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.722332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:8777e160e1c16934a11fa54c1ae6118648f4178e91f6afc5f84ade271fcb72aa

Observation 0b3cb82b-438f-4526-9b40-14faeb483442 · outbound

This paper cites remove the first and last two characters of the string.

Evaluating Large Language Models Trained on Code remove the first and last two characters of the string

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.784357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:1036269ac14c2e51c01c54d8b206e2176b4b177fee310f41bbca35fd67bc27f2

Observation c13ab40a-4630-4b7d-8e62-b24d5e35edaf · outbound

This paper cites removes all vowels from the string.

Evaluating Large Language Models Trained on Code removes all vowels from the string

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.780369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:98d3721843c667e0cb8c498845c8e65c5f7f15eb09457cb1add10dcafeeb31b7

Observation 207afd72-1ae0-4b41-bf1c-141de88ec2bc · outbound

This paper cites remove every third character from the string.

Evaluating Large Language Models Trained on Code remove every third character from the string

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.776285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:2b8bec7c2124002fb3e95a3faa860b3b06d5c326da0b01eb19c0ae4d70722c35

Observation 538c5d86-d1f0-4171-8b32-492f2e72f9ef · outbound

This paper cites drop the last half of the string, as computed by char- acters.

Evaluating Large Language Models Trained on Code drop the last half of the string, as computed by char- acters

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.772880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:f45079d1121a76143ccd67ddf5c9a53ff99a945f3241ea9df186b9b80befe8de

Observation 0970941c-eeda-48cc-8922-9bea0310597d · outbound

This paper cites replace spaces with triple spaces.

Evaluating Large Language Models Trained on Code replace spaces with triple spaces

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.769451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:f75aca1f5d4e171572206d8e859c3945d0c9c9ea6ca3124af5f5704a9e1815c1

Observation fd7e41d6-b3c5-44a9-9338-36224d03dadc · outbound

This paper cites reverse the order of words in the string.

Evaluating Large Language Models Trained on Code reverse the order of words in the string

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.766230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:4bbdecdccfaef6a1ca41157beab15b4ba15726dae57e476b1e0e6d30e2d40d9c

Observation ab3777a8-d2c2-4ca3-8f0f-7bd8d4a6e78a · outbound

This paper cites drop the first half of the string, as computed by num- ber of words.

Evaluating Large Language Models Trained on Code drop the first half of the string, as computed by num- ber of words

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.762741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:8671762cc4fc1c1beead726ddb542f97096ce5f1151019f57c38ea3cfd76910d

Observation 3370eca5-4706-4b28-bd92-49941edc7f38 · outbound

This paper cites add the word apples after every word in the string.

Evaluating Large Language Models Trained on Code add the word apples after every word in the string

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.759477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:e635c21e344d01e064ea640b1d84511592ffc4fd25b8f004db9adddf8a807e36

Observation a7d4dd1f-1f8f-4d10-b25a-efdfc37df6d1 · outbound

This paper cites make every other character in the string uppercase.

Evaluating Large Language Models Trained on Code make every other character in the string uppercase

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.755772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:678bb38a92f9badca18a2b863121963b81df8c5e55edab89fac09f4026b4c6b5

Observation 2558f076-28ad-4cd5-bb74-1da50af953fb · outbound

This paper cites delete all exclamation points, question marks, and periods from the string.

Evaluating Large Language Models Trained on Code delete all exclamation points, question marks, and periods from the string

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.752040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:3cfaef4699bd14194d58ea8692facd32e4ba90693b71b041e7d7e7c522971503

Observation 4652fab6-c693-4529-b937-a0249be20f49 · outbound

This paper cites When the prompt includes subtle bugs, Codex tends to produce worse code than it is capable of producing.

Evaluating Large Language Models Trained on Code When the prompt includes subtle bugs, Codex tends to produce worse code than it is capable of producing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.747674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:21c1b1f34d829ea087d072e42a2faedef9ebfd112e10bd702aac8cab9fd64701

Observation 8eb67d4b-80ac-40f0-ba59-a3948cb7a762 · outbound

This paper cites 10") 10 >>> closest_integer(.

Evaluating Large Language Models Trained on Code 10") 10 >>> closest_integer(

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T12:26:12.743176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T12:19:56.526100Z digest=sha256:d7702752be2efa5e59856fa90eb0750107b020ca863d85543480b8ed1f6c4327

Pith citing papers

Observation ede0bc2b-2c2b-4356-8d5b-574a40ce4c17 · inbound

Measuring Coding Challenge Competence With APPS cites this paper.

Measuring Coding Challenge Competence With APPS Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:11:36.721130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T17:11:36.563285Z digest=sha256:f574b6c60e3a7e2617cbb17291cc4f8734fd1fc74173d8a3acb7bd4d209a3ce8

Observation 5fb2b0ab-3cf0-4775-908a-77122bc48125 · inbound

CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation cites this paper.

CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:23:26.392243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T11:23:26.272169Z digest=sha256:d45e83c693031b507d09ae2215359bdfff171a47470abab9e3a940c96b17988a

Observation 8167052a-8ea5-4642-8a92-b2e5c8ed9a2d · inbound

TruthfulQA: Measuring How Models Mimic Human Falsehoods cites this paper.

TruthfulQA: Measuring How Models Mimic Human Falsehoods Evaluating Large Language Models Trained on Code

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:48:53.905185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T21:48:53.535496Z digest=sha256:3584294a3f25f2e52a259ca8c99ec6a9c7c4dd48b5ab4351bdb66f070c759308

Observation 0208c989-8cfd-40fc-b419-10a23a247863 · inbound

Show Your Work: Scratchpads for Intermediate Computation with Language Models cites this paper.

Show Your Work: Scratchpads for Intermediate Computation with Language Models Evaluating Large Language Models Trained on Code

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T00:31:41.587817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T00:31:41.558831Z digest=sha256:2c49d50552a0aca0a9f30d9d4f5f10b876c3d514f04b182f1e6d7339aed416b0

Observation eae23089-4d58-44f0-81ef-ef055ca36342 · inbound

A General Language Assistant as a Laboratory for Alignment cites this paper.

A General Language Assistant as a Laboratory for Alignment Evaluating Large Language Models Trained on Code

Reference 218

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:22:59.110008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T14:22:57.925354Z digest=sha256:57410455d98f1e7ca8e49c48a8457e1f744d7b3777c7964420bc63e737264ebd

Observation a8f5e8b6-cb14-49b3-874c-93d6f5f17628 · inbound

Ethical and social risks of harm from Language Models cites this paper.

Ethical and social risks of harm from Language Models Evaluating Large Language Models Trained on Code

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:24:30.010341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T18:24:28.835688Z digest=sha256:15d6b4a204af014f1c92fb467b93162fa5721ee74fb16afea030f740a5360a5a

Observation 9859903a-28f8-4f62-a559-12d723f3133b · inbound

Text and Code Embeddings by Contrastive Pre-Training cites this paper.

Text and Code Embeddings by Contrastive Pre-Training Evaluating Large Language Models Trained on Code

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:24:11.939164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T19:24:11.907204Z digest=sha256:e4d726fe3c31addc7ee0389633bf08039e79249efcaa5615d00a481845d9c6d3

Observation 5612270d-b438-4009-bd9a-5f6a12d4b0be · inbound

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models cites this paper.

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models Evaluating Large Language Models Trained on Code

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:54:44.729636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T12:54:44.636760Z digest=sha256:70e45e6fbeb749b0b463fcc47dd246350e5854921558e76b77df963247b502b9

Observation b3647d15-340e-40b3-b08b-3b98bd74bf7b · inbound

Quantifying Memorization Across Neural Language Models cites this paper.

Quantifying Memorization Across Neural Language Models Evaluating Large Language Models Trained on Code

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:04:59.704660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:04:59.678438Z digest=sha256:93af3913f26114afffaabafeeaead8bab048c9540f32ee0917bd376a7873e67b

Observation 826508b5-3646-4866-85e0-2ffb1d738ed8 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language Evaluating Large Language Models Trained on Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:50:00.637885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:400acadc3318b40adb0c2e533ef0f38a0239e203333b7ca8b216cb4258fc956a

Observation adc4bb5e-3c9a-4c4c-9a81-7d17b4343d4c · inbound

PaLM: Scaling Language Modeling with Pathways cites this paper.

PaLM: Scaling Language Modeling with Pathways Evaluating Large Language Models Trained on Code

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:07.121538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T23:45:06.755839Z digest=sha256:cc559aa6f22a763f37401055ab3bab01367362ca06c8d122281c83a432d10048

Observation 5642cd0b-22ab-44dc-99d7-cd88a40cb1ef · inbound

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback cites this paper.

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback Evaluating Large Language Models Trained on Code

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:35:55.981926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:35:55.949167Z digest=sha256:5520f88ddbdf29ad86f285d10ac0e473ac8f899608d1de651916e2d6a35d480e

Observation eb5cf1b4-3081-4255-a808-1e9d10bf9771 · inbound

InCoder: A Generative Model for Code Infilling and Synthesis cites this paper.

InCoder: A Generative Model for Code Infilling and Synthesis Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:21:20.482326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:21:20.438666Z digest=sha256:c79f8277fb313ef43ccbc4c1c1a4986c5e8c8ef96124aa55503380c712c8c0dc

Observation 9499c658-d77a-413c-9d6e-34664636ef2c · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Evaluating Large Language Models Trained on Code

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:34:28.341949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:2eb7b936f0c39bee409faa71eb77372502a8401f14699b214cc2bd08a64aa070

Observation 2484b4e1-55a9-4e46-88ad-a885fe98a18a · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know Evaluating Large Language Models Trained on Code

Reference 293

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:42:48.066073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:14b86a2df3341dac4d9a2a7738aed38e42d6b7b5898798f0247d149b431c05e8

Observation 01a60a53-9413-4c31-ad2c-f6edc4d484ec · inbound

Inner Monologue: Embodied Reasoning through Planning with Language Models cites this paper.

Inner Monologue: Embodied Reasoning through Planning with Language Models Evaluating Large Language Models Trained on Code

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:10:45.384454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T20:10:43.912935Z digest=sha256:b533386429bf4059d6d39f9d0662504cb2dcd4ed2b99965775f902ccd40ce7cd

Observation bb97fade-0110-464f-970f-bf1335bcdfe2 · inbound

CodeT: Code Generation with Generated Tests cites this paper.

CodeT: Code Generation with Generated Tests Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:55:14.770594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:55:14.711111Z digest=sha256:b16f25db440a107cbc6d572fa83f27cfc97e430cc9aab57dcbddcee905f1671e

Observation 06c16986-b310-4bc6-955a-669e63512121 · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle Evaluating Large Language Models Trained on Code

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:40:41.872622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:2c744773ec6d58c95729ecd3b8040e527a01c46b70af891cdb23ee9ac11aa142

Observation 74b0ac66-0753-4936-9023-05fada10f7cc · inbound

Code as Policies: Language Model Programs for Embodied Control cites this paper.

Code as Policies: Language Model Programs for Embodied Control Evaluating Large Language Models Trained on Code

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:38:02.850124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:38:02.723424Z digest=sha256:6a5aa3df8058b85f812ef8369ef800356097f86300210d180e66b53c095ed943

Observation 36599aaf-edd3-4eb9-bc9a-94c66fab9273 · inbound

In-context Learning and Induction Heads cites this paper.

In-context Learning and Induction Heads Evaluating Large Language Models Trained on Code

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T03:49:09.865109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T03:49:09.374351Z digest=sha256:2fd48a435c0f53095c97152e551b61013fc1b01ea853d1025eb31f4d151237b6

Observation d449c4a6-eda1-407a-a4b3-badd8195d43d · inbound

Automatic Chain of Thought Prompting in Large Language Models cites this paper.

Automatic Chain of Thought Prompting in Large Language Models Evaluating Large Language Models Trained on Code

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T10:39:17.061258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:39:16.997741Z digest=sha256:d8587d2a946752a534211ee39dd6c1a4d97901510a9a6c1917cc3563e9610bbf

Observation 6dccc3fc-f32c-40ab-ab2f-3159e18ec2cb · inbound

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them cites this paper.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.837037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:589abca51e67a028260a50217e8813d64ffdcc6a23ba812cf58d6a28367af9cc

Observation e5181052-74b9-4b62-85b0-888dd0143191 · inbound

Large Language Models Are Human-Level Prompt Engineers cites this paper.

Large Language Models Are Human-Level Prompt Engineers Evaluating Large Language Models Trained on Code

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:43:26.424082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T09:43:26.288866Z digest=sha256:25af40371db26ae0152fdc0276f8f751536264db5c6d1c7685a305ba71d0326e

Observation 6764f59b-4615-411b-b4e6-a484b474e8fc · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Evaluating Large Language Models Trained on Code

Reference 216

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:51:11.530545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:505652dae27132682f43bbb0897aa429d33fba191a8a5047fe2527aedaa2af18

Observation c3fcae19-54a6-4b3d-9b9f-fab6720b085d · inbound

PAL: Program-aided Language Models cites this paper.

PAL: Program-aided Language Models Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:02:50.191110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T05:02:50.117120Z digest=sha256:18700c3a150888d08dc2156b8230abd4a31cf9ae693a85450344b8c534041e7b

Observation 82c4473d-ed10-4232-b015-4be5bb83161a · inbound

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks cites this paper.

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T16:48:28.104925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T16:48:27.918334Z digest=sha256:8b096cda33138f0ff30fea07b54d477f1e34a1386f7282f1e379c2c819fb73f2

Observation 5ec2040a-7fd2-4c63-8284-31f88fc54b0a · inbound

Solving math word problems with process- and outcome-based feedback cites this paper.

Solving math word problems with process- and outcome-based feedback Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-24T11:14:23.183555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T11:10:40.864420Z digest=sha256:b789018ab7294b70e8400ba7904de45aa64210af929ace09dcd610df214c0611

Observation 52982979-213a-4ff0-be26-3f435df510de · inbound

Accelerating Large Language Model Decoding with Speculative Sampling cites this paper.

Accelerating Large Language Model Decoding with Speculative Sampling Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:29:36.316770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T07:29:36.205026Z digest=sha256:99a20328a63780355a00e0045dbce9dca672b237a7c4a8746137cc3b3cbe1a46

Observation 47e3fefa-0ee8-4987-8652-9fbb600fda91 · inbound

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents cites this paper.

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents Evaluating Large Language Models Trained on Code

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:27:40.697150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T03:27:40.524895Z digest=sha256:8db56e59cbaeb5498dba630b2479c6649459e03280062e729f3a4370f49c15d6

Observation ecc77b02-af9a-448b-8d25-13668bdbb813 · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning Evaluating Large Language Models Trained on Code

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T18:15:14.524122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:bf5281d443d6686a549ecb5acaa39bcba011bd0cdeced692e9bbeb1673e0f46e

Observation c7eae23b-fb87-4e08-8650-99cff9ad5410 · inbound

ART: Automatic multi-step reasoning and tool-use for large language models cites this paper.

ART: Automatic multi-step reasoning and tool-use for large language models Evaluating Large Language Models Trained on Code

Reference 154

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:03:06.260961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T19:03:05.597295Z digest=sha256:af72198f87774df193a6749051b0f36b1590cf061cb7e4559a7cacd7ba6e476b

Observation 4c563659-1881-4c50-b6b7-c05579685761 · inbound

Reflexion: Language Agents with Verbal Reinforcement Learning cites this paper.

Reflexion: Language Agents with Verbal Reinforcement Learning Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:51:41.983927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:51:41.915864Z digest=sha256:2eddbb6a6ef26a6e3b9fb81656eac6abd679ffd44867b136ebf03bfcc96e3a82

Observation 1538bedc-91a6-442b-98e7-90d2e42c6447 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Evaluating Large Language Models Trained on Code

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:19:46.618158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:9bd0f8be837b393440a283c21d14fc77f4134e3ae18b2e4308ed5d03746c6081

Observation 2f5eb956-cc35-4f21-b392-fb867a056493 · inbound

Self-Refine: Iterative Refinement with Self-Feedback cites this paper.

Self-Refine: Iterative Refinement with Self-Feedback Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-10T20:47:39.779618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T20:47:39.476572Z digest=sha256:6097e57ce7c70785c7df713e9fc66e61a6f2fbbdc87bcbdbfe26862eaa6ee5fb

Observation 32eca035-f54f-44af-a2e2-d5e3466ffe5e · inbound

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society cites this paper.

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society Evaluating Large Language Models Trained on Code

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:40:53.549399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:40:53.351795Z digest=sha256:1e350b40bbcf9a95056b75aeff623ddcfd087bde9279f515507aefb1613f367f

Observation b7094ea0-7540-476c-9035-18fbb3431e0a · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Evaluating Large Language Models Trained on Code

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:46:39.977136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:d6074ab42a965cb0c8add9cd601a2cacb5aa213ac3450aa89211dcab92ee2556

Observation 90bc66e9-a994-42a3-81b1-f11b4bd08253 · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Evaluating Large Language Models Trained on Code

Reference 174

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:45:17.643269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:d982bbf30cc7a949ebc5837d13e30ebbec9eca3444b00078ff4bad0351fa271b

Observation 471a6759-2deb-4cb2-b67d-8222d50fdd44 · inbound

Teaching Large Language Models to Self-Debug cites this paper.

Teaching Large Language Models to Self-Debug Evaluating Large Language Models Trained on Code

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:24:25.019335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T06:24:24.607354Z digest=sha256:07d86916f6ddb4ce82b2f3a156f825459fea8fb8dc15f6c416ca48cfeb3dc0d8

Observation ef3ea066-0f13-48cd-8186-9dc0bbfc0b41 · inbound

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs cites this paper.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Evaluating Large Language Models Trained on Code

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.220201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:537f0086dfb2559b9bfed95823cb21f13e3ca82b03e83d30594c7a653a81aabe

Observation 3142dfeb-939e-48b8-b326-b526bf136cc4 · inbound

Usenix'23 Extended Version: Smart Learning to Find Dumb Contracts cites this paper.

Usenix'23 Extended Version: Smart Learning to Find Dumb Contracts Evaluating Large Language Models Trained on Code

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:44:18.617274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T09:41:00.172667Z digest=sha256:4b3fc345c2ab818381c251e1f93fc6ccb40cf508b54132fbc9f57e4f92b7de27

Observation 6515b73d-d5ef-4b96-8183-8eef4c086a3a · inbound

LLM+P: Empowering Large Language Models with Optimal Planning Proficiency cites this paper.

LLM+P: Empowering Large Language Models with Optimal Planning Proficiency Evaluating Large Language Models Trained on Code

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:36:18.618524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:36:18.342052Z digest=sha256:b7d79f911a1042b7c7bd766614580041cdcaab35d9e824547dfffcf406029fd3

Observation 6317b66e-5a3f-4d89-9c99-943eaed98f99 · inbound

WizardLM: Empowering large pre-trained language models to follow complex instructions cites this paper.

WizardLM: Empowering large pre-trained language models to follow complex instructions Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:28:25.033116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T07:28:24.827546Z digest=sha256:e8df43c05057ee8b1691e678a3cb82fceb96b43c1c883d6c857691ff76a45142

Observation 5f0de9db-4d67-41a8-ad68-edb9b52b77f3 · inbound

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation cites this paper.

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation Evaluating Large Language Models Trained on Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:04:17.959701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T02:04:17.888553Z digest=sha256:838bbee662340b932ae9e4461581c493b92256fb83ef97942f210a1cee0ea696

Observation d0565ccb-3356-4837-804b-e3faacff4760 · inbound

CodeT5+: Open Code Large Language Models for Code Understanding and Generation cites this paper.

CodeT5+: Open Code Large Language Models for Code Understanding and Generation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:26:57.533071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T05:26:57.440959Z digest=sha256:98f497baeee4805de91c26f9afd07e224c96cb245e283002deb8dce796432aa0

Observation 17618bc9-caf6-462c-b5d3-9d1332ad2364 · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Evaluating Large Language Models Trained on Code

Reference 264

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:59:27.436795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:e71c73ac5d6f041197face853d6a7c030813693420426876075a16f4f5a2b30d

Observation 6567f0fd-bfbd-4b6f-9153-a723e5876d83 · inbound

Gorilla: Large Language Model Connected with Massive APIs cites this paper.

Gorilla: Large Language Model Connected with Massive APIs Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:22:17.241725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T23:22:16.879418Z digest=sha256:c57a9f438ed846000ad3a7c0abd52cb8af301373669dcc9ec6f8a4397ecf824d

Observation 2d3d2d01-8e4e-49e1-8a9a-d1204b262bbf · inbound

The False Promise of Imitating Proprietary LLMs cites this paper.

The False Promise of Imitating Proprietary LLMs Evaluating Large Language Models Trained on Code

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:54:31.436199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T06:54:31.175090Z digest=sha256:369eae429c4d54f4cacbf5165883002b5b913a40d8185a85d73cf6620b5d5495

Observation eb29a2bb-cb55-4441-bd02-72f755dda231 · inbound

Voyager: An Open-Ended Embodied Agent with Large Language Models cites this paper.

Voyager: An Open-Ended Embodied Agent with Large Language Models Evaluating Large Language Models Trained on Code

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:11:41.133871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:11:40.995345Z digest=sha256:47245ea153545134b3db1c45cf335ea4eea93be11e8c5a5051662fb52eba59e6

Observation 112537c4-7ddb-4d52-aa96-e2f4b5222e23 · inbound

RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems cites this paper.

RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems Evaluating Large Language Models Trained on Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:29:52.601821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:29:52.457841Z digest=sha256:211f426f4711ead3ba4c1ee242ce2ea54d26ec40f44acecf524fe36ce9f4d519

Observation a4505964-88d7-4bcd-8a49-358f75ddea94 · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.155000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:53b21c101778191f35c3f28922031a370cb4c3af971d2a0d0e390c556a79aeff

Observation 3ce7a911-ee76-488e-9550-e17528221591 · inbound

Textbooks Are All You Need cites this paper.

Textbooks Are All You Need Evaluating Large Language Models Trained on Code

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:44:03.201011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T04:44:03.148223Z digest=sha256:bc7aeba5c4f50d4467b2734ed41c22e0e9db25d6b44fef9eaf04b2a3cb43932d

Observation ea4a1be1-5845-463d-aa61-6400ae1fcd26 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Evaluating Large Language Models Trained on Code

Reference 141

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.466163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:5629c6bb350b9f0e9dd26a7bb485cfb4dfc2cd9eeeceabc5ae95ca84cf1d9e45

Observation d37fbb32-2122-4138-829c-888fe3e76966 · inbound

Towards General Text Embeddings with Multi-stage Contrastive Learning cites this paper.

Towards General Text Embeddings with Multi-stage Contrastive Learning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:33:45.988893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:33:45.855974Z digest=sha256:7d4eaa176f8771e7d369a889aaf70f822b36d42890a469cfa0316ebd963b6ed2

Observation 5fa05cf2-719f-4576-b13f-217c51a461fd · inbound

A Survey on Large Language Model based Autonomous Agents cites this paper.

A Survey on Large Language Model based Autonomous Agents Evaluating Large Language Models Trained on Code

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:03:00.572771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:03:00.340349Z digest=sha256:f31db957ff444c0877f5fce9fc6fa29cf5f652693721c2543d268c944f8a75fe

Observation da439efc-6a4a-4058-bf30-485a9a3ea435 · inbound

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding cites this paper.

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding Evaluating Large Language Models Trained on Code

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-12T20:22:10.695701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T20:22:10.482509Z digest=sha256:5cdb136ca1a82fdbcadb094dc60677d49a9be35bdb5ce4cc60d66b32cd2b5e4c

Observation 05d46dd0-c7a1-4155-8ad5-65f64a18c6cd · inbound

Cognitive Architectures for Language Agents cites this paper.

Cognitive Architectures for Language Agents Evaluating Large Language Models Trained on Code

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:33:44.289372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:33:44.146134Z digest=sha256:85eec3fd7488efd1c386867fb2a13505729e80f1ae792d3e15ede2e6b795d707

Observation 9c3163cb-1430-4a58-b9f6-7db7f94f6a48 · inbound

Textbooks Are All You Need II: phi-1.5 technical report cites this paper.

Textbooks Are All You Need II: phi-1.5 technical report Evaluating Large Language Models Trained on Code

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:18:01.336049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:18:01.244364Z digest=sha256:c03db695bda3fe45d0f0de68e1232699fdb4d06ecf3c2b9b99f52777bcf2efd4

Observation 924c57c6-2071-4f27-aba3-a8b1cca610d9 · inbound

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning cites this paper.

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:46:39.565513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T23:46:39.330438Z digest=sha256:3c80f60dc2486b8241c83f54e04581c199ee21bb11207e484e99789358222802

Observation fa3293c6-1d77-49e6-addc-9ad3407bc999 · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:03:07.724048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:07bbd927d909054f34472c6024f4a21077aa75e1a0cf84ab2ddd9d375ea64605

Observation 99d3b1b9-f0be-452d-939a-be6269844720 · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models Evaluating Large Language Models Trained on Code

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.581478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:25ee168981257231547b5484698dee332905d3344dca82bf3d9a21c0f34317b8

Observation ed5ea7e5-b033-4bab-a875-b2100c4b73a4 · inbound

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models cites this paper.

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models Evaluating Large Language Models Trained on Code

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:07:53.869141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T10:07:53.748795Z digest=sha256:c290fcbf91118411c10461d9ec877bf1058e16fcc56b146f030966ad482e06c8

Observation 4eb5ab85-a01b-4997-add8-ade633abde6d · inbound

UltraFeedback: Boosting Language Models with Scaled AI Feedback cites this paper.

UltraFeedback: Boosting Language Models with Scaled AI Feedback Evaluating Large Language Models Trained on Code

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T16:41:29.153815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:3382af2a966549fc1f2b65b6e25b2c98c8269b407560fa8ab5458aba39c5595b

Observation 486b472f-e689-4023-a6aa-ae37e5da4134 · inbound

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs cites this paper.

Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs Evaluating Large Language Models Trained on Code

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-17T11:11:21.701115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T11:11:21.460613Z digest=sha256:542d4f6b71e1d5df1de20646904a6809255af9d6edb2b59718b69f3bc5e730f0

Observation 1670ac2a-2c7c-447a-b2de-eefb37389b0c · inbound

Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models cites this paper.

Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models Evaluating Large Language Models Trained on Code

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T23:26:04.484065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:26:04.445225Z digest=sha256:571555850e691703edd14e50f74c8ab46b6d7698bba7981c9ebaf277e7c916b1

Observation c43e08e5-d1d3-4a8a-b658-9a665e2ca7be · inbound

Mistral 7B cites this paper.

Mistral 7B Evaluating Large Language Models Trained on Code

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:14:00.083241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T06:11:38.406350Z digest=sha256:5cf6c95da6432e01e2e907ad4f6396b7056d86012e7db94a2d8a8b6caa0474b1

Observation 7b5da458-943f-4ea4-8ff2-cbceb8884217 · inbound

Instruction-Following Evaluation for Large Language Models cites this paper.

Instruction-Following Evaluation for Large Language Models Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:36:00.954215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T05:34:04.002648Z digest=sha256:2d772b8312be6685b7dfb3c9a64b78106dfac6599c4644b1ff193821aa8e3d06

Observation 0fb12c4d-9e66-4641-819d-85b575b9fb02 · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants Evaluating Large Language Models Trained on Code

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-12T15:46:03.458328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:181fef567fd82e03002b97e046082f0539a050026ec526219bc576dfeffae4dc

Observation 86df73b4-9a1e-4a98-b88a-513abe34604e · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Evaluating Large Language Models Trained on Code

Reference 258

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:46:10.030820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:a3c9f7c9e27f9377426606f85df0aec2d10a7cdac6951edc19669883bfd1f5d1

Observation 2aea54e5-25c2-4a7b-92f2-0b9db08103a8 · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Evaluating Large Language Models Trained on Code

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:03:55.449843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:888c3190a3a3bffeeb863c44025025fd40ed3cfe1349c04f62ac58c48e2aa6d8

Observation 392b939c-156e-4b0a-b9b7-a620cad38e2b · inbound

AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation cites this paper.

AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation Evaluating Large Language Models Trained on Code

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T03:58:45.905016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:58:45.824197Z digest=sha256:41d560a4a120d4920e0ff312e4058298ba45f948cb96be63df2f455dc8df1557

Observation 16881213-f8ee-4503-97ac-bbe7fe8677cb · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users Evaluating Large Language Models Trained on Code

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T10:16:43.924142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:84a5106f815575059c33a5fa5fba7f84af6ebbe877ba3dca144ad0a4841aedb8

Observation b6796b5b-97de-47ea-9eff-ceb1d1ecc1fb · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:08:05.947264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:2ab5ef3a4cd38262241e201d7dee958e5623286d0a04c1a2c6a8630bd91c4410

Observation 8837baa4-7275-4218-9361-69eb425cd406 · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:13:53.749937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:b3297ca07f6f00c4e0289adede0f2957bbfd945c17c09b0c394f1b12cabce81f

Observation 34f954ac-2378-40c7-b8f0-952895f46ceb · inbound

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models cites this paper.

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Evaluating Large Language Models Trained on Code

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:50:11.045598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T22:50:06.399707Z digest=sha256:cb9e765914ff700c73f9fca510b8d6beede7ee02665243c2dada3253bae2ce0f

Observation c213a35c-e268-4bec-a6db-a6046d8994d7 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Evaluating Large Language Models Trained on Code

Reference 126

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T10:36:18.334337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:bf3e718d64166a0ad08c8c6ebb78c000141931174f736f36029dc1d3ee99e717

Observation 43f51345-189e-4143-afc2-1be62d58aca8 · inbound

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence cites this paper.

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-10T17:26:38.844119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:26:38.812485Z digest=sha256:53bfbe87505cb835c21cba397b28a3c16cc34fc73ceb0f89a5483a053e4d1626

Observation a71d9742-1d95-49b4-8097-036a2ab81432 · inbound

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty cites this paper.

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty Evaluating Large Language Models Trained on Code

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:15:49.360849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T00:15:49.303458Z digest=sha256:39e10440a5e28e8da2fae0fa22384db3c294b97fd226b99e6badb595f0e825dd

Observation 19c9ef4a-20ee-4e15-978b-e790df07fbf1 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T12:17:53.503196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:94a3c0e5998ab13f5fdbc4047163512b31b1ccbb9abed92d1e4c7055498c8bac

Observation 278ac788-5422-4d60-a9ab-f9e4c9134f90 · inbound

CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology cites this paper.

CodePori: Large-Scale System for Autonomous Software Development Using Multi-Agent Technology Evaluating Large Language Models Trained on Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:58:51.466709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T03:58:32.556725Z digest=sha256:eaa11b6f121b7c7062720b208a85e42e49add41e4b07ad9308521508b6a68eae

Observation da414c94-9568-468d-9a03-b31188e6f4a0 · inbound

Understanding the planning of LLM agents: A survey cites this paper.

Understanding the planning of LLM agents: A survey Evaluating Large Language Models Trained on Code

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:12:57.662687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:12:57.568144Z digest=sha256:fcf37d40d951670cc6027621bcf52f57ac5379ade0c28edad6bc29428be6d597

Observation 699fccd2-077e-417f-8499-5ee20b57a0c9 · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:23:49.630828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:683ae0a98f9de959653644012afef7dc79a42fb26ce1c86f8f2f3d56f7949845

Observation e6709153-ff5b-481d-99be-d9fb52509ddc · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey Evaluating Large Language Models Trained on Code

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:22:54.959354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:866080db0cbdab0d446c5f5c3fa454b977cf551f2f68d4f79dc26094fba3594b

Observation aa7b790d-53b2-4273-8abe-1a9e95f6fe11 · inbound

Massive Activations in Large Language Models cites this paper.

Massive Activations in Large Language Models Evaluating Large Language Models Trained on Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.833214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:09281839dd14c958ce9d072b8719e51a31eeb4a41d83aab46b62bbf98e3f8e97

Observation 97f96a46-0a2e-4bdf-975e-a3427da73404 · inbound

StarCoder 2 and The Stack v2: The Next Generation cites this paper.

StarCoder 2 and The Stack v2: The Next Generation Evaluating Large Language Models Trained on Code

Reference 178

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:28:22.902602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T17:28:22.353355Z digest=sha256:7b8867eb6f9b6e1c829aeb548fbc3c89b2a3e0b46f01355354b5ebac943c42b1

Observation 3090d910-ae40-4419-89ad-3daa61d6b655 · inbound

Retrieval-Augmented Generation for AI-Generated Content: A Survey cites this paper.

Retrieval-Augmented Generation for AI-Generated Content: A Survey Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T13:32:17.284337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:32:17.177021Z digest=sha256:aacc57d2b0a8fa24140b3c617c128cc813aebf128a57ac71adf600f0e120a90f

Observation 40cef2d6-289a-462e-bf7e-66a1147d7f6e · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Evaluating Large Language Models Trained on Code

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:47:27.821452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:900fa4b28f47e5b8211af2c57ba9dea7e6aeaa0ec7d558a07ba4adcaf2654bd5

Observation c386a1a3-9b7d-4d51-83e2-b0868df0439d · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Evaluating Large Language Models Trained on Code

Reference 249

Resolution
verified exact
local_arxiv, observed 2026-05-10T17:34:42.880437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:1559801131060636e283bc41eee44bdb99806258cf9a76fb205d7661dd70013f

Observation dfae1096-f9d7-4dd2-a0a3-620d0a9f95f6 · inbound

Gemma: Open Models Based on Gemini Research and Technology cites this paper.

Gemma: Open Models Based on Gemini Research and Technology Evaluating Large Language Models Trained on Code

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:54:09.024442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T15:54:08.950440Z digest=sha256:bd6610f5becc5446c6ac49b58c27766c54fe295560c22b44dcc45a550edf76e4

Observation 699041f8-d3d4-4e55-a540-635c8858ca67 · inbound

RepairAgent: An Autonomous, LLM-Based Agent for Program Repair cites this paper.

RepairAgent: An Autonomous, LLM-Based Agent for Program Repair Evaluating Large Language Models Trained on Code

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:22:16.561582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T10:22:16.434514Z digest=sha256:d16abca9c222f9cef4600ca27c81de4641516db92cea599b57b7289cc31b9dd8

Observation fd78022a-b462-4c63-9b6c-e89b5843aa7b · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report Evaluating Large Language Models Trained on Code

Reference 191

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T11:44:38.308805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:1d4bbb46082efbb30744b880a04e320c286a37380e592f316da0856ddd4eac1a

Observation 02139f0a-55b8-4909-bfc4-901f25b862aa · inbound

Jamba: A Hybrid Transformer-Mamba Language Model cites this paper.

Jamba: A Hybrid Transformer-Mamba Language Model Evaluating Large Language Models Trained on Code

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T14:11:27.193672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T14:11:27.156350Z digest=sha256:b6030791f34b89e8fed553e2088b4d8a8699c6a4f917dd4fea6bebc17bdcf377

Observation e8b9e1a8-b2e7-43ed-85cb-27d49b46c16f · inbound

Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation cites this paper.

Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation Evaluating Large Language Models Trained on Code

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:23:46.128748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T02:19:23.135463Z digest=sha256:cab3028bee05b178b4bc38824c236ada5cdb9fdda8e320594423f6337a33c887

Observation 0b3b96ea-2cd0-4073-8d08-9f28f1020137 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Evaluating Large Language Models Trained on Code

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:00:53.440386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:b6774b8ddbd417d0b74b3d3840b50fe1539056cf5923bc2703fb76ca8e7a1c14

Observation 4d430552-d115-4ebb-9305-7448586a8337 · inbound

A Survey on Retrieval-Augmented Text Generation for Large Language Models cites this paper.

A Survey on Retrieval-Augmented Text Generation for Large Language Models Evaluating Large Language Models Trained on Code

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-24T02:15:55.529980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T02:15:05.379583Z digest=sha256:8cc45104087426d3ff16e38974323aa068c79275468ea74ba11babfdae4887f7

Observation 07ef32fc-0e55-4c63-95a6-257ed1fefc4f · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Evaluating Large Language Models Trained on Code

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:39:33.320470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:660430e5c0685a3854d5d865f3b201d9524792dc099a5c647be8ad9eb0fe1302

Observation 1ef9dcd8-8003-41ee-bb8f-2402bdfc3d81 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:21:57.919153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:d9064ce54fbb6527e0c2dd8504cba53cbff90a6cd6d414cac27d14d5f7ef4002

Observation 6649110d-78d1-477a-b3b6-ee9a4fc1006d · inbound

Better & Faster Large Language Models via Multi-token Prediction cites this paper.

Better & Faster Large Language Models via Multi-token Prediction Evaluating Large Language Models Trained on Code

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:26:09.825106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:26:09.731664Z digest=sha256:b166af4124c1ce53fedb6e5d92069054c1a11f5c7ad4cdd92b705f4e94904e52

Observation 0b4d7b31-7be9-435f-bf76-02761533340a · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:36:26.993658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:42ddeca122c45f22560c193c3e8778422369fd9ea7d5728aca4b4b4fac8c4623

Observation 6692694d-d432-4fa5-99a5-fb7d40a69996 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Evaluating Large Language Models Trained on Code

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:44:49.699022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:bd652160a2d328fe049a1c20e3357b45b98bfeda492f672e71b2bcb5dda6d5ba

Observation 5a5ca1a7-5fb6-4eed-951a-31f158165c7b · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation Evaluating Large Language Models Trained on Code

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:18:06.481204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:ef9950a741f44dfdbe5bccf0e56b6d44dfb8bd40fa469a64a6e893f77f2e083a