Pith. sign in

Paper Citation Record · LEDGER

RoBERTa: A Robustly Optimized BERT Pretraining Approach

As of 24 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 100 inbound Pith citation observations for arXiv:1907.11692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1907.11692 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T04:47:43.784327Z

measured 151 of 151 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 1743 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:28:07.412371Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:47:48.334341Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42b310f6-a40e-482a-a814-4525d4b58c82 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.567388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:ccc05074fd34f6f22bb59547db1c7ec833d0633c05256db326339891a61f711b

Observation 2869ce48-f10a-4448-baf5-f7d6367d6f22 · outbound

This paper cites Cloze-driven Pretraining of Self-attention Networks.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Cloze-driven Pretraining of Self-attention Networks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.431266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:82fd0c3cb16a28fc3bbe20568f1bfefb69dd4fc957dcefa50e97d535d9def8ba

Observation e3150b9c-1c2a-4bab-b1a1-291f0251c0a3 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.577859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:666fa1e2fff102dccab702d21f0118cb0f71a1975d6aa1145ec1f40cdfb00bc3

Observation 6d93e9f6-94f2-4fbe-b439-2e086ec340e4 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.581367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:cd6641fb7c657f61876cf0ab483a21e1bd116ccf57482d8af01e9f9e2b40dae2

Observation 3c932573-ced8-48ab-9582-33d9e2c391ce · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.583299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:59aad7b99441bf062cba12f9bb6c2a49a5f413c63147d953c1c4494e803d34c4

Observation 168e1a3a-ad70-418e-a2e0-f45aa600d78e · outbound

This paper cites KERMIT: Generative Insertion-Based Modeling for Sequences.

RoBERTa: A Robustly Optimized BERT Pretraining Approach KERMIT: Generative Insertion-Based Modeling for Sequences

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.445584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:11ac4abb926db99fdfaf1f9f76a523ffd5a63f861f4dc45d34cb31fb5b5990a3

Observation 89ec54e3-e658-4f01-8309-e4a3e914bcd3 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.585139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:7e84cd055924f4d898b4b981d3975675ed3e9294190d66724cbbc2e793f28899

Observation 567b5f2d-f1f6-4a90-a0fb-862828d8ff69 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.589631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:7d2aaee83173bf0f1bb899470476b3666939a478345dcc0c42b742a6ba00caa6

Observation 62ad42ae-26ab-4308-a215-283a4bbb342d · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.592974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:b2e6b6d9f4c3dc74f29becdac1c7cb7900c0b57cf3b65c2c65dff21bda77de3d

Observation 1a6f9557-6ae7-4c7d-b8b9-7eed8ec88625 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.595306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:c74c0981ac5636f94677f6c3eef7f2f0a7f6ca93a44d00dbbd506344b37ed14d

Observation 3f5a24a9-38f9-47eb-b8ce-585ac9cabce6 · outbound

This paper cites Unified Language Model Pre-training for Natural Language Understanding and Generation.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unified Language Model Pre-training for Natural Language Understanding and Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.486013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:af0c09052e10556dd5a9ee599ec34e3dc3c23f8f19c11d9bf0c1c480038e6123

Observation 68edb794-3ad0-42c6-abd9-1ee02abe2a5c · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.608115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:403bf15bf1d57f7dc30e7bcbdc953f56b89a4e96e18e9141e16c81342da91bd0

Observation df65ae48-1b2a-41c6-841f-38e4d423b443 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.610400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:181ea6825b4e13d9d655bdb041dabcef12070b387dbdb9b1d1411a7a72ddd207

Observation e91828c3-ebd7-4513-946b-06527a2b343d · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.615108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:06f3f06a9d3964e64f3826b3a53b593e7b15386376d59b67cd7ab42d2b5dea89

Observation 116ed174-51e1-40cb-a943-576394a75619 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

RoBERTa: A Robustly Optimized BERT Pretraining Approach Gaussian Error Linear Units (GELUs)

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.408621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:c391d7827c9905a478cdb9a48142c441eaa337d7805c02bae967cc84f901cbf9

Observation 5f9836d9-fc5e-4ecf-9407-4b2ac2db4166 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.617400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:bf10514a772d309de5e6f9b8d87ebdd8e424de6861752d033ecaa1f1039a87d8

Observation caf21d04-95a5-43a9-a2c7-354bc3a6b94a · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Universal Language Model Fine-tuning for Text Classification

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.438569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:969698940ec6fda2d70dd8014539841af38d88049cdfc5ee39e5517cc09e93f3

Observation 80d780f9-fbfb-4a20-bc38-8b5cf32f6ad4 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.623451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:8298a6dfcbd4b263aaa03becc66af4e4d4a90a438aa696d97260d91011658e4d

Observation a0ca66fe-ea1f-4119-a519-3abbb8b02c74 · outbound

This paper cites SpanBERT: Improving Pre-training by Representing and Predicting Spans.

RoBERTa: A Robustly Optimized BERT Pretraining Approach SpanBERT: Improving Pre-training by Representing and Predicting Spans

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:47:44.471482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:308aec5679e8b0c53a7abd1b1fdf200f8e2a54d360d937457b80e3aa0a13ee92

Observation 61ae08bd-2c4d-4078-9bb3-b91350666f12 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.627157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:ad286c5da60b5c4dd915dc27e1f7f8ec11155de684d19d33853d8ec9b62ca5c4

Observation deaddbce-8ace-4223-a2c5-6dee7cdf9e0e · outbound

This paper cites A Surprisingly Robust Trick for Winograd Schema Challenge.

RoBERTa: A Robustly Optimized BERT Pretraining Approach A Surprisingly Robust Trick for Winograd Schema Challenge

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.494874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:69d4face55efcc048da59a0f5fe47caa2cd43cfec410124849ea00740e193d54

Observation 382a824c-6500-4467-8cb7-f38dac1fcc50 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

RoBERTa: A Robustly Optimized BERT Pretraining Approach RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.381582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:33bb11cd06739e3c237ae51406f512e909dffee4f73fe9bf0b3ef3de9b283d2d

Observation 5f06b988-92ff-44a3-b31c-534cbd3754e5 · outbound

This paper cites Cross-lingual Language Model Pretraining.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Cross-lingual Language Model Pretraining

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.396678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:88a6167bd6cd3a6fc4f6850353ccbb366c90c9aa31c3be48044d2b6a393ecba8

Observation a4d735df-503a-4d06-bd9b-2d0f6f71144e · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.630673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:ebc7114f71a6fb08d3d634f21cc9509b50542d0a2c46ef99c70eee460e534d54

Observation 09067d6e-ede4-4afc-a354-aa376cbfba05 · outbound

This paper cites Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.419233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:247319cb3a5c498b74926ba29c7989196bf9cfefa71f3e64819dd0b70d9dd48a

Observation a30af3ea-b967-4c99-b6a9-64bc633b8514 · outbound

This paper cites Multi-Task Deep Neural Networks for Natural Language Understanding.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Multi-Task Deep Neural Networks for Natural Language Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.425009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:74fd0d303579f3d81bdaa1ffc9decd3c605526640ceb8222535445b60e033a7e

Observation 04286537-edc2-4353-8dbc-abac69270a80 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.634257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:4ac994c2892905c518a13852755b5e9a9d0b33b938208adeee71127f5cb6b299

Observation 8892cedc-ecde-4c67-95fa-f93b5d390e10 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.640820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:c4ef7669753e7a2c3a2692c5921fbe28f47ab812d30e5e20c8d81683848be5fe

Observation 27445044-287e-4323-8b53-92bd6fe48712 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.643803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:41fa59b9e0d30c01810e2270e5b953808bee4f298a167e0dc5c4b1d1b5b0d1b8

Observation c415c121-c53c-4775-999f-b0e8215d603a · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.645841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:ecf799a1145289189326c3f9f90269801c2791e2962a7c4df223f07b2be337b6

Observation 5af010d9-ed9f-4ab7-8f6d-fb8717ae553d · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.519190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:66a258552d3479e46b80ab4fd5c8a3d42600dbd0c0773f4201ac2af98e772046

Observation ff3a61f6-6817-4c46-8ba4-79c9d6d80e7c · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.522217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:53034107a65ffa4e7f68ae86c1734560144cb61506630bc77081e4f3e1c7e41a

Observation 2c33d10d-fc97-4baf-8483-7efcfedaf27e · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.527608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:44ed404d45effdae20ae4bed53c634f1b03d472612891b1f27cd03770eba0084

Observation f7b0ebc4-60a4-4f48-88ec-523716b270da · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.535594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:434d5eab3bcf675b4c379ac039758b53d918d1b892044fca019fbe71b9e791b5

Observation f38e906e-8a5b-43c7-8ff6-09f83a9520f7 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.540196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:3893850c68827efe2c3a0e45288d882f63733377435c5fa594e4119bc3a66efc

Observation 0b31be48-e20b-4984-9de4-40eb5a02a471 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.543813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:4de4f591d62bc263314a26ae9206e0adbc11e03f8693975aa349f3e33405277e

Observation e4801fd6-fd3b-4ba7-8e7c-b38a30490495 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.547454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:08a1db16175cb479d703169a551fc47ad4b6471f3b77f7f6234851bab1373ffc

Observation 745733d3-1c9f-499a-a6cd-0fdfa3ddf07d · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.552186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:f9a41b1cb003e2a5b8dc10f00927bc5eadc1995bd718caf9e45e4cf9383ea5f5

Observation a0fe0227-bd3e-4b92-811e-3392ce391a6a · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.557090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:143d3967bf75d8f431dc8cfde9893e7f8d8e848f4912130eb378de0bd8f2021c

Observation 6c173e8b-0508-4e44-98b8-9d9d714b79f0 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.559545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:49b249f6973bb74ce9cf777a4b8c1a38ca6fa122588547a3cd9d58fb9f90981e

Observation ebda2bb5-a438-4aae-97b8-11362c6302aa · outbound

This paper cites ERNIE: Enhanced Representation through Knowledge Integration.

RoBERTa: A Robustly Optimized BERT Pretraining Approach ERNIE: Enhanced Representation through Knowledge Integration

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.449437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:119cd174553c4d49d55e5c68edd895d39a396d94d949910a5dbbc3c0aba26e16

Observation 79737f3f-50f7-4f74-8619-64cad06624eb · outbound

This paper cites A Simple Method for Commonsense Reasoning.

RoBERTa: A Robustly Optimized BERT Pretraining Approach A Simple Method for Commonsense Reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.459190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:f927719f62965e88a992a3546f3065115917cad245d6aad344a92971236f3612

Observation 8e49a579-998a-42e2-9ad5-b15f14a338a9 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.565018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:c9dde0ba66452d7d8afa19e27fce0edf53c167953e5fe43f8e83ef6bf3221d43

Observation d8fe8a63-2265-4180-8bb7-7bdb5cc32dc1 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

RoBERTa: A Robustly Optimized BERT Pretraining Approach SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:34:11.049162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:e0282cf7db442c2e8aca53f05f2d3b87ffbfe362233d9e070ab73853e01dfab0

Observation 7f57bd08-8109-44c9-a39e-9ac67f346697 · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.575034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:9c54ffe2b0a27c664912726ec02e0b614a3bb010251553994299f451433609b9

Observation a47ce014-e42f-4cb9-b107-931ebce220dd · outbound

This paper cites Neural Network Acceptability Judgments.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Neural Network Acceptability Judgments

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.491477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:49296d4a0f23a402b42f57b4c6867aa2610ec22236b03080a079f450faa4fa31

Observation b401f9b6-cab9-4fbd-a90e-afdb2242bd3d · outbound

This paper cites an unresolved cited work.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-09T04:47:44.569659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:b11b48133b4e27944b9471d4d8c22fc5b79dbaac1f4b15a1e7b81347aaa6596f

Observation 7f265c45-b424-4fae-924f-77db0afb45f5 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

RoBERTa: A Robustly Optimized BERT Pretraining Approach XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:29:27.629409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:4f6cd29e72cb04d294ecd558a5585a485e5c9a6270ce0b98f01353d4e2fc14d9

Observation 7ad21350-3562-4aff-94f0-c08159090a27 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:39:00.104164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:c1d467bdef0ca2f5311be13d12f842c4326e1eeb0e11554179b7ce45f52a8108

Observation 72101c75-c640-4643-afa0-fee0815a087f · outbound

This paper cites Defending Against Neural Fake News.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Defending Against Neural Fake News

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.510678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:04efcb7fcd85bfd3d3943aef808c6005e0d57e889b97d5c5f8175d919ed7b80d

Observation 50ea6f9e-85ef-4729-a8e5-a723b1985122 · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

RoBERTa: A Robustly Optimized BERT Pretraining Approach Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:47:44.516702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T04:47:43.784327Z digest=sha256:ec040d895facfab21fbafd727160378f719e5440be840e4fd82de5eba9e35f71

Pith citing papers

Observation 05ee0b1f-83b4-4a89-8768-f3e2bc07e3ce · inbound

XLNet: Generalized Autoregressive Pretraining for Language Understanding cites this paper.

XLNet: Generalized Autoregressive Pretraining for Language Understanding RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:29:27.514572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T01:29:27.427361Z digest=sha256:e26164ac6a14d873cf9403190a73fb2c3378665b101ba7ef2812f3514a8d10e1

Observation 171e4bee-62fa-47a7-ba21-135b5701fd6f · inbound

On the Variance of the Adaptive Learning Rate and Beyond cites this paper.

On the Variance of the Adaptive Learning Rate and Beyond RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T14:28:07.412371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:28:07.412371Z digest=sha256:c642dd5b2e30af4e6022f6aa6c2d6f92915838ff0da74b5680298d8175d54633

Observation d12daa22-ff2d-4e85-9343-03f53a22e7ff · inbound

On Identifiability in Transformers cites this paper.

On Identifiability in Transformers RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T13:58:53.205312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:58:53.205312Z digest=sha256:3a0d4cdb30f854493778c6bc2d79b96d895b7793edad774393ee0314e4ff090d

Observation 48b3c4aa-e5c5-4376-9009-350b33d16acc · inbound

StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding cites this paper.

StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:42:20.354787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:42:20.354787Z digest=sha256:372886be0751a50db80be1c4ad06c6566137daae24b6c5b744f514bf091680e6

Observation f9e48b4a-e4e9-4f9b-97cd-cad56cee18e3 · inbound

SenseBERT: Driving Some Sense into BERT cites this paper.

SenseBERT: Driving Some Sense into BERT RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T13:14:19.382555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:14:19.382555Z digest=sha256:3951265bbc6eebd3096afb67ace76db667ab13c8eb8bd3065f16a02b6ef28fa0

Observation 6bca7d63-b134-421b-a982-2ef12eeea35c · inbound

Reasoning Over Paragraph Effects in Situations cites this paper.

Reasoning Over Paragraph Effects in Situations RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:06:19.381322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:06:19.381322Z digest=sha256:0c85d9b3f7ce420906b23710481377d3a026328b22dc1e419f4eb4e6e7ec66f7

Observation 646cbc41-217c-4cf3-a704-543d6adf4a31 · inbound

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training cites this paper.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.263107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.263107Z digest=sha256:6f69716c8662e5d8b0a8d1f39ce2905ee0d3261ba855e9cbccd4316785259145

Observation dda9fb00-f13c-48bd-bf24-b83017e768cb · inbound

Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models cites this paper.

Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T12:41:18.607366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:41:18.607366Z digest=sha256:19fca7ecde62570580cf9f91aa2255d48ce07034e8083ed6d895ed156bc28ca8

Observation 86ee30b1-fbb8-49ed-b3cc-a34ff7f85b39 · inbound

A Multi-Turn Emotionally Engaging Dialog Model cites this paper.

A Multi-Turn Emotionally Engaging Dialog Model RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T13:16:31.314216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:16:31.314216Z digest=sha256:050cc8e1c824e4591bde6dbe0105d3c87b44b1f274dac607b6e61fa3767c66f0

Observation d4f101d8-78ed-4421-bfd3-1e72d94d96a1 · inbound

Revisiting Semantic Representation and Tree Search for Similar Question Retrieval cites this paper.

Revisiting Semantic Representation and Tree Search for Similar Question Retrieval RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:30.686995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:30.686995Z digest=sha256:52570da4532cf12cb2e9a174af06313d72543fba2b0cdb06682ffa2d30dd6315

Observation d86dbec9-c9d1-4dbb-ac88-9a26e584d074 · inbound

VL-BERT: Pre-training of Generic Visual-Linguistic Representations cites this paper.

VL-BERT: Pre-training of Generic Visual-Linguistic Representations RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:19.097865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:42:19.097865Z digest=sha256:f8d54dc356af357f53b2873f407831810389784ce5909e4c75544aaf930afe0d

Observation e280b660-4a21-4fa8-9b89-9602b510e574 · inbound

BERT for Coreference Resolution: Baselines and Analysis cites this paper.

BERT for Coreference Resolution: Baselines and Analysis RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T11:25:13.414625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:25:13.414625Z digest=sha256:df2acbab267a860e31896eda429c8260b957b73f4c2bd4ab6fabd2b223f2f6fc

Observation d2cd3a7e-b13e-401d-98e9-4a3367777000 · inbound

Patient Knowledge Distillation for BERT Model Compression cites this paper.

Patient Knowledge Distillation for BERT Model Compression RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T11:18:24.383308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:18:24.383308Z digest=sha256:e3bd33d33a9dff25f77c4b1f8bd9f5975310041558eb60beb18e8556e664603b

Observation 79f13606-6178-4a90-be9b-faabb26ba87e · inbound

Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks cites this paper.

Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:54:08.952302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T14:54:08.729352Z digest=sha256:39de562eb901a78d916fc31b4d0fa4166b0cdd51d3cf7cf50e47f6296e935825

Observation 0af27f8e-1283-4ece-bbbd-a30c734f33cc · inbound

A Morpho-Syntactically Informed LSTM-CRF Model for Named Entity Recognition cites this paper.

A Morpho-Syntactically Informed LSTM-CRF Model for Named Entity Recognition RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T10:53:38.407437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:53:38.407437Z digest=sha256:ccd594e29522a6e738be6fe8f65bc060c84d5a9edb84d577ef7364fd3d952c84

Observation b0c2676e-c278-468c-ab39-cb525b8c5544 · inbound

EntEval: A Holistic Evaluation Benchmark for Entity Representations cites this paper.

EntEval: A Holistic Evaluation Benchmark for Entity Representations RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T06:07:05.483395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T06:07:05.483395Z digest=sha256:20183a97c3007d0ea7f06f81593a27162f54dd2acd8e58511956b61c26ecd046

Observation 4009de4c-b3bf-4e11-8c73-126396f5a89a · inbound

Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations cites this paper.

Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T06:05:46.277374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T06:05:46.277374Z digest=sha256:4cd5aee67571fe3c8f75874baea8db67e9da1447357b689897674031e5878fdc

Observation 9bf0b096-c473-4c5e-a7eb-562368e9bbe7 · inbound

QuASE: Question-Answer Driven Sentence Encoding cites this paper.

QuASE: Question-Answer Driven Sentence Encoding RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T06:01:44.982615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T06:01:44.982615Z digest=sha256:9b2c538b5f89bf6a250e51359bbdf3b266a72aa7f577493a7372882a4a47b3b5

Observation 05d66b7d-5279-4d8f-9c78-72351bd9cba8 · inbound

How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings cites this paper.

How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 180

Resolution
unresolved
no resolver link, observed 2026-08-14T05:53:17.413381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:53:17.413381Z digest=sha256:d4fc968e4f04fe679958b98d54b9f8f152b975657cc6d2a01bf5db64901d0b72

Observation 6194c3ad-fc51-44b6-9e1b-6fdadd7bf8cd · inbound

From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project cites this paper.

From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T05:07:09.481176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T05:07:09.481176Z digest=sha256:0f6650d6dd2f289bfd549d982827c5ece2458994a54cf534c84fd30e62c27ce3

Observation c32810dc-135f-4230-878b-28c7dc03deb4 · inbound

KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning cites this paper.

KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T05:01:52.513710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T05:01:52.513710Z digest=sha256:dfe4db23771d28401810d2e8051065088d2852f851e555cc9574d9c24a902f66

Observation f769aeaa-9d70-465f-85f9-8d9d0a29e9de · inbound

Specializing Unsupervised Pretraining Models for Word-Level Semantic Similarity cites this paper.

Specializing Unsupervised Pretraining Models for Word-Level Semantic Similarity RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:57:43.354683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:57:43.354683Z digest=sha256:9a56f67d70ac759a53d578de663d601022e685b6771546ff045fbc0bf184f7d4

Observation bce180f1-1614-4035-8dc6-6d1a86c910b6 · inbound

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations cites this paper.

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:26:58.111796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T12:26:58.015594Z digest=sha256:43f4d033133de15b29c20edba78a5662ed0b1986c9f8505c7cdb1c91d6529e20

Observation 6be089a8-7631-4c13-bc34-ef4a42a65c86 · inbound

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter cites this paper.

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:02:34.205689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T05:02:34.076464Z digest=sha256:ddf377d4b2efe48dbfcd104a662620dc6658b1224bdcb90c7c3242f77a5eff59

Observation adecd3ce-e71a-429e-84bb-da38e9ee1329 · inbound

HuggingFace's Transformers: State-of-the-art Natural Language Processing cites this paper.

HuggingFace's Transformers: State-of-the-art Natural Language Processing RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 169

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:53:59.753399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T14:53:58.963468Z digest=sha256:7f5eda76e4618dc6b6050dccfcdddd5a181c6e5c58e9a3ab191b04ad2fd8b8a9

Observation 43d5290a-d8c3-4c1d-bd88-290077011313 · inbound

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer cites this paper.

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:37:55.805752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T05:37:55.083206Z digest=sha256:ade17923424b87539dd7769adea6ebefc4eacd236638a0d4fbc33aca1051bf2a

Observation 6ba14bed-e016-47ff-b199-8ead5c8e1ff8 · inbound

BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension cites this paper.

BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T00:14:58.230060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T00:14:58.134513Z digest=sha256:ee75326b9d4e5bfa28a7cf4eefd98cd3b4387d21c964ae4531ce1929555f53ea

Observation a15bcf44-0dea-497c-89f6-b84ce51dc3d0 · inbound

Unsupervised Cross-lingual Representation Learning at Scale cites this paper.

Unsupervised Cross-lingual Representation Learning at Scale RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:23:29.620944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T16:23:29.564169Z digest=sha256:85e62f75407635e010d9136868c2818e613e49add7da825aed67c97f5a1894f3

Observation c1ca71d5-47b9-4de9-9254-a90177e8b7f1 · inbound

PIQA: Reasoning about Physical Commonsense in Natural Language cites this paper.

PIQA: Reasoning about Physical Commonsense in Natural Language RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:55:17.868145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T14:55:17.830588Z digest=sha256:a15b1101dcacbef6c8d42f0e38319a9096e67c5e9c4258080f0d0452c1135a74

Observation 7862ea24-4cd9-413d-a7b0-d95d0bd0e72e · inbound

CodeBERT: A Pre-Trained Model for Programming and Natural Languages cites this paper.

CodeBERT: A Pre-Trained Model for Programming and Natural Languages RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:04:26.301884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T21:04:26.198288Z digest=sha256:acc3f485b5cc2ac78c6c7e6b4422531d1b6d69bc95171b86afdca7c7b3ef2010

Observation 968d8800-7dbe-45d5-908c-9afa78515aed · inbound

REALM: Retrieval-Augmented Language Model Pre-Training cites this paper.

REALM: Retrieval-Augmented Language Model Pre-Training RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:59:16.208479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T09:59:16.120886Z digest=sha256:2f568e2e3ec4562b8755fb8f9216c19a73374adf64336a00c5e72962da3dd71e

Observation 288a2b82-ffb4-4e48-8790-4ecdaee8fd46 · inbound

How Much Knowledge Can You Pack Into the Parameters of a Language Model? cites this paper.

How Much Knowledge Can You Pack Into the Parameters of a Language Model? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:00:28.136236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T02:00:28.055865Z digest=sha256:249ca54cb9d83af03e09e233c31c4299e947c3e0e299cf356c6ba6101795ee9d

Observation bfd1728c-9130-4292-8df8-14286f61db30 · inbound

ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators cites this paper.

ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:26:47.631973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T10:26:47.593122Z digest=sha256:d1e515e14c57306619e6c6bdb940fc49820f9309e1429cbd1c531213082142bb

Observation 60f64f82-3fb1-45dd-8af5-ae50414ee364 · inbound

Longformer: The Long-Document Transformer cites this paper.

Longformer: The Long-Document Transformer RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:29:58.773293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T13:29:58.719341Z digest=sha256:c6f30e2b852b594f3c0edeb32aa6104439786657dae1befabdb310c33126f291

Observation f1eb84eb-3f1e-4d4e-9c99-85be86e57dd0 · inbound

Language Models are Few-Shot Learners cites this paper.

Language Models are Few-Shot Learners RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:05:38.310615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T12:05:38.045330Z digest=sha256:78108be385af6356c3d8c7b90a0f08e0705c5ddf51391d27c1ba1533b2254a17

Observation aa4f25a6-5599-4f67-b9a7-369e05bfe402 · inbound

DeBERTa: Decoding-enhanced BERT with Disentangled Attention cites this paper.

DeBERTa: Decoding-enhanced BERT with Disentangled Attention RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T04:50:53.652209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T04:50:53.587891Z digest=sha256:a3f18eec1fb2c63eb6c4aec05ddae757bd4515e83f2e2e0bc568f4154f4f13b6

Observation 20c51158-4973-40bf-8cef-7339efed92f5 · inbound

Linformer: Self-Attention with Linear Complexity cites this paper.

Linformer: Self-Attention with Linear Complexity RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:37:42.306109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T00:37:42.175821Z digest=sha256:26b121bffa1495df8ade4a573f73b561b80a9b9fccf65117fa874d867a2408ed

Observation 92d47afd-a349-4591-aa06-9df8dcac3499 · inbound

Aligning AI With Shared Human Values cites this paper.

Aligning AI With Shared Human Values RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:41:26.579097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T14:41:26.484482Z digest=sha256:d10ee3f9c0b9d6b9669ffcb0fe9bd6850722f613da9c20d5b5ebd3b386ea87d0

Observation cbd652dd-6c6f-4ade-93b1-30ef4c74e2e0 · inbound

Measuring Massive Multitask Language Understanding cites this paper.

Measuring Massive Multitask Language Understanding RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:43:44.466442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T12:43:44.359247Z digest=sha256:65e7f74110d0931f5dd877691f16c1c7bad7abe498b3f85c5e862b0947ffcd39

Observation 97aec4ec-d04b-4848-8324-a1ea857a1996 · inbound

GraphCodeBERT: Pre-training Code Representations with Data Flow cites this paper.

GraphCodeBERT: Pre-training Code Representations with Data Flow RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:46:10.968140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T08:46:10.892060Z digest=sha256:130d24f79cd14df3a02a5e9e88fe8a2c2cd8d3af603ada2e71141390d3f5e747

Observation 2329a960-8cb7-4f8a-8101-57f7deb6abee · inbound

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps cites this paper.

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T07:39:58.406542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T07:39:58.366423Z digest=sha256:13d5604c4f4d6aec254f81023eb14daad2035394ff8c9efc4c495d2819788d62

Observation b01c135f-f45e-4808-87e3-afecc4db4d6e · inbound

The Pile: An 800GB Dataset of Diverse Text for Language Modeling cites this paper.

The Pile: An 800GB Dataset of Diverse Text for Language Modeling RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:35:18.851942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T21:35:18.513342Z digest=sha256:2efbc7d6938e3bf83763b9238a2f12775e6a18c68439f0bdcf8c1b68937bc116

Observation c49c00c3-2675-4b64-8ba6-0d67683d73a5 · inbound

Prefix-Tuning: Optimizing Continuous Prompts for Generation cites this paper.

Prefix-Tuning: Optimizing Continuous Prompts for Generation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:57:25.078348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T16:57:24.816495Z digest=sha256:37528a16b10d326154effc675f63b25d6af32e41de366fb46b9f4c990d33d926

Observation 481c9497-9c3e-4f10-bf1b-c62ab7471204 · inbound

Scaling Laws for Transfer cites this paper.

Scaling Laws for Transfer RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:58:13.678055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T00:58:13.116663Z digest=sha256:f71d7ff9f24dd74366d660bcb905dc9be08669e4843d82e1ac0d49bea8112493

Observation ff1d9dd7-adf5-41ef-8f1b-8fd6bb712c39 · inbound

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation cites this paper.

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:40:02.694785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T11:40:02.535705Z digest=sha256:f5f6e3642f9ecfe18d0f0588c58395668bab1184a411e19de5b3b1de03038c28

Observation 344b46e6-8fb0-4619-8864-b008b9831455 · inbound

SimCSE: Simple Contrastive Learning of Sentence Embeddings cites this paper.

SimCSE: Simple Contrastive Learning of Sentence Embeddings RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T07:48:23.431202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T07:48:23.186825Z digest=sha256:77a7423f7133bfc61fa9c03093f8e163f63f015239c5d5076637247f557b4f3f

Observation 31d75374-748a-45c0-863e-9d6e6fe51930 · inbound

Perceiver IO: A General Architecture for Structured Inputs & Outputs cites this paper.

Perceiver IO: A General Architecture for Structured Inputs & Outputs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-15T19:47:13.954700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T19:47:13.831505Z digest=sha256:aef93d6a8de3f127e723ae3894afa2ef09e28d4d3e74bc472daa4391c4a929fd

Observation b72bb6f5-77de-4cf9-93f6-495b2aa184ca · inbound

CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation cites this paper.

CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:23:26.373511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T11:23:26.272169Z digest=sha256:39b9d6f5c9b5661530f7045d5d646366d1f8b698ca169be955c79e92fdd075df

Observation e7a923c0-3bb6-4c42-89e1-ffbf286e753a · inbound

BBQ: A Hand-Built Bias Benchmark for Question Answering cites this paper.

BBQ: A Hand-Built Bias Benchmark for Question Answering RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:55:11.543523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T19:55:11.431020Z digest=sha256:bf880adab6e5088c977e0738f25e320bb64ab1638a02f43cac948cfd5fdb05bb

Observation 1767c08f-60e1-4dca-9bef-1e8e5449ef3e · inbound

An Explanation of In-context Learning as Implicit Bayesian Inference cites this paper.

An Explanation of In-context Learning as Implicit Bayesian Inference RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:22:08.030085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T22:22:07.932850Z digest=sha256:3297bac7f3f2e7737569506df2a20410caa7a6766dffd8b225e16953831f6ade

Observation 79b02f68-5813-4951-9ae6-89f10255f95a · inbound

DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing cites this paper.

DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:49:19.466157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T12:49:19.418548Z digest=sha256:0e304452804b385a8c13007bf2764b6bd41a926c890a82e1fa15cc97cb314365

Observation 6b3ef818-bd93-4082-a260-11f72878d21c · inbound

A General Language Assistant as a Laboratory for Alignment cites this paper.

A General Language Assistant as a Laboratory for Alignment RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:22:59.750866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T14:22:57.925354Z digest=sha256:f3d16a56d48f5c89510e04b31bce85fb32e87eb8f4396a42b7603cdb7682b994

Observation ef7f74dc-9afd-4d81-b31c-3ee52a86f798 · inbound

Ethical and social risks of harm from Language Models cites this paper.

Ethical and social risks of harm from Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 169

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:24:30.304103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T18:24:28.835688Z digest=sha256:545143f2b89f01617d3066bcbd41238eb1c499fb3f2c89f0cd934cffdc5ece7b

Observation 3d294c37-3cc6-47f6-a936-0e5ffd212e51 · inbound

LaMDA: Language Models for Dialog Applications cites this paper.

LaMDA: Language Models for Dialog Applications RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:17:32.343076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:17:32.272353Z digest=sha256:5aaa70de946bb846e031abe131a78d17636e7d5b310c4b992fef042ee1223b17

Observation 4eb41430-3e27-4438-a046-df04adb8ab8e · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:14:26.713341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:454bae1653c4422d24ec56cc03a922001cc7b13141115d5aa571322b56dfeefa

Observation 92d4ad6b-4809-4772-9c4f-d0f23723751b · inbound

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? cites this paper.

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 217

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:51:46.860899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T09:51:46.701149Z digest=sha256:3adcfda9440e696ecdf998b26f8612703cd51310419ed8b1744c854deccef9c6

Observation 6066c43b-5730-471b-964d-57f0ed3735cc · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:50:00.675448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:726a64a4cd9d143220ad9f79069e3f11f22f5f3a46732daaf49d7de08a5d1f14

Observation 4c5bec25-687c-454a-ba22-1ec586171bf5 · inbound

MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning cites this paper.

MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:31:08.354783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T07:31:08.266737Z digest=sha256:b869a10be54c6ad7402ef777cbd60f050939289c6cd6239752ff69d10a291769

Observation d3f698f4-2e85-4d37-8709-be57258d6c02 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 128

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:53:17.488684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:ee431ab08973f5839bbf7b88d4fbdc0c8e9da7e065117778f506cb0555bef11e

Observation 5b4ca4c1-2337-473f-a29e-e10e4f698439 · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T20:54:07.679215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:10aa763087a21c4d2cdc782dcf49c9bfa67720c2ff0df51c712f6cd403df1127

Observation a49acf2e-3a23-4476-a862-09e47c138e07 · inbound

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness cites this paper.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-12T16:22:08.976943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T16:22:08.801066Z digest=sha256:1c4676bdf5f7ef017fb23b1866317e93c789d86f0efac5ce34bda2293de9473b

Observation c5f442e9-129c-4219-87b6-447251aa381b · inbound

Language Models (Mostly) Know What They Know cites this paper.

Language Models (Mostly) Know What They Know RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-05-10T15:42:47.797397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T15:42:47.274448Z digest=sha256:9a7b7c8d25d735fde79c30b3bdbed9d61714e89653a3cf5c37783041dd32d9ce

Observation 86451164-a68b-4b69-bf9b-5cd7629706be · inbound

Inner Monologue: Embodied Reasoning through Planning with Language Models cites this paper.

Inner Monologue: Embodied Reasoning through Planning with Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:10:45.515398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T20:10:43.912935Z digest=sha256:4a0d3178a58af3a005f68219f026cf9ced9c9f8c614b89f925534464c5976955

Observation c5d9d2cb-4620-48d1-9ab3-a7531029f6cf · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 146

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:35:36.118237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:6b8abab885be4be36f7669973e76ff856adc98135f56e37416a69f398e0eb2e8

Observation 4c3c38e4-878e-46b6-9587-417542c1c42d · inbound

Make-A-Video: Text-to-Video Generation without Text-Video Data cites this paper.

Make-A-Video: Text-to-Video Generation without Text-Video Data RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:13:03.319494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:13:03.213358Z digest=sha256:a86bb6d424126b0f10cf41610703ea29a864c3426b553949152e9f79746f327b

Observation 1c860e65-bb2f-4329-b537-9e68038e3201 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 269

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:51:11.408748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:87b767f340482fcff14c4c7511d70e8984856901dbc5ebdcdcce14f91d7daa88

Observation cea95ff9-6a54-430c-b18c-5b42053049d5 · inbound

Text Embeddings by Weakly-Supervised Contrastive Pre-training cites this paper.

Text Embeddings by Weakly-Supervised Contrastive Pre-training RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:54:03.944796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T04:54:03.524365Z digest=sha256:542cec77629ee93adb92a9f7973516056658a85ae1ef25e3f9a0d1f998742561

Observation ac205680-71ff-44e0-804c-a37f7454ead5 · inbound

Discovering Latent Knowledge in Language Models Without Supervision cites this paper.

Discovering Latent Knowledge in Language Models Without Supervision RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:34:08.327577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:34:08.207848Z digest=sha256:1c4187c1b78e2172e74eddc67c3d37346fa41ecb38a09aa6948d72ae7287aa88

Observation ac8ae89d-8bba-41b4-af8b-0fe342fc279d · inbound

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers cites this paper.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:31:07.062215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:cf3b555874582bfe5a2bdca1704a95dee55405e403d0b0dba001adf95b4365af

Observation 52d027ef-b15c-4ece-b5d5-ed164e24b9bd · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:32:22.957247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:acaccf3a21d1dcf2d51f5749fdcdbe1468d6c8cb77469261b1004e36e9bc964a

Observation e8cbbafc-2bd5-4f05-b8cc-09cfbc3e149d · inbound

Eliciting Latent Predictions from Transformers with the Tuned Lens cites this paper.

Eliciting Latent Predictions from Transformers with the Tuned Lens RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T16:54:37.606938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T16:54:37.382049Z digest=sha256:fc0aa7ee4f99ba224fa88ae0c2af493c9f4caae3ff34da6db76f4b90cdf2344f

Observation da5a0cee-fd4b-4bce-bc88-dbbdd54adc33 · inbound

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models cites this paper.

SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T15:11:21.851099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T15:11:21.589770Z digest=sha256:97459a040b743d42121122d16951ca8aa9fe76fc2797a1b0b4202326ef289627

Observation 46337eb6-0818-409f-bc46-9786fd46a782 · inbound

ART: Automatic multi-step reasoning and tool-use for large language models cites this paper.

ART: Automatic multi-step reasoning and tool-use for large language models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 90

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:03:06.094667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T19:03:05.597295Z digest=sha256:e8fb10869dc78b8b92aefaecf16e1a18c62beb917d1a85552b2b76b7945159e6

Observation 88d39381-0811-4451-8c72-d1abd95e17e3 · inbound

AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning cites this paper.

AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T21:11:32.297194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T21:11:32.186676Z digest=sha256:e37c2b73e724eae970c533116be202d24adab6fbe275a646bc4f8c5ed40be84c

Observation 9fa662b2-767b-4cf6-a21e-0600bfce9821 · inbound

Can AI-Generated Text be Reliably Detected? cites this paper.

Can AI-Generated Text be Reliably Detected? RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T19:29:45.463076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T19:29:45.315242Z digest=sha256:94838143dd777cb244e6d5e413e704b0f470934bd90bca3a303472caee58d8dd

Observation 7ece3712-0ca9-4b37-a513-c2c821f44d2d · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-14T23:07:42.963015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:8a80a9a44c71ee531ed76b13785f97e31b3d429441b81572dc789cff157f197d

Observation 3c6f3636-3638-499a-a487-0e91323b042f · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:46:40.153349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:6a93cb3c09372d8f6affa9660f3a97a6cb6c58ee1a7239ee0296bfaf0ab40268

Observation d0f7507d-1895-4871-9f0f-31435aeec32c · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:33:00.586348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:399f563d57dac850ac7a019be1ac68f742c071dc83328d0eaa469f633ab661b0

Observation c267a245-13e9-4358-846f-ade85d1eb177 · inbound

CodeT5+: Open Code Large Language Models for Code Understanding and Generation cites this paper.

CodeT5+: Open Code Large Language Models for Code Understanding and Generation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T05:26:57.505376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2ff60779c60d4d0e180c61e2bc22bd9d388e024d832c7ae2de2d83b7ea121280

Observation a4f31d89-4d9a-4010-b621-f17e683c1c79 · inbound

AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction cites this paper.

AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-24T08:49:13.942504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T08:47:23.231930Z digest=sha256:a1f2a984ecbeab18bdb9332a894db53d533530b2d2d4ee608b357ca1b7744d7c

Observation 1f9477d7-9778-4275-aaef-37da67568215 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:29:53.553914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:305366798899107f016a3b39182b614fa063f73849d1f8e2ca0b1201446f2423

Observation a4f2559e-b27b-47c8-862f-426c20e01408 · inbound

The Curse of Recursion: Training on Generated Data Makes Models Forget cites this paper.

The Curse of Recursion: Training on Generated Data Makes Models Forget RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:04:58.322139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T14:04:58.271168Z digest=sha256:add703838029f92515d9faa987e08f845feecb1dc7d052bc6d76343a29dc0aee

Observation fa90ea14-bb78-4e92-abf6-4dd757aa837e · inbound

On Diffusion Modeling for Anomaly Detection cites this paper.

On Diffusion Modeling for Anomaly Detection RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T08:54:15.953699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T08:51:14.163621Z digest=sha256:b6dd740346fb628e63f86da4048b0ca265fc5bad2b43b49319da351609531f14

Observation 50d33e14-78a2-41ed-906c-ba3ec880d4e2 · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:43:45.861688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:8c2deb9ec5a4a122fb17c923e9366f14116d01925e5840100d1fa7313e12ee56

Observation 167bc454-7fb7-4dec-ba75-a270e15993ec · inbound

MiniLLM: On-Policy Distillation of Large Language Models cites this paper.

MiniLLM: On-Policy Distillation of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T17:40:27.925798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T17:40:27.827493Z digest=sha256:5d743fe2d7281c07d1030026c33c9eb0a7ea9080ce3455436c81107893d023ea

Observation 4ab4ba32-40c3-4997-9fb8-2dd2edc23213 · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-17T18:00:50.545885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:d053317edee6125fd3bce6a3ff213b52a09a09391e8c989552e284f1dc43c176

Observation 7128ea1a-92a1-45fe-99c6-d41cf05aaeb9 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 299

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.299710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:f26d9c6233290b22d1a45fa625eb5f50806e5a6787728e688b696701c1d11ded

Observation 92f0be19-f331-403b-8196-c89c96ac4279 · inbound

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models cites this paper.

Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T14:21:16.517753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T14:21:16.453610Z digest=sha256:2e0a4204579ea89d56d1e1be55e437a502d89bafa71c1cb27d59d9f13fe985da

Observation 6ad8b803-3ac5-4cdb-840b-4aa7199d205e · inbound

A Survey of Hallucination in Large Foundation Models cites this paper.

A Survey of Hallucination in Large Foundation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 106

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T15:21:00.841132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T15:21:00.778049Z digest=sha256:821ebe50a3515a7a69deeed532b7478712fb268b2b42f64132ba3b0006f9deda

Observation 2e20841f-11f6-4864-9b5f-0afa446ac21a · inbound

C-Pack: Packed Resources For General Chinese Embeddings cites this paper.

C-Pack: Packed Resources For General Chinese Embeddings RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:24:32.229997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T13:24:32.084878Z digest=sha256:857ed381f038c0c82d1a52f62d4edb9c7678d5028e7d45d793a805f1ab44e7df

Observation 146d8f17-ea43-4389-8177-ed127a83c899 · inbound

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts cites this paper.

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:25:21.128869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T06:25:20.966510Z digest=sha256:1efd01a1e7d5019cbe4397d3e3d7d1c8a8993b2e76bc9fa090a5caa6f590ab63

Observation cedaf0e7-8209-4190-b948-8e2bbcbadc02 · inbound

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models cites this paper.

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:07:53.852306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T10:07:53.748795Z digest=sha256:63d68cc6cb3c7f726a1e1b1be9d9e4dc036bb72f26d7be1af73c75b08ec59384

Observation 509e3640-f91d-42be-81c4-dc17ba7956c2 · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:07:22.326651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:aff97e05a9bfad20933022047771b9169597ac6895ab2b633c78da2589908511

Observation fd91268c-fd01-43ae-bef6-d086622be3c0 · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 172

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:20:20.428956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:0d633d4aa6a7d065b5c8b8ff5e9dcbd8d33c407309cc9f02b718d77e2b5bfb1d

Observation 6b784b93-b640-431f-b7db-2a1dcc6810ab · inbound

Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs cites this paper.

Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:34:01.117192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T06:33:48.456209Z digest=sha256:d8895d3e6061e67fd019a3889497d69d34675da6b0243080a101a75fed47f9d9

Observation a9224d95-b0ff-4a05-93d2-22b848bde2de · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 192

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:06:44.665657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:4f47fb7a0520fcdebcfc61a8350e99fc96d6e0512f4076c37807649e91a0c254

Observation 75929a7f-247d-4c2f-b627-408901531cca · inbound

Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models cites this paper.

Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:13:59.942765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T06:12:38.139005Z digest=sha256:1c0693660506ef4c8e844983a13d3e08154539d812e4efa0d2992f8b25572bc0

Observation 7caea12f-9a5c-48d9-b5a7-51033808566c · inbound

DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration cites this paper.

DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:33:56.709121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T05:31:30.519530Z digest=sha256:8225ab7077033c23427bcca052970b8e4123a5801d766c1152f04211f0a55d0f

Observation 3de1caeb-e433-4ac8-9b79-80f1e83329c2 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 199

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:46:27.764667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:df3b1a0a664f666eb41d875a396b55bb46c6ba5969c311b8045f4f47c5979ef4

Observation c98604fc-563a-44f0-96c6-80e0252eced8 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:46:10.018483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:cfaadd2015937925d7a446b7386936a8277b7233081a9bb9c761dec34d33dc4a