Pith. sign in

Paper Citation Record · LEDGER

Dynabench: Rethinking Benchmarking in NLP

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2104.14337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.14337 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:21:08.222059Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b61e2e05-8d89-44cd-a85b-3293219884a3 · inbound

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games cites this paper.

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games Dynabench: Rethinking Benchmarking in NLP

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:21:08.222059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:21:08.222059Z digest=sha256:b6fae5d39e72ca4b4cdb18409852999d9881b6511964f631d2d66a3820091aa5

Observation f15163c3-9e9c-41e5-a705-858294c34b73 · inbound

"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations cites this paper.

"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations Dynabench: Rethinking Benchmarking in NLP

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T14:14:09.168262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:14:09.168262Z digest=sha256:ae172a057e11f84dd6ca68b5d399bac7ef675b7c48fba37a9b08dfe7fba9f695

Observation 49ea3421-f8b6-4f38-b7d7-95ecb42e611c · inbound

CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++? cites this paper.

CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++? Dynabench: Rethinking Benchmarking in NLP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:17:02.381818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:17:02.381818Z digest=sha256:6b9059c4a83209b59d85e19d55a6e9527895acf56c647df8d55e818ba895aad3

Observation 17846667-5c48-497e-9546-8845f1211c2f · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Dynabench: Rethinking Benchmarking in NLP

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.090971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.090971Z digest=sha256:12c697553166aa94239f9c4a37281a86d6b4ea10df12cac26b3f95d81eb0225f

Observation 10374146-0427-4a94-9227-3d428057b41f · inbound

What makes a good metric? Evaluating automatic metrics for text-to-image consistency cites this paper.

What makes a good metric? Evaluating automatic metrics for text-to-image consistency Dynabench: Rethinking Benchmarking in NLP

Reference 18

Resolution
malformed identifier
no resolver link, observed 2026-08-11T12:38:06.604780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:38:06.604780Z digest=sha256:4a974eff6e99c00b75b716a417a50dbaecf4106a991a8f4b80aefca79d3f73f4

Observation bb36d922-8591-4171-8ea2-9d9ce961197d · inbound

WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI cites this paper.

WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI Dynabench: Rethinking Benchmarking in NLP

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:31:35.674183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:31:35.674183Z digest=sha256:e457f4f5bc2e7e1a10c16aa3eee4655aae3d74450642649765f78e26245685d2

Observation 422a061a-ceeb-4687-91e2-7c693a4ba007 · inbound

AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI cites this paper.

AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI Dynabench: Rethinking Benchmarking in NLP

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:16:43.259108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:16:43.259108Z digest=sha256:e09a99a19c84d64d65cf159507fa1084b9ab8ba2e5e010de77d179c5b78fb937

Observation 3bfebf8d-203d-47db-bea3-2d1bb1b32492 · inbound

Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities cites this paper.

Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities Dynabench: Rethinking Benchmarking in NLP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:20.900119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:39:20.900119Z digest=sha256:e25811fe5bb7bc9245b098b87403d750eb70191f539ef0df67e2492f03c1e929

Observation 39a34301-2c01-4c36-8d7d-52060d5d47d8 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Dynabench: Rethinking Benchmarking in NLP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.375124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:b42ae4dddeea587924fdfdfc861236b02bd660ef8f09ae6f63f8d954dc543100

Observation 4b504f72-8c3c-41cc-9278-205d85bd7d02 · inbound

When Incentives Backfire, Data Stops Being Human cites this paper.

When Incentives Backfire, Data Stops Being Human Dynabench: Rethinking Benchmarking in NLP

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T11:47:48.971268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:47:48.971268Z digest=sha256:3598e9a80baed02b4abee717785921887701807cec9163a2216f96787c4a54af

Observation 3b929376-432b-4282-9211-7a2ad3be4b8f · inbound

Thinking beyond the anthropomorphic paradigm benefits LLM research cites this paper.

Thinking beyond the anthropomorphic paradigm benefits LLM research Dynabench: Rethinking Benchmarking in NLP

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T22:21:52.595165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:21:52.595165Z digest=sha256:9e7b6abf93a3ce0c18c141e662a940b010babcf6f4915c90bdf1fc6a654b112c

Observation dc4387f9-e218-4af0-a674-4f10df6c376e · inbound

LLM Performance for Code Generation on Noisy Tasks cites this paper.

LLM Performance for Code Generation on Noisy Tasks Dynabench: Rethinking Benchmarking in NLP

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.067226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.067226Z digest=sha256:d8f954e321888fb308ccbbd8a95d61087c9a1960f964d918c835519eb8a50ed5

Observation 6c072be9-1d61-4f5e-97e5-5ac8b700aab8 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.708684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.708684Z digest=sha256:16c8845a93416a9687685f0b0d1f199c8b7769f1b38e07362d4732743a669eaf

Observation 011ad500-ec58-4f4c-8ec8-45ea96369d38 · inbound

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models cites this paper.

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models Dynabench: Rethinking Benchmarking in NLP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:15.386951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:15.386951Z digest=sha256:ce1ceeb785afda18968e49af2e259415b4e8a2f637184c67423bdbdf074d0bc5

Observation 2282e2fc-22ea-45e8-86ae-724b13ad9049 · inbound

Agentic Web: Weaving the Next Web with AI Agents cites this paper.

Agentic Web: Weaving the Next Web with AI Agents Dynabench: Rethinking Benchmarking in NLP

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:37.910504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:05:37.910504Z digest=sha256:e83f064a5347dbf024b3e5d7338370c1951bfb59fa94c3561997750b5bd446e4

Observation 6cf568cf-ac71-422f-8d4f-e25c6af4ebed · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.364886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:978be077306fefd5e3fb66a99916356906c5c4f7b6070171819e53ef8f150824

Observation b6ad4775-1b20-42e1-9058-1c72ab173879 · inbound

Private, Verifiable, and Auditable AI Systems cites this paper.

Private, Verifiable, and Auditable AI Systems Dynabench: Rethinking Benchmarking in NLP

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:59.083960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:59.083960Z digest=sha256:565bffdcf8d10355f83f6163e1be9d8aeff19a2b9751156d2ba59f5b1c630ca2

Observation 91451675-098c-4626-8d54-ab661566828f · inbound

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M cites this paper.

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Dynabench: Rethinking Benchmarking in NLP

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:22.922729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:22.922729Z digest=sha256:2dfb61c45f0e5ffa5a13f7e5727b2ded3058031437175990eee9cfae542574dc

Observation 0c5adf30-311c-4f75-b382-796f0467f70f · inbound

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation cites this paper.

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation Dynabench: Rethinking Benchmarking in NLP

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:29.823956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T08:12:02.449352Z digest=sha256:88de5d468f190d92dae7e73c03227f72d37d4f29cab8c7d9777d57b0af171be8

Observation 12d933b8-0880-4d12-86f6-c66322eed828 · inbound

EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning cites this paper.

EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning Dynabench: Rethinking Benchmarking in NLP

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:21:38.373016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:21:38.373016Z digest=sha256:25a7d6b15969f0d0d2ce7e1c2b4615516b7c5906a2694fef7b47d99d393c5a26

Observation 7fa4cbc4-7498-4013-9463-71b8fa7e0024 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Dynabench: Rethinking Benchmarking in NLP

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.268582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:a6bff973ed030f4d12d9b2b23e20b9903d99350b44eeeea689c1f6c185458eac

Observation 142def31-9698-490a-a6c7-a17862a62603 · inbound

RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains cites this paper.

RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains Dynabench: Rethinking Benchmarking in NLP

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.133066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:46:08.897540Z digest=sha256:970647e481d3f06248a258dfb21ae3d5ec84de6ffc0436aa020e7199f1e8e954

Observation 04341ae4-8797-4e6b-9c8b-15d01ebd9d6c · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Dynabench: Rethinking Benchmarking in NLP

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:42.876444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:29:25.830962Z digest=sha256:0222faa21195dd58b78fd4956a64cfaa943cbcdfbd49f52f56b2917b6967598a

Observation 1382ce5a-86b0-4d76-9a13-c56ed3d8f611 · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Dynabench: Rethinking Benchmarking in NLP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T08:26:49.097626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:26:49.097626Z digest=sha256:3a37bd13c8b485b5a5526f702f6332424994647b4e833fddf4e3df4fb7a39099

Observation 00e9179d-a40f-4902-a1e7-dfa5a8904107 · inbound

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction cites this paper.

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction Dynabench: Rethinking Benchmarking in NLP

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.232510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T08:02:50.603020Z digest=sha256:2c09cfb3b70836d8ef50af606378f5ae1a3c138546358c3c20d620cfbf9c5bca

Observation 8cbb151b-15f8-4fd4-999b-69396e71f8ff · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Dynabench: Rethinking Benchmarking in NLP

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.611935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:8f30a508d5c3fd56461ade58b440f86c3034136f2d4f66aa434ed6419e92abe3

Observation b8ad867c-3b0a-4c9f-b96a-8b3ae646d7ba · inbound

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation cites this paper.

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation Dynabench: Rethinking Benchmarking in NLP

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:17:04.499790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T20:13:06.376037Z digest=sha256:a1ef6833b0945a874b0f580418c27016d9c682c579cbf9baece7553f64e51387

Observation 8c63ef9a-691e-4399-b0ed-37d88cdde605 · inbound

Analysis and Explainability of LLMs Via Evolutionary Methods cites this paper.

Analysis and Explainability of LLMs Via Evolutionary Methods Dynabench: Rethinking Benchmarking in NLP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:11:05.214722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T20:37:40.932811Z digest=sha256:89a6d83ed15aec7449f5b5b9c580baa3fbfb3f3c370b9d1f24d11b2babba5458

Observation 10663faa-a7ed-4d7c-b0e8-5359b4e29507 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Dynabench: Rethinking Benchmarking in NLP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:17.250381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:52fb0ab19ebb65305711bdee80d4782d6b603e0cb00ef477d95013394521873b

Observation 3f039523-3804-412f-a2bc-39b04018becd · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Dynabench: Rethinking Benchmarking in NLP

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.769739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:b92edfe074c8cb7a6144bfcf8c5d3b90d9cb3e9a6bea1025e115bca8e697b261

Observation f0e8d7ea-342f-4d28-bfcd-1ae36ec07607 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Dynabench: Rethinking Benchmarking in NLP

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.040649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:dfd1e6e6b2829c4713945e3ed6266e313243a893eb7cd8389b26e38f20fd1430

Observation e6026c7b-c1ff-4fa7-b33c-fd7ec7ec5ce6 · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities Dynabench: Rethinking Benchmarking in NLP

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:43.734758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:eb5882126454d3fec98076403769994b65806d63cc76d284a2d8cad33e44c448

Observation 9137d3af-0e49-425f-865c-49b6bf62c79c · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Dynabench: Rethinking Benchmarking in NLP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.468788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:2944fc3195aa5391ae03b92036052387d1bf9bc378a9d41a18c600eecfdb04cd

Observation 3610d9eb-e494-4f55-87ed-ae935d5d4b59 · inbound

Meta-Benchmarks for Financial-Services LLM Evaluation cites this paper.

Meta-Benchmarks for Financial-Services LLM Evaluation Dynabench: Rethinking Benchmarking in NLP

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.686773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T14:08:20.432931Z digest=sha256:b2c6d53ed70bec61b585feb3344ba277ab95a3859202e52867d7f31e6b0f7b78

Observation d18b0e1a-4d83-4284-a968-2162598b044d · inbound

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation cites this paper.

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation Dynabench: Rethinking Benchmarking in NLP

Reference 306

Resolution
unresolved
no resolver link, observed 2026-08-03T00:24:07.971141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:24:07.971141Z digest=sha256:49ccb218c7449623c99b70598a294e02b65652bc9825f0fdb2c9f0710d591b27

Observation b805c310-4960-4f48-ab76-5432e72d93d6 · inbound

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks cites this paper.

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Dynabench: Rethinking Benchmarking in NLP

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:33.194776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:45:33.194776Z digest=sha256:1096ca0393ba1256c9651bc828458a25e8384226987e891c3a34d3f53d6bf2b0