Pith. sign in

Paper Citation Record · LEDGER

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

As of 5 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 59 inbound Pith citation observations for arXiv:2406.11939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11939 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:54:06.577486Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 95df49c6-7ed0-4fd1-8578-c9df2887dc29 · outbound

This paper cites InternLM2 Technical Report.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline InternLM2 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:03:38.978285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c031c033bc42fc208a4c00e126c45f7cfbb9f48e512e043fe026c331de9d1c4e

Observation ca8c94d0-bb92-4d0e-ba17-fd0a9cad37e2 · outbound

This paper cites Holistic Evaluation of Language Models.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Holistic Evaluation of Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T00:03:38.982197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c11950bbd543aa010c259639f3c4a646eabb161dc6e80fc0b75600744391235f

Observation bbb83d99-d44f-4eaa-b5e9-aa62b71fadc0 · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.137116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:12b17f63dda1f6426b21da30bc4716640ee793f30fe153c2b3c851302af883c0

Observation 91627f94-15f5-4c9a-8d42-17858dd5fa38 · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.097331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:4e193fa555c858f05c7d7cd0854a223e4a65872f59277ecce6c940a18088d275

Observation b811a5f6-6b23-40ed-bde8-474e447f6844 · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.143797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:4f58918ae4724f5cce7d8ad7539e21fe9262c7bfa53344039594cd57efcfcf76

Observation a2150172-ad39-4717-bbd2-7c697f03238d · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.100273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:13133cbf13ad2f9670b399ce7c06728555b89802bf8d85f0d0d7b15713d1be82

Observation f5fa58e5-0868-4f1c-b0a7-2891005932ee · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.140443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:a2bc7b30a6045dd49274ab6c79183c8c44cf878bf0ea1e036db67612f194de1f

Observation cca12022-6220-4a7f-b922-115b0fd342ab · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.147138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:a6d7580f7d0f2b759aebf5977c26b1dea1cba34539724d7b620bc1bf56a65056

Observation 5de70c9c-91f2-4cb6-a749-d7955c607bc8 · outbound

This paper cites Criteria Satisfied: [1, 2, 4, 6, 7].

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Criteria Satisfied: [1, 2, 4, 6, 7]

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:58:42.133553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:7681f07556fbf8e289928434aff285535ac76f166cf37bec80496ff1319c1323

Observation 187cc44f-b6e0-4718-842a-82fff15b785e · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.155492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:dcf1fd5066f788a3d8e31de4597b75ae39def0e1172205de6742aac860de482c

Observation 361fe78c-86c9-4022-b54b-6ba9c95cd031 · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.151249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:6ef2326e94ce6ca3c3e2548ed13bdda28da28c495fd0cb638284f86401581f60

Observation b7a2aa7f-e2c0-4835-a245-390962c2382b · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.129844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:2d7b06a47fd1a2ff72d8cf5ca99f19db0bc81c9fa4e3b36acf8bbbb5c2c07196

Observation a3412354-876c-43bf-9da8-efd5f3ae23c5 · outbound

This paper cites an unresolved cited work.

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:58:42.120990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:7bf68bd8010a05d28fb4a0c6cd32817e99f3f6f3650d47b82f459017ebb26582

Observation a97f0125-6542-40cb-b5d6-0747d397d2e0 · outbound

This paper cites My final verdict is tie: [[A=B]].

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline My final verdict is tie: [[A=B]]

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:58:42.125529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T23:59:46.467362Z digest=sha256:4af6d0610bc3b8d3fdbc1363bb5248325aec01d3fad0b2089a56702dc429b410

Pith citing papers

Observation 150b4d26-5c42-4ba8-94fb-d21ca507c1d3 · inbound

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback cites this paper.

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T22:38:32.638248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T22:37:43.230753Z digest=sha256:afd22b198f39bec002d1028dce8b456a35e84f6c681dbc7f4cde6117291b1549

Observation 3a66b36c-cc5e-4138-9e7e-22e305edfa49 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T06:25:27.947989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:6b3070a712784841bc95c91575277c9d90423dd75e4e4e5e60863c1dadf1bdc8

Observation bfffb73e-a91b-4d2c-9bf9-2f37f60fd8bc · inbound

Qwen2.5-1M Technical Report cites this paper.

Qwen2.5-1M Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:26:00.110871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:25:59.964619Z digest=sha256:3021e1c24cfea7e30ead3d976ceecf83b3d714cc219fee302df6cdd31efbfe74

Observation 915dee21-dd03-4714-be41-7f53fe0437df · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 133

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.052021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:fc0389c7a1681b0b9131963810c968a6718b48aac576136efa7c91b551a445fb

Observation 55aeaf3a-e86c-4f12-a57e-ba7685092167 · inbound

Sustainability via LLM Right-sizing cites this paper.

Sustainability via LLM Right-sizing From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T19:25:03.708397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T19:24:58.690462Z digest=sha256:922dea2984cab66ea35d2b92391e6b84cc73445b005ee92703782ebdbbdf9945

Observation 09202a5b-c0cc-48c2-9a0a-a4cd1f7c70c5 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:40:25.802248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:faeac33fbb15fe5c2727eb3b7764557a3df6c4fc632025935c2061a57a032845

Observation eb5e6986-17d3-4e71-9ca3-c7408c504eac · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:35:28.553855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:6af0505d82dad3efb984a05f378cd96e1ae21c47f704349f6a1134e3290984a5

Observation 8b5ed6e6-6b0c-464c-b770-dd1aafc19f5c · inbound

LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning cites this paper.

LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:09.341813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:51:35.381534Z digest=sha256:33b414b6e55e629f029d67262bd95142561f975ea68e5d14d2e467fa0b757564

Observation 3d33e227-dec6-4543-86e2-fbec5e4d4459 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:49:28.169833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:cb5f68cbb4be484c60ee66515cf3aa41f2ab139cce7dce408c3d80c1d7983ea9

Observation cb8280bf-0f5e-4da0-8588-7113be70383b · inbound

Proximal Supervised Fine-Tuning cites this paper.

Proximal Supervised Fine-Tuning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:42:50.861827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T20:42:14.423836Z digest=sha256:c0b8a6084abdd82cd1671d2e174a0e4b8e2c2aa46e1d8e1cffe0923f4ceaeff1

Observation d6fb6c37-593f-41ed-9d42-74e73f74a884 · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:16.737789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:16.737789Z digest=sha256:5653d2e9f8424ff48dbe6be596b81d7b5fbd0216fb7a5b84724e60f9877bd64e

Observation 03a7378c-57aa-44c9-b17b-fe7aa3bd19e9 · inbound

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? cites this paper.

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:51.162616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:51.162616Z digest=sha256:9e372e974678103bb376845301ad086166233e793d7a208e66682baa266fecb1

Observation 3d12ec89-33da-4276-9312-b039b62a3923 · inbound

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention cites this paper.

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:54:06.577486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:54:06.577486Z digest=sha256:752e0a3e93981991ebb1764106227e925ed7e949c89525c3bf0379a91b468ddc

Observation d8a6185a-fc02-4d71-8a46-355ab7680a42 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.334341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.334341Z digest=sha256:f0bb4c09f9cf8758cf00d6b1ac2552f4d4327e653ba68f9b43c271ee6f806549

Observation 54fdf2cd-095e-4648-8f58-8eb2e51a44ed · inbound

Contrastive Weak-to-strong Generalization cites this paper.

Contrastive Weak-to-strong Generalization From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:56:37.225419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:56:37.225419Z digest=sha256:27cf4707601ff0e0244d4effe38f94fe142f8f281ac3b207447012eb5b6c33fc

Observation bd9fb3b9-dc58-47a8-bd0d-15bc26f08b35 · inbound

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization cites this paper.

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:28:12.935875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:28:12.935875Z digest=sha256:4aefa86385970326f1a1650644ad49bb46ab1806af9851c08b90038a27032118

Observation dfe06954-e8e9-47ef-bd0a-e0287188fce7 · inbound

Evaluating Large Language Models in Scientific Discovery cites this paper.

Evaluating Large Language Models in Scientific Discovery From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.399906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T21:47:09.588941Z digest=sha256:9f6b990d199634f5026340e25796ade9c71b587bad27ee87f8418f4a548349af

Observation 703538fa-5ddd-4fc9-8533-c7e19f6254b9 · inbound

Ministral 3 cites this paper.

Ministral 3 From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:24.737310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:4609b72fa5106eaf0df589ecdf2b39b2ee079cf8620753c930a66532deb448ca

Observation 5581cca9-6475-4b63-9926-4bb09db415d1 · inbound

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill cites this paper.

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:31:02.345859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T17:36:10.725278Z digest=sha256:23d7b7fa48cdb2e01ba8b15cfd293945a49641e4706df6541b2b3dfb32d7cae0

Observation 089bb7d7-f02f-4bdb-b5cc-f0e1f45ecb87 · inbound

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering cites this paper.

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:32:00.320382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:29:03.994957Z digest=sha256:b9ffbcf9ccf23ddce8575e4de4ccdd561a493ab43bc54fd3442f46fd8edd5ea0

Observation fa042987-fbf5-4573-a1b9-ed7d3f261a59 · inbound

Hybrid Policy Distillation for LLMs cites this paper.

Hybrid Policy Distillation for LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:49.037705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T00:41:21.984760Z digest=sha256:372c29717834f080d16940c6aa22b67e43978ef97799e3ef9639b875691d002d

Observation 9c0fd436-e9ab-4a91-8650-a5d6a8307059 · inbound

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity cites this paper.

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.395945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T11:45:54.181019Z digest=sha256:efd105ee6f1d6c995f63592c55b50c44e644358dc5a396307b1e630149b1bb51

Observation 8871aed4-4f44-4ffc-a9fc-5b617c51f5e6 · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:23.131233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:e6341d9be9b27f5914f2056611659fad494a1d3d61f2c1603d5a1d5a92bc2ceb

Observation c1b00215-5dcc-4705-b4fb-ade79a9b547b · inbound

Submodular Benchmark Selection cites this paper.

Submodular Benchmark Selection From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.262305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:23:51.649257Z digest=sha256:be47816bfc3517ad4411eee84988bf20cec682dbaa0c1312454e1d9e20067fb0

Observation bb0e6e62-fdbf-4a77-af71-1215e91bf39c · inbound

Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone cites this paper.

Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:43.241242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:11:42.971428Z digest=sha256:de00f4656bda13df829b4b03f2e72954709fc6f4254a7da08d2cccb1135b97d3

Observation 47b82246-5b1e-4394-989d-8d235c1ed755 · inbound

From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMs cites this paper.

From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:55.889914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:35:11.605353Z digest=sha256:0321b9529fc9e6d664f38f584cf825d0fb7d9514970ee2c01bd6d546463173fe

Observation 43bb4391-584d-4cbf-91b5-6565d03035f7 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.402757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:b2ad939a6d83d8618a6e8dda526033a2b72d72c1431d597e885a8d049f8b0322

Observation 48717c06-398c-4f3e-8631-87c7c59b50fd · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.930353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:540bff1d325c5e1a1615fc40f8df9f79569e1487e377126212f5d2fa7409f946

Observation f3bd7a5a-3a1a-49fb-88ae-b90854a4275b · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:39.616715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:28b9dcacc474e36ce46775a6a3784ab3ac24f833de8343179139775037c0e4e3

Observation 182e70c3-1210-4a7d-9b83-492096619b90 · inbound

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models cites this paper.

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:32:36.155297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T14:32:01.939211Z digest=sha256:8faf9f20a353d0835febe77bda550210c58d616373923481810a94c766e16387

Observation 23c182b7-74d4-464e-9b92-f868d5b26276 · inbound

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models cites this paper.

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:23.428842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T05:44:59.785603Z digest=sha256:153ec4ff2b82fa1aba49a65abf78dfaf2ec107ff84c10fcc4d04839aeb0ff555

Observation 7401ba63-e9cc-4b6a-8633-98ca2216df68 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:38:12.088921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:93b61f7be9ab2f381614cffaa34da5249a0c649aa9ad0d5e34275d4e0f5176ad

Observation e10e47c4-ad36-42df-954b-a0861f3837fc · inbound

Evaluating Multi-turn Human-AI Interaction cites this paper.

Evaluating Multi-turn Human-AI Interaction From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T08:33:24.478464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T08:33:08.019966Z digest=sha256:f482fe5c7506701d65e3f1f3a6d85b980cd61922c1d94d3ecec5898af717b72a

Observation df75fa0a-181a-4f21-947d-94df2b9cf2a6 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:48:17.732214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:49a9d64b9bdd4a687628534ca73d62600a7916c44547e5130afbd4a0d727d6fd

Observation d539ea44-739e-4de8-bb4c-be49c2195121 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:54:02.908526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:a206b3515a5dd4234568de810d845193e669b8adfdcece3f8c000e90349542fa

Observation 591565b8-f68b-4c8d-b1bc-d8cd14fa2332 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:24:45.730861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:203de8b4daabe439c1b4159d4a345f1a1fe8d34c0b591a108d950834fcb49307

Observation 5297ffe4-e718-43af-9689-c1f2154f1c20 · inbound

Dynamic Model Merging Made Slim cites this paper.

Dynamic Model Merging Made Slim From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T15:18:25.364526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T15:16:24.651868Z digest=sha256:991b8b5df5d4e2b1a6b4a31455bf777f5c936715b726ae036380fbb66c9fde7a

Observation 319eb073-2cef-4b97-b614-42ae03fdd845 · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:18:11.754889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:b86c2dbb6de02d09fa3261b69cdda836e6399ecad3fd9b778066d94bdeb2dc98

Observation 23c382d4-2f54-499e-aa10-ad8029cfa605 · inbound

LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control cites this paper.

LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:03:57.885621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T05:02:31.689194Z digest=sha256:cab171d8ee67fea47737e3bdca54cadcd3f7ec45a4bccc9069e7afe30420585a

Observation fac8b699-6e06-4fa2-a028-7ea85e13cde0 · inbound

Convex Optimization for Alignment and Preference Learning on a Single GPU cites this paper.

Convex Optimization for Alignment and Preference Learning on a Single GPU From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 108

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T05:06:38.363145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-25T05:01:31.560963Z digest=sha256:616eda104274a7cf56e38e4829ea5cd81fbe56d47bb5bc6efb1bafe2f5103ee7

Observation 0b55b1ef-e633-4ecb-9a32-28ec230b4a15 · inbound

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains cites this paper.

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:23:28.534671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:13:57.012665Z digest=sha256:bbb81c88265120324746e491bbbe2228312f8aae36409bf0dcda8d17f83655bb

Observation 3e6b3b05-f573-4c49-aca5-e052215d7380 · inbound

Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Level QA cites this paper.

Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Level QA From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:03:12.895998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:00:07.102425Z digest=sha256:f80754c7ed0cf1caac79f89c496af5ece09865ffd4b1d29f6dbd83ddd2ea7bd2

Observation 3ced6e07-2f39-4cba-8b02-21f44c4b0f33 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:47:38.285243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:dce39712e3e5612239acca9a68caed05b433fc62f526062c712b6165b9f15db9

Observation 53eb0049-3436-400a-923a-e60e5496545c · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 140

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:57:41.357282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:46883c6c1bf46058331d8e17767aba23add64798c80058c498b3ff68fd758dc7

Observation 6e7a757e-2aea-482a-a65c-7626f779eeab · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 202

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:17:25.498883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:3405bdbd0e711cee523fced03987c6dc07c16d68a5c3e30e889a57184c596764

Observation 49ffdf7e-6211-4214-b304-609968163117 · inbound

MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision cites this paper.

MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:48:46.040647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:43:37.671325Z digest=sha256:d6df456deb59fb6af6f207db5b8c56fa66d15e8fe879a47e8f1f375cb57ab3ad

Observation a84e85a0-e0ce-4ecf-8883-28cce51764c2 · inbound

MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors cites this paper.

MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:08:56.728217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:33:18.750759Z digest=sha256:a24f99786fd4a47cfc984dd72bf3cfcfa6aa9ab4b06de3fade62fd681e1d304d

Observation 55cfe3bd-196c-4919-abf4-5e08550773f7 · inbound

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation cites this paper.

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:49:58.280215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T01:14:26.356160Z digest=sha256:47cb924546a7b66ee2289de1d698f68d5845ac1d71cf253ac88ce5dd91697afb

Observation 056f5a06-e39e-4dbb-9fbc-67ddff6aa74e · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 206

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.579558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:300b8149ae383d3892b2fe3b063e2d3eee31beee3000723b1ee7ac5bd4d5f89e

Observation 259eff52-fcb8-41da-b6b4-1af6fef2230a · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 206

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:fa9796fec9f496d2493dff8c042395a792cbbb05db2ec74fd0ffbabf7ac5c1f0

Observation 97ec935e-ccad-4508-8472-c7839b1c644b · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T12:53:50.140692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:590d51bfcc00c1d141c8479cf040688cda7aa321b1058c4198657197001b5a39

Observation 2114fef4-2e7f-4078-bc10-19361bce13f5 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:3f281988955e0e7c230843efcca6338bd02075679e0928df70e57155f3edf69c

Observation 9c0c2c1a-3e45-4cab-a8b8-87cc5ffaeeca · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:877d7b1546550d286cb88c1e513481dc22485d47e0a7d02520ee3f5bbfbf5c0c

Observation 67cd5d86-b2b8-4b58-bc41-af628a67289c · inbound

Scaling Point-in-Time Language Models cites this paper.

Scaling Point-in-Time Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T15:39:38.560563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:39:38.560563Z digest=sha256:b95a47625e0859dcb84eeb8890860d3036d876eac1e1d0786a2bb86c8a3777e8

Observation aa2f134a-7b0c-48a9-8009-f1286765e82f · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.224184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.224184Z digest=sha256:b52db6548844d1ac2a3c08bf5f47b70672afe3b0029ef6169a84c5cd418b428d

Observation 1647cdb6-48cf-411c-936e-adf1b43b68e1 · inbound

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction cites this paper.

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T13:53:46.138863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:53:46.138863Z digest=sha256:11942ad1463875bd128e853e374e8b485267f74412ca0a674a042d93e75f525a

Observation c2e4800d-3335-4c2f-a66d-d3be96e90501 · inbound

Economic Evaluations of Language Models cites this paper.

Economic Evaluations of Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:06.050155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:06.050155Z digest=sha256:8ba7fe0d23b1bc2509dddc4b678886a724b3b4a5748449e79cda8d55154f0050

Observation ff7d2a22-4f31-4eb7-b487-4ca02e191efc · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:53.993819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:53.993819Z digest=sha256:8a7905c9fd19c3f8e19582dafa4a8cd13c8940e808fcb41bba37cd7e6c238a8e

Observation badf8a55-2f7c-42ca-ad32-774b65927750 · inbound

ExplainBench: Evaluating Code Explanations from Agents cites this paper.

ExplainBench: Evaluating Code Explanations from Agents From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:36:22.838882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:36:22.838882Z digest=sha256:65f4ea3fb06cbf9e7ffb9eb4510028b8c30b26721c21d4de4514bf22e6e893d7