Pith. sign in

Paper Citation Record · LEDGER

WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2406.04770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04770 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:21:48.754749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:56:41.068880Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 333df622-23a6-4a25-b5b4-a4682dc55436 · inbound

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback cites this paper.

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:38:32.538349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T22:37:43.230753Z digest=sha256:1e06e1928262e82e2b2a80546dc3700d9d963e28f7fad7406e135febcdd61f6d

Observation a03e328e-ab9a-4da2-823a-1dddb6be934a · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.233353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:a7cf90bfa24c633fb9304fb056e5598c7ce3f9eafff4780c8b74ac6e264ca5ce

Observation 2c261279-5211-4122-8ebf-592572d9f32b · inbound

SLR: Automated Synthesis for Scalable Logical Reasoning cites this paper.

SLR: Automated Synthesis for Scalable Logical Reasoning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:52:13.798019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T08:47:17.008113Z digest=sha256:90bbe67e33d4d0d60d3d448cb00fd2a6a47d95518867ec3bdc8d323e1ccefa40

Observation cd8a3431-2352-4886-9242-4b74c792ef2d · inbound

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries cites this paper.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.754749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.754749Z digest=sha256:9fba3d95bd60d27a0db355e68c92b6275c980c603254681f3626762dfc132f00

Observation 430098f2-be50-4b18-a573-9a7fa46550ed · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:40.016367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:de9cfb9eea790e7f8fd728973192279cdeee318a683d75f003082f820e5aca04

Observation 9365063a-f02d-4276-8ad2-93eabb46581e · inbound

A global log for medical AI cites this paper.

A global log for medical AI WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-04T11:34:16.004834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:34:16.004834Z digest=sha256:26d77c38e7576162cb0d369c7b353ec2487d72135986801b3446d95089679cd9

Observation fca7aa40-b80c-4670-830c-3d39af3f28e3 · inbound

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training cites this paper.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.139966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.139966Z digest=sha256:59f672b80a12a0feca76e8b0ce5b95fdcdf8f800653a5c3d0dd55176c91307ec

Observation 966c419a-95d0-47d3-94f5-173d34f5809f · inbound

Ministral 3 cites this paper.

Ministral 3 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:24.743284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:bbccd599f152719ca6b73b60ea9a61b6d7d921aa2c4c68237d37eece17ee4019

Observation 2e57e716-744a-4c35-a42c-965a345511b7 · inbound

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE cites this paper.

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.582310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:34:48.524592Z digest=sha256:97aa543846237a36de5fbecddefd1315680bcddf476691e87045de28f932daf8

Observation 486f3654-1f72-4add-af90-0b07c8808378 · inbound

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks cites this paper.

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:09:26.247244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:09:26.247244Z digest=sha256:ddc722de56324d93c71e1f493e6d29b22f19bec4d4bfc590fd9b1e60b803a683

Observation 5d1de721-441c-48f9-97e7-bb3e31d73ef4 · inbound

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks cites this paper.

SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:09:26.124170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:09:26.124170Z digest=sha256:29eebff76d576bf6be124a3bb5d37faccc38cfad221f8e373f6afa1b7654c882

Observation a3f08183-652c-43b0-ab00-651e72041cec · inbound

SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology cites this paper.

SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:03.627283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T22:20:42.183878Z digest=sha256:54c22d0119d24c11aae8ca652b583c48012514b28703962c1fd1c201394bc01f

Observation 9d4c6217-117d-4b38-a557-c4b0b1dbf71d · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.492065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:a4cd560329ca602d3f536b6c96704d8c1d64fcd57eea1aaa7c8c6877588876f0

Observation 5c809790-9671-44a8-8359-bf613eaf5d8f · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.432902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.432902Z digest=sha256:5b91b82cfa46f8ac077f20dae710a56f4f847c097006bd844fd8ea3f2a496e39

Observation 4ddca33e-c1f3-4791-a5b2-43df81a94cae · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.341882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.341882Z digest=sha256:9c2a2fa08c0321138ae70d9cbd6f721751fb278d882f50f386f5518cff1332d7

Observation e6301820-d308-4e8e-a821-8f120c81072c · inbound

SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility cites this paper.

SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:15:58.269538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:42:57.745612Z digest=sha256:6383ec71bda44bea6985f0a84dfba6aea8ea33aaf01e668754008f371a0c89e8

Observation 2d604029-404f-459e-b770-e98a57a70e77 · inbound

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems cites this paper.

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:20:10.373403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:18:09.127345Z digest=sha256:8bdfa7855ea95e1025d2c386673d50e27337920371e4a40bbab83253c084c889

Observation a7e803ed-fad8-4877-b38d-803aab63ef84 · inbound

Submodular Benchmark Selection cites this paper.

Submodular Benchmark Selection WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.309294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:23:51.649257Z digest=sha256:dd5003b7cfb06e5f17c9fce9accac81e02eea8eff853b503281bc96b96765f0c

Observation d3401c76-3a18-4044-9471-62c4fc37d666 · inbound

Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone cites this paper.

Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:40:43.319500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:11:42.971428Z digest=sha256:41c16abb337d69e75073c7a9bc82dcfdff628b3694486f840323621759cbbfbb

Observation 6f66d758-382c-45a7-9032-dbd91a7ef947 · inbound

ProactBench: Beyond What The User Asked For cites this paper.

ProactBench: Beyond What The User Asked For WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:16:16.058678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:14:01.145443Z digest=sha256:431a6670397ec5d5689e2f2e3789e7621f7de38f2d20fc5119402453f73c75e8

Observation d068d95e-3fba-4153-979b-5947cc41ca9e · inbound

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants cites this paper.

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:48.587483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:28:13.317630Z digest=sha256:cd131651d63799a803ab4c17d430cc010aa943a5882c905dc58a5c71e11bf0d2

Observation 0d5e2a77-4f95-4d5d-bfb0-53968831b645 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.744570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:a4f88445734860a6ff858d87c8f10542f1d1e811d409914f699930f78e5c4cbc

Observation c7de08d9-9862-4ecd-9fca-5c26f3d31a63 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.863252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:859a29a0ed2e21580d362b6a15e0f02bad9076d793244a2da687443afec8252d

Observation 22e02d68-cf31-4d4a-870d-0d1869249cd8 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.688249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:361e6f2a22028e07923ff5626f7cffd9e89330c002f54a3e771cf77631fcec3e

Observation f3e2b446-5cc0-4b78-ba83-21ff1527939a · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:43.678968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:d545a03f6f5a313742746193c64670d0afdde7111a459ca36b7fc9c38db395be

Observation 383d3279-6317-48ae-88c4-f430615758c6 · inbound

LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control cites this paper.

LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.894432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T05:02:31.689194Z digest=sha256:4c48bf89b88dd046db2f58a78d391faadd8f684556341faa846cd62cb7c4825d

Observation d634685b-3815-4c48-8c0d-7db91b7bd372 · inbound

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation cites this paper.

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:56:41.070852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:51:14.922841Z digest=sha256:29d27c62e6bbe5092339cc64a0130da2adab570d868991e8cdab91188d5a4965

Observation d1bae9ca-a40d-4ca4-8d85-52e438fb6a62 · inbound

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning cites this paper.

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:31:54.729032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:31:54.729032Z digest=sha256:5d08e804cc7230d5a0819b08fc4b28f99347d61b7a8d381f4cda8dab6678083b

Observation 45081487-b398-4748-b9b5-2dd1067f557e · inbound

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning cites this paper.

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:50.931706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:49:50.931706Z digest=sha256:85285b32652873cfe2cfa55c3331858a91118ae8bb109f2f2fe0f82946e9d26f

Observation 79d822aa-4ddb-4df3-bd1e-ec4c29589558 · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:24.341223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:24.341223Z digest=sha256:af3e423ae1d82b9969070f34ce38e4487cc14bc859e0fd4f3acdb2e142b3dd60

Observation b4c8ce04-d87e-442f-8b76-86ac69f15ed4 · inbound

Response drift across frontier large language models cites this paper.

Response drift across frontier large language models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:46.097954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:46.097954Z digest=sha256:7426863f75940fb88bff41c72cb2e098d89ef757c7364f5210dc1af23e20e976