Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2406.04770.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:21:48.754749Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T01:56:41.068880Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 333df622-23a6-4a25-b5b4-a4682dc55436 · inbound
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a03e328e-ab9a-4da2-823a-1dddb6be934a · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c261279-5211-4122-8ebf-592572d9f32b · inbound
SLR: Automated Synthesis for Scalable Logical Reasoning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd8a3431-2352-4886-9242-4b74c792ef2d · inbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 430098f2-be50-4b18-a573-9a7fa46550ed · inbound
Evalet: Evaluating Large Language Models through Functional Fragmentation WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9365063a-f02d-4276-8ad2-93eabb46581e · inbound
A global log for medical AI WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 127
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca7aa40-b80c-4670-830c-3d39af3f28e3 · inbound
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966c419a-95d0-47d3-94f5-173d34f5809f · inbound
Ministral 3 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e57e716-744a-4c35-a42c-965a345511b7 · inbound
EvoESAP: Non-Uniform Expert Pruning for Sparse MoE WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 486f3654-1f72-4add-af90-0b07c8808378 · inbound
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1de721-441c-48f9-97e7-bb3e31d73ef4 · inbound
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f08183-652c-43b0-ab00-651e72041cec · inbound
SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d4c6217-117d-4b38-a557-c4b0b1dbf71d · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c809790-9671-44a8-8359-bf613eaf5d8f · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddca33e-c1f3-4791-a5b2-43df81a94cae · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6301820-d308-4e8e-a821-8f120c81072c · inbound
SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d604029-404f-459e-b770-e98a57a70e77 · inbound
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7e803ed-fad8-4877-b38d-803aab63ef84 · inbound
Submodular Benchmark Selection WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3401c76-3a18-4044-9471-62c4fc37d666 · inbound
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f66d758-382c-45a7-9032-dbd91a7ef947 · inbound
ProactBench: Beyond What The User Asked For WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d068d95e-3fba-4153-979b-5947cc41ca9e · inbound
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d5e2a77-4f95-4d5d-bfb0-53968831b645 · inbound
General Preference Reinforcement Learning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7de08d9-9862-4ecd-9fca-5c26f3d31a63 · inbound
General Preference Reinforcement Learning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22e02d68-cf31-4d4a-870d-0d1869249cd8 · inbound
General Preference Reinforcement Learning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3e2b446-5cc0-4b78-ba83-21ff1527939a · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 383d3279-6317-48ae-88c4-f430615758c6 · inbound
LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d634685b-3815-4c48-8c0d-7db91b7bd372 · inbound
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1bae9ca-a40d-4ca4-8d85-52e438fb6a62 · inbound
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45081487-b398-4748-b9b5-2dd1067f557e · inbound
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d822aa-4ddb-4df3-bd1e-ec4c29589558 · inbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c8ce04-d87e-442f-8b76-86ac69f15ed4 · inbound
Response drift across frontier large language models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.