Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 91 inbound Pith citation observations for arXiv:2403.04132.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T04:18:53.743378Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
51 of 51 outbound references displayed
External citation measurements
322
pith, observed 2026-08-05T02:28:24.338817Z
Observation 3b852332-53e7-4b07-9e7f-5afdd21562e3 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Training Verifiers to Solve Math Word Problems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d142e643-0925-47b4-9744-8455131e05f0 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Statistical behavior and consistency of classification methods based on convex risk minimization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 787e3bc8-c303-4384-aae9-92022ddcfd50 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference emnlp-main.608
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a3589a1-94a7-4f14-a295-2b476f74e970 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 35aee26c-adcd-40c8-b846-637c71cf6786 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b8f8976f-bf49-416e-aa11-65afed9d7222 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Travel Itinerary Planning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a3a773f7-2951-4bdb-b665-571133c1c05e · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference It houses collections of European paintings, a medieval and Renaissance collection, ceramics, French sculptures and more
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff5e5e2e-da36-4b7d-9be5-474fc07dabaf · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04aaa260-9393-47be-8f71-d584be635ecd · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7c6ceb9f-17b8-4e34-bc67-3cb6bddda03f · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2956a5a4-2fa4-4f63-945f-3f35be0c6caa · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b60a4b1-ed88-4dd8-8c66-3867ec5d765c · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88f3a4e1-4bf7-4b77-bf07-7d0a3c52fa44 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9c0c2fa-ce7e-4ad9-a487-74682a1c16f9 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 59899fac-74b0-4cb2-952c-7f4fa5da14cb · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ed6558b-229d-4c0d-bde5-41cceede1b47 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dce4dfaa-a604-42bc-bd6a-068216c35d74 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2d465c1f-d45a-4de6-bc5a-c1c8ea01fcfc · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f2c0d46-da3a-4aaf-ac5f-1d054d136ad5 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation abc05d99-10c8-4086-acfd-e241ee92cd02 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5e625f3-9fc7-447b-97cc-a324ac44f331 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Remember to check the opening times and any COVID-19 restrictions before you visit
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef59e6a6-e601-4c3a-88d8-a99d3c1029c3 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a11181a5-6c78-4a8a-8e76-b6af720e4c12 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7e98b0d-fc78-495b-92f9-61310a62788e · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8403948-795c-4bec-a05a-a42863595158 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19aa2493-1f99-4adb-8086-7d40831e636f · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd4c204e-295a-4d04-a2b9-db2513e28cf5 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96c83534-083a-46b0-a5f4-c2bbba8428bb · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97c5cc0b-79c7-405c-9704-04de5d1b26d2 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75062dbc-d0a3-4fc1-9764-fc9711c1425f · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0ee6e5ad-c4cc-4d44-a8c6-2b093683ecde · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2b302ad6-b485-4af5-ac58-48afc5f9b1d1 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98978a21-05da-453c-bb0f-3ca2679f8132 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f7cd57b-7194-449b-b645-da23a8cee567 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 48c6805a-20b5-47f7-b74c-b5b73aac8339 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 795c00ba-9007-40af-a9c1-191090fe20ad · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference These are just a few ideas to get you started
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 068acaa6-f5a5-4991-b206-0b2471cd8873 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a6f4ef4-59b7-46ed-9c85-7658b832cefe · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f170670c-4834-46c4-a5b3-9e3fda7ec112 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1d8eb5cf-b8d5-4c00-a1cc-59afb239712c · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28a77312-fa14-4b52-bc9f-4afbf4aefaf6 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference My final verdict is tie: [[A=B]]
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3691fdaf-2448-44c4-b6e1-5f885dce0214 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference 27 Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Continuously gather and incorporate customer feedback into the product development process
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7de7b00-0f13-463c-8e3d-9ef264470775 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Align your product’s features and capabilities with its value proposition to ensure it meets the expectations of your target audience
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fab17ebb-364f-4e64-bc43-81c393a82421 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Validate assumptions and hypotheses through experimentation and user testing
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 691382cd-cec5-41b0-995d-813bb4a3d71c · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Be open to pivoting or making significant changes based on feedback and market response
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 736c5688-4e45-44e9-b9f2-d1741812e124 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Establish key performance indicators (KPIs) to measure the success of the product and track progress over time
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a1938a07-32ef-46e6-8c6d-115900e32fd5 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be obsessed with their customers and be willing to put in the effort to understand their needs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff7b0c08-f83e-4cde-944c-007c14c6514d · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be willing to try new things, test hypotheses, and iterate on their product based on customer feedback
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb6c5f2a-218f-4afc-b89c-3f40e719a856 · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be able to identify and prioritize the most important features and functionality that deliver the most value to their customers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a053b572-8a53-4e6f-9303-f719c5f490bd · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be able to work effectively with these teams to develop a product that meets customer needs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf8ed48f-61ae-466a-acff-af8de76fc54e · outbound
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference This includes analyzing customer feedback, usage data, and other metrics to inform product development
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f6666ca9-1e4f-4bdb-b702-6f659369aceb · inbound
SGLang: Efficient Execution of Structured Language Model Programs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be987ecb-0262-42e5-baaa-7a61352de344 · inbound
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a07d66b-ea87-4efa-9588-0d8643c0835f · inbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d53d4858-397a-42ed-a2d5-1225a7f1286b · inbound
Benchmark Data Contamination of Large Language Models: A Survey Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1c431ca0-cfef-482f-85f4-a53d9773a37c · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f619c2b6-d3c7-47a0-9f86-70eb839dc106 · inbound
LiveBench: A Challenging, Contamination-Limited LLM Benchmark Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94c915b5-1495-4b1f-84e9-3f91c581087c · inbound
Qwen2.5-Coder Technical Report Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 200f367c-3777-429e-bee8-0ac5ee91faeb · inbound
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ed17aba-e7da-470c-8e1c-0dccf8518699 · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8802a63b-d80b-4ce6-b2f5-eb01ac5bf8ef · inbound
Multi-Agent Collaboration Mechanisms: A Survey of LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b32ffa91-5b0b-4bec-8a41-5755d99743f5 · inbound
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fe1e8af-6e58-4e29-8e82-be82ed0ee751 · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5c39fabb-892a-4e64-8eb7-441385c8b786 · inbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f90d5d43-c538-407f-9790-41783aabd005 · inbound
The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04aaa514-edc1-417e-8aa5-9e22085d8c99 · inbound
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9c9a3ff-c413-4d5f-845b-cc6fadf33474 · inbound
Rethinking Human Preference Evaluation of LLM Rationales Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d00c505-e7a5-4e03-89aa-5c8ffbab6b7e · inbound
TSVer: A Benchmark for Fact Verification Against Time-Series Evidence Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 20479b87-f7c6-4495-b5ad-676257f27f08 · inbound
LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 508110b7-512b-432c-9656-451603ea0949 · inbound
Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88bf4585-498d-41dd-a5ec-289dfb7042e9 · inbound
Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57c1488-86c7-4bb4-ab30-95242d27c7db · inbound
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988ad47f-da48-4ce5-a6b1-6a13ba1ffa12 · inbound
Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 23dbe2e5-3896-41a3-aa64-3dffed36e444 · inbound
Internalized Reasoning for Long-Context Visual Document Understanding Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5696c388-44b9-4c34-95b3-4aa9c37d11a1 · inbound
Internalized Reasoning for Long-Context Visual Document Understanding Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9159a9c8-47fd-4f0b-91d5-d525b6c81d28 · inbound
SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4e90c56-fcb3-4843-80dd-63f663e666f9 · inbound
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation be803fee-5694-4c59-ad52-6e571d2bfe0e · inbound
LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 20bd82eb-43f9-4303-a8c1-eb40cf2202ab · inbound
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9492a5d2-a2f2-4c93-91b7-894c9fe32c06 · inbound
Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71ceab0a-c40c-4aba-b195-8f6c45c1ee4f · inbound
Confidence Without Competence in AI-Assisted Knowledge Work Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e07459e-a70e-4115-a1d2-e45dd3cec6f4 · inbound
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain? Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e5f4802-f1cd-4094-a3f2-f662c884fa60 · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42f5b3fa-cd29-4735-95a6-bbb94f22fbe9 · inbound
LATTICE: Evaluating Decision Support Utility of Crypto Agents Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f95d9cb-6daf-4e65-98fa-1eb09ce21f6c · inbound
When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4717e454-09b4-44c0-ab02-08fc9d983156 · inbound
When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4ad8f3b-56fa-414b-b30c-5fcfc3798edd · inbound
Analysis and Explainability of LLMs Via Evolutionary Methods Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 77f57b36-a566-443b-b2e3-887b9e40e599 · inbound
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bb8e06a7-9b25-46f0-b8aa-ef96652c4b41 · inbound
Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cfff2cd-51ae-4974-9bcd-58b9d64ccb60 · inbound
The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ff60c3b-d11a-49a5-88fe-faf44a00eacc · inbound
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76c3cf12-7ab9-4eb5-a36f-b95dd1038667 · inbound
ProactBench: Beyond What The User Asked For Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e8ea536-a7ad-4355-99ab-6f25de15f9f5 · inbound
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c8001ea3-a8bb-45a2-9ad0-bbc0f631dabf · inbound
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a208242-f5c2-488a-aa63-ec42280d6d09 · inbound
AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49371657-65f0-4fad-9a63-384681b87a43 · inbound
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dfbc38e3-72e3-4714-8fc4-92c05da62fdf · inbound
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fba3910-90e3-47e3-8633-da161ca5bf27 · inbound
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 302d5554-ffb1-41ba-b394-9d0bd85461d8 · inbound
Interactive Evaluation Requires a Design Science Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58d3fcb2-1f2b-421e-a77f-3b0e5ac76424 · inbound
Engagement vs. Commitment: The Economic Trade-Offs of Polarizing News Content Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03363f7a-2574-4f4d-8b89-4f6839531ace · inbound
SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd68035e-29db-4470-a76f-de948ef291a3 · inbound
Design and Report Benchmarks for Knowledge Work Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb50ee5f-f3f8-446a-bba2-571b487fadd8 · inbound
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5dddc29f-a976-4ff4-a8ae-fb3b6f52ebb3 · inbound
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e1651dd1-5a50-4a9a-aa2a-022c2cd756a8 · inbound
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 371314e1-cd82-4728-9da0-8b5567f94400 · inbound
Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69bbf51c-54e3-4b1d-a7e6-31ef517638ff · inbound
Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e2e5a9c-e207-4e35-8b97-0e67991d320c · inbound
Pairwise Reference Alignment as a Model-Level Ordinal Observable Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f2425e7-1786-48d5-96c4-41f0fba894f9 · inbound
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cff335cc-ace1-47a5-929b-402bda8803e9 · inbound
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adc9565e-4185-48aa-bb93-651eb5b1c336 · inbound
A Finite-Calibration Regime Map for LLM Judge Panels Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d58710c6-de9a-497d-b2c0-3b34d70bcbed · inbound
From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 915bf9c3-9d18-4d79-9a82-343735faa7cf · inbound
RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29a0f2b7-83f8-4a06-bbe1-5e1890a7a32d · inbound
Characterizing initial human-AI proof formalization workflows Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 65b99aeb-3ec9-4ba0-99a1-81f7c3d6c7bb · inbound
Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ed6470c8-fb15-4353-b068-1ea25f26b8ab · inbound
Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5233192-3592-4e6b-b655-325d90713e30 · inbound
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6540e42-6e1c-41a7-a179-3745e148d073 · inbound
Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dbcd1090-bd88-4a5e-82a7-4448e0ef5fd0 · inbound
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01eb3df8-1d6f-472e-b7a4-229b3150dc6a · inbound
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 26ca8f70-4bbd-4c25-9af6-baf3bdc742f5 · inbound
BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a3c0d88-b33e-4a02-b932-f269fcf3946e · inbound
Evaluation of Small Language Models for Arabic Language Processing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fac787c1-5121-4ab7-aa74-f87fad71e94e · inbound
What We are Missing in Multimodal LLM Evaluation? Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5d2ea79c-9aff-4a06-bbeb-4fca48a15fcd · inbound
CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c91fc324-c48b-4935-b971-07f1eb83269b · inbound
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 667234eb-960d-40bd-80e8-99a2519905b3 · inbound
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0ec5ac9-52c8-405b-83b4-d5bdbd98e582 · inbound
Open Problems in Constitutional Preference Reconstruction Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9886c68a-673e-4848-8c85-c76604d6b443 · inbound
FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c8a0f39-84f9-4b74-b096-afe663c86cb6 · inbound
FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b132ae3-9955-43be-9d57-3142f906bbcd · inbound
Meta-Benchmarks for Financial-Services LLM Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8c96aa6-ab97-44c2-9693-257764ebc289 · inbound
Dissociating the Internal Representations of Sycophancy in LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7c91878-0858-4415-8d2e-646c3601dea8 · inbound
Dissociating the Internal Representations of Sycophancy in LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6f0dd8-f2db-48bd-b322-a6e7c8081cec · inbound
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0eb6e539-f6d7-4683-a86c-f3d8776d222e · inbound
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 145
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3695e37-570c-45da-9ddc-0488959a07f7 · inbound
RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc02a58-a58b-4efd-80a9-a74db86a856e · inbound
ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afa2c31-87b0-4c15-858f-f118faad26df · inbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b015e0-5c66-4bcb-bb08-a301f121422c · inbound
Economic Evaluations of Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecea5ad7-4b16-48b1-9baa-cd0ccf4f2c87 · inbound
CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03550489-49bf-4da5-b5e5-d9163a366e47 · inbound
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48237d8a-376c-4304-87b9-3336c346a1e5 · inbound
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d1caaf-2e4b-46f2-b385-791fb0520842 · inbound
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.