Pith. sign in

Paper Citation Record · LEDGER

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 91 inbound Pith citation observations for arXiv:2403.04132.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.04132 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 142 of 142 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 91 of 91 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:18:53.743378Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

322
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3b852332-53e7-4b07-9e7f-5afdd21562e3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T15:13:25.347366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:5fdbbcc3aaeb7e78f2959e0bd8ab22722ee1ceebc229dabcb59062e5872cef5a

Observation d142e643-0925-47b4-9744-8455131e05f0 · outbound

This paper cites Statistical behavior and consistency of classification methods based on convex risk minimization.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Statistical behavior and consistency of classification methods based on convex risk minimization

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.353915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:0a2c3ad09d2acb3eb440d4176ae27aaf454feea01972449925f33e7984e106b9

Observation 787e3bc8-c303-4384-aae9-92022ddcfd50 · outbound

This paper cites emnlp-main.608.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference emnlp-main.608

Reference 3

Resolution
verified exact
doi, observed 2026-05-13T15:13:25.357448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:d8985909567f8e8ff99fe940f1b8e1deb78d0bfdf29e1d8d8fe14c3e9d5a927d

Observation 5a3589a1-94a7-4f14-a295-2b476f74e970 · outbound

This paper cites GPT-4 Technical Report.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference GPT-4 Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T15:13:25.362893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:6ae3b92aa4bb6a88722e74ec68880caedb3066e42d5eff3a9e471b051941ae28

Observation 35aee26c-adcd-40c8-b846-637c71cf6786 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 5

Resolution
malformed identifier
local_arxiv, observed 2026-05-13T15:13:25.368965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:25cc58cc74b2dc34c6f2b8a1f4a24b1fd05ae246d948f6aa5fdb1ec57b91f9ef

Observation b8f8976f-bf49-416e-aa11-65afed9d7222 · outbound

This paper cites Travel Itinerary Planning.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Travel Itinerary Planning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.372973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:f8fc44aea9f1ae05c88af036001c25c9bf7cf7bcfdde63ed9ad79f6db751b2d5

Observation a3a773f7-2951-4bdb-b665-571133c1c05e · outbound

This paper cites It houses collections of European paintings, a medieval and Renaissance collection, ceramics, French sculptures and more.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference It houses collections of European paintings, a medieval and Renaissance collection, ceramics, French sculptures and more

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.377708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:39c1924d569a300b0fa63d773ce8194df591083672aaed67e4e803c1094990fd

Observation ff5e5e2e-da36-4b7d-9be5-474fc07dabaf · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.381544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:e97bfad6d80277ef5f3886d62270bddf7e8b3c3854776f3b3467d1326b670250

Observation 04aaa260-9393-47be-8f71-d584be635ecd · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.384969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:37a29d6617d81048768a03c8fea86a0bb0063ea3b83600a3ac3c92a4c8520408

Observation 7c6ceb9f-17b8-4e34-bc67-3cb6bddda03f · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.389165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:4ef1afaa2470515f1a951418cd7b16ba0c1a22c6b70aaa343d5bc5282f34f6f2

Observation 2956a5a4-2fa4-4f63-945f-3f35be0c6caa · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.392667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:7eac8f809ddde308adcc45e9b0b88624922717b75afac3c05e1446aa80bb10d5

Observation 3b60a4b1-ed88-4dd8-8c66-3867ec5d765c · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.395924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:5855ce8610f189ae8c88844dd06d139e99735a796d47d162aa5f55f91bba90fd

Observation 88f3a4e1-4bf7-4b77-bf07-7d0a3c52fa44 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.399098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:6d5085ef7ab2a7cca6c65e5e4ecc93fb2320c3c6179a8014d496a3d35b00fa8c

Observation b9c0c2fa-ce7e-4ad9-a487-74682a1c16f9 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.402166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:9fd3106094368d86f91f9c0e3b2287d416ff7a04082b3b6498d116a8c3d85e4e

Observation 59899fac-74b0-4cb2-952c-7f4fa5da14cb · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.404987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:10b45584b5de7d3e5cb828ab9838ed803f9afd8360345364bd72632e6f731c50

Observation 8ed6558b-229d-4c0d-bde5-41cceede1b47 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.408873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:ca860d2e9afff8175a0a66e052a9c0e8bd754a25c8d4657aa4bb48b696f7557a

Observation dce4dfaa-a604-42bc-bd6a-068216c35d74 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.412629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:dfb54b623fd2713dab5769a2bcbac78fb51d730ba089f65138547ee4d3c587cd

Observation 2d465c1f-d45a-4de6-bc5a-c1c8ea01fcfc · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.416107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:989f059f63dba4033607688dffc5a1ada2a43f0d00d15abf771db73727cb03a4

Observation 3f2c0d46-da3a-4aaf-ac5f-1d054d136ad5 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.419406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:71de3579b6a89300787743d9f2bd5f2ccfc523345ad3ec02eded7fb74e2d45eb

Observation abc05d99-10c8-4086-acfd-e241ee92cd02 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.422286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:1894f9beb91a4c663a0fa4092cbffabe0b9f8703144861cba6d65bcde63a1a18

Observation a5e625f3-9fc7-447b-97cc-a324ac44f331 · outbound

This paper cites Remember to check the opening times and any COVID-19 restrictions before you visit.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Remember to check the opening times and any COVID-19 restrictions before you visit

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.425537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:ac908778ae362d51e537396d8734205baabc77a99c168f4025449e3f66f96cd8

Observation ef59e6a6-e601-4c3a-88d8-a99d3c1029c3 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.428646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:6c41d054777b4812e08d1e44fb69fe5fc296133494cf3cad7dc874b316f4ba8e

Observation a11181a5-6c78-4a8a-8e76-b6af720e4c12 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.431553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:315904270b34c44a3706b2277c7cee7ee5211204b612da32cb657516db5dc1cb

Observation c7e98b0d-fc78-495b-92f9-61310a62788e · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.434909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:a14e973675e7710866239ba8b9c41ec9c098b0eff1ae5df5b54b39cf0dda355d

Observation d8403948-795c-4bec-a05a-a42863595158 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.438581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:e746b7c088857fafd4431ded5362d5d7c9f463198c8019b92fdc2943f4961a17

Observation 19aa2493-1f99-4adb-8086-7d40831e636f · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.442218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:4726ba0bee857dd3c946df6fd6c05f12a22fca0724cf0dd197798179874a88e0

Observation bd4c204e-295a-4d04-a2b9-db2513e28cf5 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.446063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:17e0de16eabc54188b8cfc175eed0d3b45b4a70d1439f32336580841f9023177

Observation 96c83534-083a-46b0-a5f4-c2bbba8428bb · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.450297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:be29007ad39392571d2fcaca24c64fb9f7f988d5296783fd66a460fc10549940

Observation 97c5cc0b-79c7-405c-9704-04de5d1b26d2 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.454143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:0ad06181a7f9ef8ea3b9984af2ffe765417e4602049c809c87971d420d6a6bed

Observation 75062dbc-d0a3-4fc1-9764-fc9711c1425f · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.458259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:d30af5e871a9f2b81240b7b0156b5c51242413bc91a43e85736a2a4b423b1cec

Observation 0ee6e5ad-c4cc-4d44-a8c6-2b093683ecde · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.462987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:76876e49c247ef38baaaeb0ee0fccc33dd5a68d9876fade97ce60d46c46a3859

Observation 2b302ad6-b485-4af5-ac58-48afc5f9b1d1 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.466618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:3d39d2e2fb5f28669e5fdbf23e77c80bf8d1919b4dc3e03e55d2c5e7a0f7a7b6

Observation 98978a21-05da-453c-bb0f-3ca2679f8132 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.470473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:ad3c5b8bf30a7c081f48c9460695e20638937444ceec8318c08319bf678ae1f6

Observation 6f7cd57b-7194-449b-b645-da23a8cee567 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.474517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:0576a269567b699281f4d24eb032305da5ce07a36e807bf253a224d5eb405d46

Observation 48c6805a-20b5-47f7-b74c-b5b73aac8339 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.478159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:a101d6417b716894f4bbad33ecabbe9545c9c6faa217a991afa7d118e4a8a25d

Observation 795c00ba-9007-40af-a9c1-191090fe20ad · outbound

This paper cites These are just a few ideas to get you started.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference These are just a few ideas to get you started

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.481889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:f2f82188ccca91f3a685f828a69fee82e43ec493b1f287dfe4166f5b935931b8

Observation 068acaa6-f5a5-4991-b206-0b2471cd8873 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.485104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:2c215ee34d8285811223e0a4d2a2351a17a02b9d1ef1d139ef4efbde72fc4652

Observation 8a6f4ef4-59b7-46ed-9c85-7658b832cefe · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.488773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:a479303142879729062d1601b6bfe496502047583c5b9fbd760d84083c728245

Observation f170670c-4834-46c4-a5b3-9e3fda7ec112 · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.492031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:cb2fb18072b3a3dd564e0c2f31a90705e4b734c00812d1c3216fe1bf16dcda8d

Observation 1d8eb5cf-b8d5-4c00-a1cc-59afb239712c · outbound

This paper cites an unresolved cited work.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-13T15:13:25.495280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:3c2bb2fa0e8a09bbbfd800d2e306c96cb6a69a5c8c94b7e261382b4927110fa8

Observation 28a77312-fa14-4b52-bc9f-4afbf4aefaf6 · outbound

This paper cites My final verdict is tie: [[A=B]].

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference My final verdict is tie: [[A=B]]

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.498802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:d55f8442d57aaf3fda6d4f185925b3d9d63a3af553306b98e6e2a0fa60694ba5

Observation 3691fdaf-2448-44c4-b6e1-5f885dce0214 · outbound

This paper cites 27 Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Continuously gather and incorporate customer feedback into the product development process.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference 27 Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Continuously gather and incorporate customer feedback into the product development process

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.503334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:19b7dd31d67348db9d1bb668be7ff1a14e65a9f147038e4b4be5e8867307c3ab

Observation c7de7b00-0f13-463c-8e3d-9ef264470775 · outbound

This paper cites - Align your product’s features and capabilities with its value proposition to ensure it meets the expectations of your target audience.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Align your product’s features and capabilities with its value proposition to ensure it meets the expectations of your target audience

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.507261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:354063cfaa7b0e896f7f140efde2daa5e392c7d30f02b034ab79cdcd2b156e73

Observation fab17ebb-364f-4e64-bc43-81c393a82421 · outbound

This paper cites - Validate assumptions and hypotheses through experimentation and user testing.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Validate assumptions and hypotheses through experimentation and user testing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.511278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:1b85a6313ed06533209a76aa071e256b75e63aa3de1baf2c00b1aaf314a6c74c

Observation 691382cd-cec5-41b0-995d-813bb4a3d71c · outbound

This paper cites - Be open to pivoting or making significant changes based on feedback and market response.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Be open to pivoting or making significant changes based on feedback and market response

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.514912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:d9159010e9de49294517af5ebf75b11085a3d3dd1a4a2c36db3474d4616162fe

Observation 736c5688-4e45-44e9-b9f2-d1741812e124 · outbound

This paper cites - Establish key performance indicators (KPIs) to measure the success of the product and track progress over time.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference - Establish key performance indicators (KPIs) to measure the success of the product and track progress over time

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.518950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:5b2e0cee5dee3a43129a44706cd76c30584470dd5f5e4f1c7f1c8f5d1677353f

Observation a1938a07-32ef-46e6-8c6d-115900e32fd5 · outbound

This paper cites Founders must be obsessed with their customers and be willing to put in the effort to understand their needs.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be obsessed with their customers and be willing to put in the effort to understand their needs

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.522664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:12545442c7c6d9aedb94e304168bfc16bb3f4ef6cf242b1a5bc291b905c66d88

Observation ff7b0c08-f83e-4cde-944c-007c14c6514d · outbound

This paper cites Founders must be willing to try new things, test hypotheses, and iterate on their product based on customer feedback.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be willing to try new things, test hypotheses, and iterate on their product based on customer feedback

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.526414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:cc61d03f44aeee457374b06ca86d7771d32d95752a4c93113bbc2ce71e9c929b

Observation fb6c5f2a-218f-4afc-b89c-3f40e719a856 · outbound

This paper cites Founders must be able to identify and prioritize the most important features and functionality that deliver the most value to their customers.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be able to identify and prioritize the most important features and functionality that deliver the most value to their customers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.530386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:585b958813d4b169dc0bfe9000b0a0c4c1f0440d840f8c1208bfe3d40a528a44

Observation a053b572-8a53-4e6f-9303-f719c5f490bd · outbound

This paper cites Founders must be able to work effectively with these teams to develop a product that meets customer needs.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Founders must be able to work effectively with these teams to develop a product that meets customer needs

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.536031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:12bfdc95ef13873566c3102eb063af58af4cb9300814ab146c2e2397ac652b9f

Observation cf8ed48f-61ae-466a-acff-af8de76fc54e · outbound

This paper cites This includes analyzing customer feedback, usage data, and other metrics to inform product development.

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference This includes analyzing customer feedback, usage data, and other metrics to inform product development

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T15:13:25.540637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:13:25.320641Z digest=sha256:ed685755227cab156cbd6421271a7f45d8536b1c7cc14b23643d70327ce207bd

Pith citing papers

Observation f6666ca9-1e4f-4bdb-b702-6f659369aceb · inbound

SGLang: Efficient Execution of Structured Language Model Programs cites this paper.

SGLang: Efficient Execution of Structured Language Model Programs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T08:20:01.011625Z digest=sha256:91f6766924388504e543c176ad989f92e8eccd6b0d8e2ea14ffaa8539e8a4813

Observation be987ecb-0262-42e5-baaa-7a61352de344 · inbound

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents cites this paper.

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T12:06:13.697487Z digest=sha256:296328da82f2e8ca6e23a3fa1547f9241a8fe08f75cc771d918475cab458e89e

Observation 8a07d66b-ea87-4efa-9588-0d8643c0835f · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:432f781067684c38ad4476452dd981009b772770bc4e90f926406f6b32bd767b

Observation d53d4858-397a-42ed-a2d5-1225a7f1286b · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T23:10:40.889832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:3e1307bdfbf875138177d4b48f9d7d23a174cebfab6aa8f0c91e1acaf26274fe

Observation 1c431ca0-cfef-482f-85f4-a53d9773a37c · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:05:03.840660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:838612b86cb0b97de67458b4adddbba583b1a69945068ed5ea949f16166b5fc9

Observation f619c2b6-d3c7-47a0-9f86-70eb839dc106 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:48:26.439008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:2c78230ce97cb923c03dd02f9c42e4a9549db7ba9a9394f34cd06d2face06e95

Observation 94c915b5-1495-4b1f-84e9-3f91c581087c · inbound

Qwen2.5-Coder Technical Report cites this paper.

Qwen2.5-Coder Technical Report Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:33:38.867604Z digest=sha256:6f050bfe8c8755a8c4f355e98259ead54c18e69c6e9f8503a064651277710b43

Observation 200f367c-3777-429e-bee8-0ac5ee91faeb · inbound

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models cites this paper.

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T09:09:15.078017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T09:09:14.884516Z digest=sha256:bf83aab5119b45c42da80e86bf240be72d3479d93b5f6ab3212c002feedba043

Observation 9ed17aba-e7da-470c-8e1c-0dccf8518699 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 185

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T07:51:13.377092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:337bed4d0faeb8253eac715e64ddb5d94b04c5e41a965d94a88005d1e39299bc

Observation 8802a63b-d80b-4ce6-b2f5-eb01ac5bf8ef · inbound

Multi-Agent Collaboration Mechanisms: A Survey of LLMs cites this paper.

Multi-Agent Collaboration Mechanisms: A Survey of LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:54:54.330650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T15:54:54.146003Z digest=sha256:cc675d5f1c8e058f6d79be1b400feeae3999431b3b34887629296ba04dc6bb62

Observation b32ffa91-5b0b-4bec-8a41-5755d99743f5 · inbound

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong cites this paper.

Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T05:37:36.068625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T05:37:22.955895Z digest=sha256:73170e13c7703fdd71a51a8afcb9322be08fcf69bb94753c1a09bb5099d4e426

Observation 0fe1e8af-6e58-4e29-8e82-be82ed0ee751 · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:36:58.884817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:a37dde797784dd35844e34b09c576d31c86ef833fc7ca7dcd56f7c089056f1e3

Observation 5c39fabb-892a-4e64-8eb7-441385c8b786 · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.076312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:eeb6f6af3d710cf012e656e54a3ca553cb21fc822a0c627296fab5dad2d4c9ae

Observation f90d5d43-c538-407f-9790-41783aabd005 · inbound

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering cites this paper.

The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:38:55.329835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:38:55.159673Z digest=sha256:113b27faa8b611059d8d9ef8009076eca356606accfde3df228267097805975a

Observation 04aaa514-edc1-417e-8aa5-9e22085d8c99 · inbound

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection cites this paper.

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:31:50.077992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T20:27:34.917353Z digest=sha256:c48edf06152ff8f35709c6739d885e06323b2f5bcc8bbae96551adc7aa5c98c5

Observation b9c9a3ff-c413-4d5f-845b-cc6fadf33474 · inbound

Rethinking Human Preference Evaluation of LLM Rationales cites this paper.

Rethinking Human Preference Evaluation of LLM Rationales Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:27.863520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:27.863520Z digest=sha256:d28bda84d0c58bd04ec4725b68b09c084c8435a0e0ffcd0bdcee50f5787d076a

Observation 0d00c505-e7a5-4e03-89aa-5c8ffbab6b7e · inbound

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence cites this paper.

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:00:34.071723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T00:56:33.806425Z digest=sha256:43b029e6253fc46953a784a1c43cc409c30a8a61bd19cc8c5780a604fa70ac73

Observation 20479b87-f7c6-4495-b5ad-676257f27f08 · inbound

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems cites this paper.

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:53.777048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:47:28.248540Z digest=sha256:3e5b5083e80ad3bdb75ac94bf2adc801c33747d6fce58bb3e815c8e0b3402f68

Observation 508110b7-512b-432c-9656-451603ea0949 · inbound

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests cites this paper.

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:20:53.151926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T11:20:20.651794Z digest=sha256:c8dfd47a3ca0c5e0f768e88f8bf55598d6dff124a5300f7d6e9f34944d028e7d

Observation 88bf4585-498d-41dd-a5ec-289dfb7042e9 · inbound

Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing cites this paper.

Dynamics of Learning under User Choice: Overspecialization and Peer-Model Probing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T20:24:29.500934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:24:29.500934Z digest=sha256:3e28b5faf8cdd9def71722788fa288a4a8a99826a2cd9c3583c6f0d4db6d92a8

Observation b57c1488-86c7-4bb4-ab30-95242d27c7db · inbound

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines cites this paper.

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T17:52:54.735172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:52:54.735172Z digest=sha256:f213a602766ac7e3ff778fee4771b73fe52ca36c8ae0071a53804e7bdf8f2d49

Observation 988ad47f-da48-4ce5-a6b1-6a13ba1ffa12 · inbound

Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini cites this paper.

Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:18:21.868181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T00:17:41.692833Z digest=sha256:02f8b1edc85ddcc261e225dcf948c991d9506fd227f162775d635aacf8c6e3ef

Observation 23dbe2e5-3896-41a3-aa64-3dffed36e444 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:53:28.175114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T23:53:19.148407Z digest=sha256:b0621d1aa1d9c77622f0eed0f30f4f153a86886e5f280e89686f8ae064560fc6

Observation 5696c388-44b9-4c34-95b3-4aa9c37d11a1 · inbound

Internalized Reasoning for Long-Context Visual Document Understanding cites this paper.

Internalized Reasoning for Long-Context Visual Document Understanding Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T15:50:49.083652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:50:49.083652Z digest=sha256:097e86e8ded0b66bcd715e848951469231744d82b8cfb00080d5aaa0cd93924a

Observation 9159a9c8-47fd-4f0b-91d5-d525b6c81d28 · inbound

SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics cites this paper.

SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:32:25.959711Z digest=sha256:e140794adbabe650f5a64236c85ce195465e9b9b094fea01265cdd48c6f20ded

Observation c4e90c56-fcb3-4843-80dd-63f663e666f9 · inbound

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads cites this paper.

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:21:27.428360Z digest=sha256:cdd9a603bec4ac0d80de3a3d310abb206b369c8cc155e49a5129895d3a4e9d9b

Observation be803fee-5694-4c59-ad52-6e571d2bfe0e · inbound

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency cites this paper.

LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:39:49.792609Z digest=sha256:9ed846f0d7f6cfec253547e0fe584c544600ad0cbb189e14244d69af1b02b6d9

Observation 20bd82eb-43f9-4303-a8c1-eb40cf2202ab · inbound

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks cites this paper.

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:49:32.983778Z digest=sha256:e26d272f1ff4eb702b9a2c0f2d9f8eb414425bb1cf2144c4cd8f3ada6c3763e6

Observation 9492a5d2-a2f2-4c93-91b7-894c9fe32c06 · inbound

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models cites this paper.

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:28:25.839101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c274a5cee07978fafa6dfb099d6ed57f065517a55334fb0d0b97e06a52130801

Observation 71ceab0a-c40c-4aba-b195-8f6c45c1ee4f · inbound

Confidence Without Competence in AI-Assisted Knowledge Work cites this paper.

Confidence Without Competence in AI-Assisted Knowledge Work Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:02:54.050021Z digest=sha256:0b7c26ef33007fd9ae9a8a5cbcdaa609a08c5dab05f813b0288f345c0efc0cb2

Observation 4e07459e-a70e-4115-a1d2-e45dd3cec6f4 · inbound

Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain? cites this paper.

Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain? Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T02:28:50.416827Z digest=sha256:cfa070edc0960be76837dfd85e265b7365ec140c442e3ad1bcf71fc3e936ed0f

Observation 2e5f4802-f1cd-4094-a3f2-f662c884fa60 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:632009096b901d1c90879db01d2d0b20188ef708c1f10c60c0a4b6e4bddffd2c

Observation 42f5b3fa-cd29-4735-95a6-bbb94f22fbe9 · inbound

LATTICE: Evaluating Decision Support Utility of Crypto Agents cites this paper.

LATTICE: Evaluating Decision Support Utility of Crypto Agents Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T13:30:46.523784Z digest=sha256:58385e8d3dcee0b93baba321e96c087878b6d59c848acd62ee9f8a5742dff280

Observation 3f95d9cb-6daf-4e65-98fa-1eb09ce21f6c · inbound

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems cites this paper.

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T02:13:17.642291Z digest=sha256:085660122e1700f19e6d5aa71d33a3ad87e34c1ae50e5c5d42b4f87820ab3f02

Observation 4717e454-09b4-44c0-ab02-08fc9d983156 · inbound

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems cites this paper.

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:29:00.554110Z digest=sha256:01e6432b3cdfc87a1c526f6e5151aba0fef0f6733373b1242845c5d240ba858d

Observation c4ad8f3b-56fa-414b-b30c-5fcfc3798edd · inbound

Analysis and Explainability of LLMs Via Evolutionary Methods cites this paper.

Analysis and Explainability of LLMs Via Evolutionary Methods Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T20:37:40.932811Z digest=sha256:d3c783844b72d86db3f83351f046b0b209ee26a0119a0a7c05ac2b7eab81252b

Observation 77f57b36-a566-443b-b2e3-887b9e40e599 · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:d3c999de25e23aec29f8e8c817b1cd89ba12897d0e140c9764642e625a43bb5b

Observation bb8e06a7-9b25-46f0-b8aa-ef96652c4b41 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:58aa17042a425faf8c4e8ad933fe7b5d45f59a5eba1027bc5762c71fef7f4b29

Observation 0cfff2cd-51ae-4974-9bcd-58b9d64ccb60 · inbound

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models cites this paper.

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:44:31.583936Z digest=sha256:c7bd182a23fb63959a10219ce74cb95f1995ea62421f97152d81b796d91434b1

Observation 4ff60c3b-d11a-49a5-88fe-faf44a00eacc · inbound

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators cites this paper.

AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:54:36.868103Z digest=sha256:ad76e6f18a9e459296c0dc02a6df35df49b36fb68cc78e7c75555f270539e083

Observation 76c3cf12-7ab9-4eb5-a36f-b95dd1038667 · inbound

ProactBench: Beyond What The User Asked For cites this paper.

ProactBench: Beyond What The User Asked For Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T02:14:01.145443Z digest=sha256:be791d249fa6e7df036d83f3a5c3801d49eea1e12a675f5d6e55177c76e29b4d

Observation 4e8ea536-a7ad-4355-99ab-6f25de15f9f5 · inbound

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants cites this paper.

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:28:13.317630Z digest=sha256:086f68359eefa5425ad7ea4970214510727de011e9e5c6d41765446d42919460

Observation c8001ea3-a8bb-45a2-9ad0-bbc0f631dabf · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:d6031116878104f7394db159da88eb4c809be1eba733f42fd060d09df72bc737

Observation 6a208242-f5c2-488a-aa63-ec42280d6d09 · inbound

AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents cites this paper.

AI-assisted cultural heritage dissemination: Comparing NMT and glossary-augmented LLM translation in rock art documents Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:15:04.643040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T21:07:08.026581Z digest=sha256:90318b7982536c0f501d806ac43b6b2030b59b302b1aac0ca4d1a956d1181be0

Observation 49371657-65f0-4fad-9a63-384681b87a43 · inbound

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation cites this paper.

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:08:58.379385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T03:06:54.507594Z digest=sha256:868940f68e67b04227baa458e1a1373b416a239d68c3a450391a444006eff8c9

Observation dfbc38e3-72e3-4714-8fc4-92c05da62fdf · inbound

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation cites this paper.

OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:53:43.622699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:51:47.393589Z digest=sha256:c8d3fa02a96da2e2eb89edf39a73c99da600fd8fa13e238fdeb9aadc5b2b0eb6

Observation 0fba3910-90e3-47e3-8633-da161ca5bf27 · inbound

StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video cites this paper.

StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:33:50.709916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T23:32:38.332828Z digest=sha256:59946c1a66bd3fd99e6a9ba7071a0235f3987fafb26bd15e3a0ca5416cab0d9f

Observation 302d5554-ffb1-41ba-b394-9d0bd85461d8 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:58:14.097172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:029fe26fc9873f3adfc64d128ce47602966186fa07957ee6607c8b90ded15e3c

Observation 58d3fcb2-1f2b-421e-a77f-3b0e5ac76424 · inbound

Engagement vs. Commitment: The Economic Trade-Offs of Polarizing News Content cites this paper.

Engagement vs. Commitment: The Economic Trade-Offs of Polarizing News Content Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T23:32:52.791540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T23:30:19.352088Z digest=sha256:437af1986436c4f7e9748f3fa2067c986dba9bec1b3b9690cc4f5bb41e1164ab

Observation 03363f7a-2574-4f4d-8b89-4f6839531ace · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:38:12.069195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:0d578d19b89c840dc845e4b7e46be69954992f0defa3696e2daa89e387e003c5

Observation cd68035e-29db-4470-a76f-de948ef291a3 · inbound

Design and Report Benchmarks for Knowledge Work cites this paper.

Design and Report Benchmarks for Knowledge Work Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T04:40:23.466577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-25T04:39:14.319133Z digest=sha256:5083002a9f2e18f2273c3455c8d29ba038f57a573e9efa483fd4c32aa8969bf8

Observation cb50ee5f-f3f8-446a-bba2-571b487fadd8 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:34:48.434672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:d2ee239109ef36ddb8f8d262790d30b9deeb94ef3e28fc16b73c96b0d51d54e8

Observation 5dddc29f-a976-4ff4-a8ae-fb3b6f52ebb3 · inbound

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems cites this paper.

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T00:24:04.104979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:d1f8ea56f9dd67f4712a5405373b54e545a57dbe12d13e6470457410af205f03

Observation e1651dd1-5a50-4a9a-aa2a-022c2cd756a8 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:03:26.793885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:ae03992392f9196265bc739fd535bc68f7fe527c7c9a48e1d489e9a2ad5944c8

Observation 371314e1-cd82-4728-9da0-8b5567f94400 · inbound

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons cites this paper.

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:53:32.261037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T06:11:48.446361Z digest=sha256:0a8cd634f6a4a628e0d09203c126ef20cc18170e967ed8fa455416bee51dd591

Observation 69bbf51c-54e3-4b1d-a7e6-31ef517638ff · inbound

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles cites this paper.

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T06:03:08.867810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-29T05:54:39.899517Z digest=sha256:845090ff5ba92ec989eb3193fbca60c2eb160979a63cfcefbe670afb10bb32df

Observation 4e2e5a9c-e207-4e35-8b97-0e67991d320c · inbound

Pairwise Reference Alignment as a Model-Level Ordinal Observable cites this paper.

Pairwise Reference Alignment as a Model-Level Ordinal Observable Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:25:59.686950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:46:11.450812Z digest=sha256:ef2dbbd5a25d60f6f2e51c1b61491530b5f72a9296282a29d216b1fa2c46ac13

Observation 5f2425e7-1786-48d5-96c4-41f0fba894f9 · inbound

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models cites this paper.

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:22:37.608482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T19:53:36.990878Z digest=sha256:ef1cde1611c5bd4f5d344354df5c3abeb432c95dc61fe958175ff12c89405c9c

Observation cff335cc-ace1-47a5-929b-402bda8803e9 · inbound

CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences cites this paper.

CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:32:37.621567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T18:36:46.381564Z digest=sha256:b5603e842e1031ae61988de6bbe87fd311921149c6b60bd2c3de3e1693449046

Observation adc9565e-4185-48aa-bb93-651eb5b1c336 · inbound

A Finite-Calibration Regime Map for LLM Judge Panels cites this paper.

A Finite-Calibration Regime Map for LLM Judge Panels Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:06:13.143303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:34:28.258222Z digest=sha256:4b1ddad0ffccdde58c516e4f230414498d71846400785fd896ccf49e6388473b

Observation d58710c6-de9a-497d-b2c0-3b34d70bcbed · inbound

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication cites this paper.

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:12:24.309494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:11:12.419269Z digest=sha256:03857eb9529bb8fa5a7833cc768fdacb6fb92df4e6feebd79dda02f072fc667e

Observation 915bf9c3-9d18-4d79-9a82-343735faa7cf · inbound

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions cites this paper.

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:36:29.207143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T09:56:36.860369Z digest=sha256:c9e4bb326de8da75669798c3731e2d13150d0f48c4cfb3b77c09dcf0c6c69eaa

Observation 29a0f2b7-83f8-4a06-bbe1-5e1890a7a32d · inbound

Characterizing initial human-AI proof formalization workflows cites this paper.

Characterizing initial human-AI proof formalization workflows Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T04:06:34.849422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T09:29:50.282874Z digest=sha256:b2b1243d6466e7ae2aa32371c5bf4cc3fc64ab7ac1bdf3650697503cb829caa8

Observation 65b99aeb-3ec9-4ba0-99a1-81f7c3d6c7bb · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T08:36:47.673444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:4fb4199a6849ce2ff235f722e800b63d0ed0e2635e789091800defd89ed7f793

Observation ed6470c8-fb15-4353-b068-1ea25f26b8ab · inbound

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios cites this paper.

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.169551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T02:49:01.527149Z digest=sha256:01decffaa66da6c696294439394863d287be82fa1f7666c183de9222412691ba

Observation d5233192-3592-4e6b-b655-325d90713e30 · inbound

Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies cites this paper.

Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:27:22.393053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T20:25:42.500234Z digest=sha256:9d2c163d37e4cb1c90fb3fd660f961d53498e8c82b78b58e734fc93f7b2a306d

Observation b6540e42-6e1c-41a7-a179-3745e148d073 · inbound

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces cites this paper.

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T00:37:29.845249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T17:07:39.645843Z digest=sha256:5387f29697aebc88bba9e07424e38b41c5c39331a98b085dec0c60fad971b1fd

Observation dbcd1090-bd88-4a5e-82a7-4448e0ef5fd0 · inbound

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System cites this paper.

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:08:02.942190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T09:43:12.633039Z digest=sha256:8ad4ea2d20c8740e27e1082bb1baf5dfdebfab6eea02ffcaa571afd38c435f09

Observation 01eb3df8-1d6f-472e-b7a4-229b3150dc6a · inbound

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing cites this paper.

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T05:39:40.119444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T15:48:26.303462Z digest=sha256:6c2a0710b31fc531c64848ad1600681970e67c49ffe18bfdfa0cbffc7a78b02e

Observation 26ca8f70-4bbd-4c25-9af6-baf3bdc742f5 · inbound

BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models cites this paper.

BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:39:29.446244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T17:53:46.206401Z digest=sha256:32547af350a9c266a2322c9654993ef85c74cdf3eccdd01f7113d8332497c8d0

Observation 8a3c0d88-b33e-4a02-b932-f269fcf3946e · inbound

Evaluation of Small Language Models for Arabic Language Processing cites this paper.

Evaluation of Small Language Models for Arabic Language Processing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T06:39:36.969233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T14:22:37.152936Z digest=sha256:d6eaf7ec565bee03c6bac136254fac84bea39a6052d84a04da15cbccb75726f9

Observation fac787c1-5121-4ab7-aa74-f87fad71e94e · inbound

What We are Missing in Multimodal LLM Evaluation? cites this paper.

What We are Missing in Multimodal LLM Evaluation? Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:39:56.690059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:31:35.307417Z digest=sha256:8716721df6ab55679eead9f2d6c880418564597a35dbdf3734f2d4e0f301572e

Observation 5d2ea79c-9aff-4a06-bbeb-4fca48a15fcd · inbound

CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval cites this paper.

CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:24:38.797790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T11:07:31.861143Z digest=sha256:291b28ffaf069ca1d1018b0429c8e5abddaa7eb3444a1065301ef8762fcc540b

Observation c91fc324-c48b-4935-b971-07f1eb83269b · inbound

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries cites this paper.

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:34:34.718843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T09:30:00.467709Z digest=sha256:cab47e38606a3a5f4f9238a1b9cd197920b0e03bfd4681324884562a2518524d

Observation 667234eb-960d-40bd-80e8-99a2519905b3 · inbound

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data cites this paper.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:54:44.499146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:5ae8f8fca833c9a2bcd1df53d26cf88378020e77446ed874b15f0b352cb4f065

Observation e0ec5ac9-52c8-405b-83b4-d5bdbd98e582 · inbound

Open Problems in Constitutional Preference Reconstruction cites this paper.

Open Problems in Constitutional Preference Reconstruction Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T06:54:20.510382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T06:49:24.255622Z digest=sha256:eb84bed159d53323d781cbe00be1b6550a6236dc9120188860109ae1ae01a47e

Observation 9886c68a-673e-4848-8c85-c76604d6b443 · inbound

FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks cites this paper.

FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:24:21.045668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T07:22:32.638556Z digest=sha256:a58e388f1e7b448f2ce4b61b3fe8310ccd4544977c13ae337b7b5d9bdd6a7a6c

Observation 4c8a0f39-84f9-4b74-b096-afe663c86cb6 · inbound

FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks cites this paper.

FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:15:29.983182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T07:13:43.014623Z digest=sha256:449fa727b2d43ff1d2d1480ff93ff5126c91b561ec2c5b24a1d86438e5a9a263

Observation 1b132ae3-9955-43be-9d57-3142f906bbcd · inbound

Meta-Benchmarks for Financial-Services LLM Evaluation cites this paper.

Meta-Benchmarks for Financial-Services LLM Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:18:22.689464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-03T14:08:20.432931Z digest=sha256:50513ba13fe59034bf96b0f6df5e5963832bf00142865a596e1977050127d124

Observation f8c96aa6-ab97-44c2-9693-257764ebc289 · inbound

Dissociating the Internal Representations of Sycophancy in LLMs cites this paper.

Dissociating the Internal Representations of Sycophancy in LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T22:06:35.339752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T21:56:32.778349Z digest=sha256:795f675a90f13c886f73377001ec34b2e54e969a02a0b6339590826cb80c6566

Observation f7c91878-0858-4415-8d2e-646c3601dea8 · inbound

Dissociating the Internal Representations of Sycophancy in LLMs cites this paper.

Dissociating the Internal Representations of Sycophancy in LLMs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:11:39.471149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:11:39.471149Z digest=sha256:0178fc330e4614984842bdac39817573b666dccc71bb5240db80d3210754a41e

Observation 2f6f0dd8-f2db-48bd-b322-a6e7c8081cec · inbound

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation cites this paper.

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:56:41.111012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T01:51:14.922841Z digest=sha256:d4fc04a54e60c5289757b5a537ed501302a15a692eee821d7b429fff2039d870

Observation 0eb6e539-f6d7-4683-a86c-f3d8776d222e · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:06.034941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:06.034941Z digest=sha256:7d7ce1f33fd795ed524cfeded800f66ca3e801b0cb9f3e9e8729adaf6e331cc9

Observation d3695e37-570c-45da-9ddc-0488959a07f7 · inbound

RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar cites this paper.

RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:59:51.029265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:59:51.029265Z digest=sha256:cb314407d5370b7f06c368a839b53ff57276abf2d19bb20fc8f3f0c81b6d4c93

Observation 5dc02a58-a58b-4efd-80a9-a74db86a856e · inbound

ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing cites this paper.

ContinuityBench: A Benchmark and Systems Study of Stateful Failover in Multi-Provider LLM Routing Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:06:31.303455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:06:31.303455Z digest=sha256:1cde8243a31f6f3c7a6adebb94464d218f673bb3150648fe318c88c2faed620a

Observation 1afa2c31-87b0-4c15-858f-f118faad26df · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.040725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.040725Z digest=sha256:6655aca8db68029b8a2205d476c8777b7e3bdf5ee7c31cae9a621a730a68a794

Observation b7b015e0-5c66-4bcb-bb08-a301f121422c · inbound

Economic Evaluations of Language Models cites this paper.

Economic Evaluations of Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T09:51:06.039948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:51:06.039948Z digest=sha256:1b23368a122c387c7b7e1ab169558c8676404b73ce5306507f935bc261616408

Observation ecea5ad7-4b16-48b1-9baa-cd0ccf4f2c87 · inbound

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation cites this paper.

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:22.182908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:12:22.182908Z digest=sha256:6e69d7cecd8e89a58aadde5123e720f233aaa06e5439bac93798251d7fdaaee0

Observation 03550489-49bf-4da5-b5e5-d9163a366e47 · inbound

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI cites this paper.

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T05:03:44.347887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:03:44.347887Z digest=sha256:f96b6fbe8dae2773cf35e4afa32e730ce88d408933662b716301ed87f20f537e

Observation 48237d8a-376c-4304-87b9-3336c346a1e5 · inbound

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset cites this paper.

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:40:15.001814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:40:15.001814Z digest=sha256:abd4cf26f93dd206a48f118da4171e007f1b611b0ba29b43a184bf77f34b4fb7

Observation f7d1caaf-2e4b-46f2-b385-791fb0520842 · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:53.743378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:53.743378Z digest=sha256:50380494edda0b37d6538100f6757ae0cb6aea865c9f4ef12de5f63a8bb1d603