Pith. sign in

Paper Citation Record · LEDGER

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

As of 5 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 100 inbound Pith citation observations for arXiv:2306.05685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.05685 v4

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:52:59.033645Z

measured 152 of 152 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 100 of 453 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:34:40.410309Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact27
  • verified fuzzy21
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

8605
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c3f9150a-0bf5-48f7-88b3-e3e822ae231e · outbound

This paper cites PaLM 2 Technical Report.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena PaLM 2 Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:59:27.667882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:ea9e98a7dca7058bf197e57ffa8c465e9ea316efb2a3986f58c833d1f2b9490d

Observation 7c33d917-af12-4745-be3b-a0054a21a150 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.063023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:a951cc23a0228d24b7c46357fd704d29b1aa0c2514773186f6b4566d35d4519b

Observation f56df5af-05f3-4af8-bf15-aeff97581859 · outbound

This paper cites Position bias in multiple-choice questions.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Position bias in multiple-choice questions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.181414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:7fad458f6d902df095de1b313564196ad89c62a87263d4df43315e06e983a94a

Observation a87b7471-459c-422a-9790-db258d8c2609 · outbound

This paper cites Evaluations of self and others: Self-enhancement biases in social judgments.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evaluations of self and others: Self-enhancement biases in social judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.184039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:b588f9ee4a81b4294436ebc26f07090b487e8a1be5837bcbe6379da76ff56bd8

Observation 09baee0a-afe2-4161-8531-0431dc543526 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:44:34.276011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:800dc193a20ae68dce93ee2c896f6ccb24374d02e9e34d7a8e5366b2605baa66

Observation a4505964-88d7-4bcd-8a49-358f75ddea94 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.155000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:2932e962a15a54af4636d0b2a5deb43f65f0fc86382cb0f0327400e9adc9aeaf

Observation 814939cf-b7f6-4969-91c5-9e1bec2f57ff · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.159832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:9ed413c2c0c72caac04d667e4312670df83ef1d2dc23ca6c65b7fdf8e0f70492

Observation 0d93a60b-5759-4767-8d3a-d98958c459d9 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.194175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:7dffa35cca3595e09410593418781603581fc999c193e6f56b1494cd959fc206

Observation 9ab99aab-cafc-481c-bd8d-885eafaa8def · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.163477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:2e5fe8a9690157f240a732bfda361ec645a1874c627838619856452997e716a1

Observation 9f2f3665-ceb6-4f9b-8182-c9f2431e286e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.167028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:3cfbc1d4f16bf74c868c4f36a24f322b94b90ffc635d950690469c54868d818d

Observation 7b6674a3-fecb-4a1a-ac7d-254840df5b89 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.202270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:7f5f61e895e7c2b9fc3a00631e8230a6b5d00e7525f2418194a45bddcc636ca5

Observation 60135306-7518-4a01-a77c-02b859d36201 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena QLoRA: Efficient Finetuning of Quantized LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:29:54.058781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:9ff06e2372b265d20b0e59afa48bb53dfb30537b147a3c569669ccab73d7f7b5

Observation 5e41114d-4857-432e-9fe4-75e731cf2093 · outbound

This paper cites LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.174675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:d2acb8e37340af1b64ef02090929cb5c777fdfa420265fac983371f075d62107

Observation 41c776e8-d9d2-4e35-a3a2-e329d193b2fa · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:52:59.178783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:8bc3bfaf0f049faddce4b954adf603e02f94e52d742619db859a1cd4ebe2ad80

Observation a702fc8b-546c-4aac-a4d7-b38999c4ffe1 · outbound

This paper cites MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.067341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:b7aa89ae09672c68eb1f521781504a8e3ca368cd2eface05ac941d7d72dbf4da

Observation bf0492e9-4aa1-481f-b6ab-0344a36a23cc · outbound

This paper cites Koala: A dialogue model for academic research.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Koala: A dialogue model for academic research

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.215697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:63d2c2958723ea6f742caece428f5de247677208d7904d47f38c2bc263660f65

Observation 882919d0-3361-467d-8025-cc4e9587a3b4 · outbound

This paper cites ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:52:59.071717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:288fa46203afe6e2926d5420badccca17edbb7e467170becfa0452de4c42e137

Observation d19eb9a3-a7a6-4d82-b093-2fd66a707a2c · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena The False Promise of Imitating Proprietary LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:54:31.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:e246b70f13c7ef84dc47a24f15490f6419bfa2970d7f1d1c430003ac3e527344

Observation 92c0b87e-7c8c-425d-adda-f0ffa9fccc7b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Measuring Massive Multitask Language Understanding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.079924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:7cd01f1d2366f1fc8d527824d2eb7fd9b0a8c7a243d92a0074cc7780691f6840

Observation 276d5d61-adce-4dee-a3d2-4613ec4588f7 · outbound

This paper cites Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.083672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:dba308cb5931b0bc261983f6b6d80ac746523d45911fb6da4298d810c6ce9bd0

Observation 2ef3ec40-fe00-42b7-8282-502c13017f18 · outbound

This paper cites Dynabench: Rethinking benchmarking in nlp.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Dynabench: Rethinking benchmarking in nlp

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.233936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:de3fe5d4c5c82214dba5c70e7d5c136eaaeac3ac0cbc070c547916aa69589e2a

Observation de397f39-9f19-4e9f-bcd2-2b16d88d5019 · outbound

This paper cites Look at the First Sentence: Position Bias in Question Answering.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Look at the First Sentence: Position Bias in Question Answering

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.087290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:8316c87cf3380b1f4276f2e3b877d9764ab17dc2399236904bb65a8718df05ec

Observation d0b41416-d58c-4176-9f98-fb9772bc5685 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:52:59.090974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:d1d3bb16a6715ad4f43799c6be5bbffe333c09b09bd7c986ad110962c132bd79

Observation 4c4b2457-6247-403d-a0b9-40662455db54 · outbound

This paper cites Holistic Evaluation of Language Models.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Holistic Evaluation of Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.095035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:22df3ac5c0b1993919ee23d96e506194213aebaa70f3e2a2780c279c2bc7ad62

Observation 56ddae28-18e5-4c78-ad64-cccf5a2c424d · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Rouge: A package for automatic evaluation of summaries

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.189235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:b6feb5a91efa5c57c18312d3d20408f83d4ac9b6d2499e4d7db79f2192f67ebf

Observation 6420962c-9527-41b1-a684-e98a96cd9d4c · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:48:54.871535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:78113e45973fee642cacf3e4fd9cbbf45e0254425085d88b8d1f74c120b03056

Observation 5de40e0e-d688-4fbc-9ca9-36e30c1e56b6 · outbound

This paper cites The Flan Collection: Designing Data and Methods for Effective Instruction Tuning.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.104420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:821ef84aa4cbded4c0a6b34314f56751e54f23348b34e75ee296db2302fc0912

Observation c6f0fdbc-bd59-490a-826f-8bad7e6cd33d · outbound

This paper cites Cross-task general- ization via natural language crowdsourcing instructions.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Cross-task general- ization via natural language crowdsourcing instructions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.199605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:3c8b9f5596a1ad562bbeabd5170899bde94572f548799bbbe8fa52fa062e72d1

Observation 023175b4-86cf-413d-8ee2-5c0fdef23333 · outbound

This paper cites Evals is a framework for evaluating llms and llm systems, and an open-source registry of benchmarks.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evals is a framework for evaluating llms and llm systems, and an open-source registry of benchmarks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.205005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:b3063aea6404c6a324714a49cbe44d1a24c43fb992004d7ccd95829f4fb89339

Observation 1b604ac3-fa49-490e-80bc-e49732fbd078 · outbound

This paper cites Gpt-4 technical report.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Gpt-4 technical report

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.207449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:920edec0811816c3c39e35ade87a2f4d425619827b53ae6dbbe56170bd26229e

Observation de3513db-3987-4583-b27a-f726addd9802 · outbound

This paper cites Training language models to follow instructions with human feedback.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Training language models to follow instructions with human feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.210440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:28245883f572eeedef8a2c523ee090a20e656c51eba5151f30ecb51f0f8f48e0

Observation 1cf3c607-128d-4481-a6ee-b71b6901bd4c · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Bleu: a method for automatic evaluation of machine translation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.212883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:596333a4e0ba075ce4c5f72d917f9e89e62ecf027c341bc3c47f9dfd729a4d65

Observation ec220355-10e0-493f-a74f-774e3da1878d · outbound

This paper cites Instruction Tuning with GPT-4.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Instruction Tuning with GPT-4

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:04:18.193148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:a43561ddfbc97d386869507ac06fda1a3504f3392e6638d5b448083e96cc7907

Observation ef89e6ef-394c-495e-be24-673beebaf236 · outbound

This paper cites Center-of-inattention: Position biases in decision-making.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Center-of-inattention: Position biases in decision-making

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.225077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:15e16165f393c44f50c5bbf7a9633478c3aae1369cb2c209c07c7247885b7db8

Observation e7dd84eb-7ec4-4080-bb57-e394a44036df · outbound

This paper cites Coqa: A conversational question answering challenge.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Coqa: A conversational question answering challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.228327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:f5538372aca9c841f3eff03b82ae7cdf7e3f33480f8aea406514a13f988fd69f

Observation 54fc39ce-0e59-4560-b018-97fb962b959f · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Winogrande: An adversarial winograd schema challenge at scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.231283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:d754089ef72970c9a6091f0e60d56f2d8ec90b7a8b94ecac60b553485c429de0

Observation 7e1adb2e-beef-4bda-b5ac-906e746478fb · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:26:25.883762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:169e4d08a8468f285aacde9004c4497660f8155085ce5dc6fae8c62d1cd8336a

Observation fc722651-daa7-4fec-b78f-d007ae6f3224 · outbound

This paper cites Hashimoto.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Hashimoto

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.239419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:a83d41f33066e0aea4a60a6e833bd73392f5bbb198b0995c4fa115c936165ffd

Observation b379bdfb-6423-4660-acf0-23794153ae71 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.117544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:42753c9bf65ad857686da2eab5e897d04d396d83576a64a09977928cca1135db

Observation 58e12cdb-2880-4458-a142-0261ec8e81d6 · outbound

This paper cites Large Language Models are not Fair Evaluators.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Large Language Models are not Fair Evaluators

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:10:42.624770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:7bd4b8df6043454f6d13117e4c0eafcdb6ed7635001b0c66861a81c90fb3516a

Observation 08fa3ac1-8d8b-4e67-95c4-57a59c7ebdb9 · outbound

This paper cites Position bias estimation for unbiased learning to rank in personal search.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Position bias estimation for unbiased learning to rank in personal search

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.196898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:6145eafedfe45cf5d3958e7ec92d262328abbfd4931d221c469ebf8cda88061d

Observation 99088c72-458b-43cf-9221-c81e0fe11a8c · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.218309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:3bb184ce9059ae009567373456ba0f4031bba57df7a3c3d0e8db83af2f379770

Observation c4054626-5488-4272-9046-d0d7255c4bf2 · outbound

This paper cites How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.124927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:8899b9e5639e095e0b927b7071e18e9a00228a23751871733f89c56ab00757ac

Observation 042d7b39-f694-44d2-9384-80ecd97ac6f6 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.186766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:1e30cd60eeb9d2211505dd4ebc4d8bbb613325004744af62773898de084107fa

Observation a8da3b96-f217-4222-963f-415b8e76bddb · outbound

This paper cites Super-naturalinstructions:generalization via declarative instructions on 1600+ tasks.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Super-naturalinstructions:generalization via declarative instructions on 1600+ tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.191912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:31e8b2b8e459625d9a337b01e713b11fb2da9572100eefdb21a5b7d128617a19

Observation 2f529359-2d9e-40ff-b5f2-39e47f9b7ab0 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Finetuned Language Models Are Zero-Shot Learners

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:14:13.786069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:8db7e5a7beba750412c801693b1fb56c44ed18d5021650b8bb1d30c713dbc42c

Observation 6685e195-3cf2-4ad8-b95a-195380b67c1e · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:52:59.131510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:2de4de0f8c9e518842f9b4bc71c052498cd805c0de169260736ec81172eef924

Observation 0ae0f4c0-46e2-422c-96b8-525f61e08af9 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:28:25.169141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:69e9cab386019e6a103cb2705c30ab104657b011ddaff87aa10730d246663064

Observation ac1331a5-b308-4285-9989-f21ed41ea6ae · outbound

This paper cites SkyPilot: An intercloud broker for sky computing.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena SkyPilot: An intercloud broker for sky computing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:52:59.236715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:20b4d4f24277760396806d84c25e19df61f8597723faa82fed657d18a3bc3cf3

Observation 26a16eed-6774-4b15-80bd-9a99e39b7af6 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:56:24.661063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:3d048365d5ef7bdc5553148749e582cd75ffa43011bc2014c121cf06dad3126a

Observation c1b92dad-805a-47d2-943c-dc989a98b5ff · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:60cd944db62e11263a8bb9a8ae2cf5f7e42b1efa03afb2ffd11e417a2c91bbfc

Observation fb658417-ddde-4e1e-85b2-65ebdd99d0de · outbound

This paper cites LIMA: Less Is More for Alignment.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LIMA: Less Is More for Alignment

Reference 52

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T11:34:13.088019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:dce366b9dc7bc894dc2ef0ce534119e59340dd52f67624768052419884b66111

Pith citing papers

Observation f72bf196-212b-4c85-b3a9-b24f04886e8f · inbound

WizardLM: Empowering large pre-trained language models to follow complex instructions cites this paper.

WizardLM: Empowering large pre-trained language models to follow complex instructions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:28:25.016549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T07:28:24.827546Z digest=sha256:40db256cc9d1450335f7f43ef5a4c8eb5408e5d533ec63f4ea040f08bc4ccf2a

Observation 1367731e-1131-4f97-ba09-7c5979034c87 · inbound

Large Language Models are not Fair Evaluators cites this paper.

Large Language Models are not Fair Evaluators Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:10:42.345226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T12:10:42.248005Z digest=sha256:74a78354f07046071996f84b209738ca09d6a2f15415e5f90f777257a7281c7a

Observation cebe60e5-147d-4845-8fa1-b8dd4a3d0d07 · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.450809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:5e801b8a9fd9159071d3d40ba137f84864bb977cc456c2de89520901097ed9a0

Observation f90ff1ce-b535-486d-8492-d0cfeb3e2748 · inbound

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate cites this paper.

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:03:18.789065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T13:03:18.765496Z digest=sha256:28138ec83b4c5c588e95f7356ccd1eadcac9ef9848f00fb2d56ba0c651420cfd

Observation 052242c2-f41a-4fe0-bc60-810a6b579599 · inbound

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding cites this paper.

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 132

Resolution
verified exact
local_arxiv, observed 2026-05-12T20:22:10.678903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T20:22:10.482509Z digest=sha256:2eda12d36e91a2ba274928fbdd3610a0a7ae8e5e58f6621675d07f727afe2a9f

Observation fc8959be-4c8c-47f3-a70c-3a5cda64fcdb · inbound

Textbooks Are All You Need II: phi-1.5 technical report cites this paper.

Textbooks Are All You Need II: phi-1.5 technical report Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:18:01.272005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T19:18:01.244364Z digest=sha256:1982021ac2eff0c105058ec75f4f6cd259fbaf0560eefd844e97183ce50361b7

Observation f0c63962-c1af-4c2e-bb0e-710a0372b3d7 · inbound

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning cites this paper.

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:46:39.656638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T23:46:39.330438Z digest=sha256:5bf56ea7d15345c141305be7c15de7e4a27a7e94d85759815071e9c56c278df0

Observation 3130017a-da2d-49d2-bf03-33ddd3fade96 · inbound

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts cites this paper.

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:25:21.165188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T06:25:20.966510Z digest=sha256:5bb434964959cf6a2e0b31600a04e2d95090cdee6b2fe32b0d0eaf259ac670da

Observation 5e2d0391-945c-4c1d-a43e-124be978a2cb · inbound

Studying Lobby Influence in the European Parliament cites this paper.

Studying Lobby Influence in the European Parliament Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:44:02.430659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-24T06:40:16.620051Z digest=sha256:268c229ed72d1373c9fae9546e25e0b805a551c846d286720de76777cf8b8519

Observation d6224d4d-b0d8-40d9-9432-72ac74903836 · inbound

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models cites this paper.

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:46:52.841020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T22:46:52.791128Z digest=sha256:c4ffb10a4a356eaec44bcbd2b944348f3fc4932d833b1ca271b4d85afacc46cf

Observation dd725577-172a-4ad1-8c24-ae36a7cc700c · inbound

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! cites this paper.

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:58:35.770758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T08:58:35.714394Z digest=sha256:7ba40c6f934b51a3ba85393c95cf6875070f754b4e22dba8577499f67e835aa6

Observation fc576694-5a01-4e56-becc-416091abdec1 · inbound

MemGPT: Towards LLMs as Operating Systems cites this paper.

MemGPT: Towards LLMs as Operating Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.240297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T12:27:29.041352Z digest=sha256:133b26c4d6b2b8cc503a8f82bcbdd91482f9a8cd858117c8e7fef62192178924

Observation cdf31227-10da-489d-9165-4cfaf1e39177 · inbound

Zephyr: Direct Distillation of LM Alignment cites this paper.

Zephyr: Direct Distillation of LM Alignment Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:13:57.426351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T10:13:57.361932Z digest=sha256:85c6751139685cf9c093a4fc5aaf12c20e4aba08197718f205bdbd7752a353a1

Observation b12c085f-27ff-4366-8209-095acbc4c3e2 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 203

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:46:09.880853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:fe0f2fdbd40bc98020e286ae79b9cbd68285fe0710f7daaa3b498fbc62a3e5e4

Observation 088c9e89-3ce7-410e-8745-03961d61da6b · inbound

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations cites this paper.

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:34:15.895133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T22:34:15.638114Z digest=sha256:bbfc84ffe5909626f26c092bd4713d0fffe31fdf8173a04c1c57d67b006f514e

Observation 196a7187-8552-4dc3-bc31-ac0d35994c4b · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T10:16:43.891325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:3d3833e422e665a57f18887afb820df251d191e454d1305bb9192beb6f803a25

Observation 0c706940-5785-426b-8f50-4b8b46b1b738 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 185

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:46:10.270145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:ee06c4ab8b8f8a49bbe091ee58fd71890ab4f5924967135d240c93ba1afbd0a8

Observation 27a7750c-c8d6-48be-99e3-9f2371b8a217 · inbound

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices cites this paper.

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 133

Resolution
verified exact
local_arxiv, observed 2026-05-16T16:35:38.191524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:35:37.937462Z digest=sha256:3f1edbf5d2b337dc54f27405f85598c9bec393fa4bd426177053a33cb0d56f1c

Observation f2032401-9541-4524-aaa3-25603ee12505 · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:13:53.885898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:c2aa1e3f993d6d010d0bf1f4ae1dfea867384f84325c46199889796ef7479e80

Observation 39d1a7cb-912f-439d-9acd-bd67ad725d2d · inbound

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty cites this paper.

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:15:49.438354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T00:15:49.303458Z digest=sha256:3a93ceee3ed8599f164f5c748b9c6fd9e2c4e9186f1b799206fcab21f989d09f

Observation 0c8235e4-cde4-4e2e-9794-3495b8e7a912 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T12:17:53.549452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:3589f5bad9b15f61418252889d8c93686786e6b332a0f50eeb192b1a12769823

Observation feda9130-605b-45b9-9ae9-89e7d424b270 · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:36:57.248655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:e83ff759e5a21ac7fa21fe8a70f5093003b8de69c0a155b48746b8441af40fbe

Observation d468840c-4596-427c-ad01-6ce08cc3c3bc · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 99

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:04:44.457086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:c70a019e451a4987b2cf8e7041e00066568a7888df03f83ac072c6605f8cdf01

Observation 0326ac5c-d92f-4a0b-aafa-749f459f01c5 · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 137

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:34:04.516226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:7c5abf48a81756fbaaccaf54b3aa9e3fc9ac7f03562fc569fff1dd85cde98758

Observation 528933f3-c0ec-4b08-849a-5ea1f4c09bc4 · inbound

RouterBench: A Benchmark for Multi-LLM Routing System cites this paper.

RouterBench: A Benchmark for Multi-LLM Routing System Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 122

Resolution
verified exact
local_arxiv, observed 2026-05-16T10:47:31.106113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T10:47:31.006944Z digest=sha256:00022f7edc7972ea346682cf946c90662d02dd6d729b69ff562d2a1b2912e3ec

Observation cd42278d-1978-49f0-81df-f7c053c8745c · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:08:05.546872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:410e44936d5b8ff4cc662e52eaf9ed0ace31399f1ebcbe0056aef0f57093dcb1

Observation 95b35164-7faa-4f42-a8a7-fac2463f6dc5 · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-10T20:19:27.284354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:155efaccb015365206077bc5bbf092e176176ef3da7e8fa7cc7e358ae3696bad

Observation 71f63490-e1d6-4585-b818-c0693bd39714 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 92

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T18:44:49.614200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:c09ad15b1039ebbe4f4a8e4a5ad73da8963fadd34d3840d152181ac5f01b772b

Observation 3cb9f199-ad53-4cca-ae65-d2774ebd1ba3 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 273

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:37.117555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:ad0a9137ab395de7a135c374042368da425c09e99463b15e32daa24f61861308

Observation 737e8445-6ce5-4896-920a-334a477100e6 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T21:23:27.447863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:9976d145e6d364583c29115887277a7286aafb33748ded5cfa15c92fe94aabdd

Observation 8e5d07f4-d36f-483e-9df5-acce550ee1ba · inbound

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback cites this paper.

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T22:38:32.597717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T22:37:43.230753Z digest=sha256:7319a9cf6c36a2738f66a1821870aeb5db2bba708c0285a8b100695b871f9127

Observation 13dd085b-36e2-49f1-b5c6-9e8b5f9fdc4b · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.574522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:e358a79e05cd17526dc1609f667c7985ddb8664ed093128cc01a5b67c99a6aa9

Observation 9315eda1-2d6a-4f23-9543-8901b8611b92 · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:42:58.153412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:8a320319b51ec5ec4f986bffb35b51bf0e5ff18454414077542e87b020e7a06a

Observation 43e780d3-cafa-43ff-98be-628f5a5392bd · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:36:58.774117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:f15037ed8e80b920a448133f9a99b44138c64dbea3c1392723d363ad33e502ae

Observation 240eb6cb-23e8-4411-b9be-1ccc02f26069 · inbound

TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models cites this paper.

TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T19:01:57.820004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T19:01:42.307514Z digest=sha256:e5d4d12d8a519206308e816ca6b6d718eab9355d8832ce9dcc599a15fdd396d8

Observation 82c6491b-073d-48cd-9408-004033b0667d · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 178

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:26:05.911342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ec8dfa0edcd0226d94301f31f25da9eff00a8f2a8316d363e132b0d36f32cbab

Observation db66ab1b-95a1-4555-9462-5ab75a143a43 · inbound

Tuning Language Models for Robust Prediction of Diverse User Behaviors cites this paper.

Tuning Language Models for Robust Prediction of Diverse User Behaviors Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T13:52:19.790722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:51:21.147668Z digest=sha256:5d9573f8fb41cb6a147acf2034b9e3deeaafc320cc7dd5ac51aee29fcbb0420f

Observation c52f11c7-9882-45bf-87f5-182aa1952a75 · inbound

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots cites this paper.

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:32:19.305119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:29:34.151546Z digest=sha256:ff36ec5016a286bc1b6067b7c815bb47dc283ac6daa94b22129ba0c3fafabb37

Observation feb105f9-173b-4e0e-9f53-991e3a26dde0 · inbound

Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation cites this paper.

Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T12:57:17.804573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:55:31.950743Z digest=sha256:dc973d3090695db33282c6771faf2845213b36a8ff5c33dcd4adcda39ebe7208

Observation 1041e806-0c87-4455-a0d6-e4349fdc1a2f · inbound

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration cites this paper.

Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T12:52:18.035155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:48:44.324236Z digest=sha256:96482ad488fbf7e9bc0b57cecbfa7d21f5c9388520e3775476f9b6c3e3415383

Observation 464ecab5-6583-4867-a904-ba5f6ed32fac · inbound

LLMs Judging LLMs: A Simplex Perspective cites this paper.

LLMs Judging LLMs: A Simplex Perspective Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T12:42:18.514994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T12:40:37.816449Z digest=sha256:c72db515731dc242f34e8ce42b3a6567a54226f7a8a05f4524f9a3d1cbd15599

Observation 9c4b1a50-88ea-4461-944b-2f7ed61a1caa · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:07:39.536246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:d39581ee9b4bad5a5a89d19a2077cc7010f42e703e22e64a1e1db5bc23c4c70f

Observation 0f794dd2-4430-4638-b7e0-a2e529df5015 · inbound

Listener-Rewarded Thinking in VLMs for Image Preferences cites this paper.

Listener-Rewarded Thinking in VLMs for Image Preferences Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:42:09.208655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T07:38:52.273903Z digest=sha256:f167e6c706a9b732ed6cc094ec52ea690fd3efdb2494103a8a919d6f2846dd6b

Observation 314c7ca6-8ba2-4e65-85ea-e43f7ebea27e · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:52:07.707421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:4b35d19bd74043f117390946b8d44a91d483877f0d2ff6a208d85657806688b9

Observation 4b4852fc-c0e4-496a-bfd9-8c4fe01178ae · inbound

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors? cites this paper.

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T22:32:52.366733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T22:32:11.401201Z digest=sha256:a43b514614d0c0db2b48ff1136f5a57aa93ebf3b73dd841bd8ea80ef2064c5fa

Observation 929cbaa5-cc58-4f21-b1da-598553038c49 · inbound

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models cites this paper.

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T22:41:53.055375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T22:41:26.047957Z digest=sha256:ed88c7a752d5b61938b9008663a178cef3af0e56d445e1a55e59153cd721dd9b

Observation ab01ebc0-fa6f-48e0-b1f0-92babacb6e70 · inbound

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling cites this paper.

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:41:51.709441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T21:39:02.560962Z digest=sha256:0b9969faebf6e8d9007268c5f083e19289859b5953368b1207f393f215260b7d

Observation 9bf28a27-842f-47d5-be4f-26a3a1b8ddb3 · inbound

Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation cites this paper.

Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T05:34:40.410309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:34:40.410309Z digest=sha256:df7d82edaa5c9815fbb9914a9d43b98a2efed265dcae76dab0f552190c0e0705

Observation e8de947b-8b25-4cb3-871d-ecce4390806c · inbound

Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks? cites this paper.

Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:49:34.834441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:49:34.834441Z digest=sha256:c798f010c6d3b3d8ac0cfb62726c4bb6449ad575e9d9a177748cb241c818846d

Observation edc29ab5-a9a5-4367-a3c0-04ee523c00e2 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.423390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.423390Z digest=sha256:197bf9985bbb9be830b082055520ace6dca91bb94681054b846ff7820ea67b53

Observation 005b982f-aff8-4b26-92a0-123ac3976a19 · inbound

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector cites this paper.

Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T17:42:46.082416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T17:39:17.456350Z digest=sha256:b8267fb67b89e55e5de86d456aa8b704f0075f82221323cfe5a4f263fc4fc510

Observation 963681bd-0677-4864-9bc4-8c316c747d66 · inbound

PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability cites this paper.

PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T20:04:29.442012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:04:29.442012Z digest=sha256:7dd90457da8dbb7dbaa47c310c2639f21555e69900dbd8cbbb836d505e9efb57

Observation 1d55773f-5807-4f6a-b8e0-4bb875b9df02 · inbound

How Small Transformation Expose the Weakness of Semantic Similarity Measures cites this paper.

How Small Transformation Expose the Weakness of Semantic Similarity Measures Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T23:35:05.545669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:35:05.545669Z digest=sha256:cf71d70c2f6e9d6588e1d4bc9544e25ce6f938eb82e2f9829e12c0844f460589

Observation 6b02f4cf-8c24-475a-8010-0af7e54faa1b · inbound

Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization cites this paper.

Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T18:37:55.129913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:37:55.129913Z digest=sha256:c5795ca2a3ca90695b29fad9d6e5f6d3f500bc9324d75ca9a6afe7c39df9eef5

Observation 9a58348b-4a58-4cc3-ad89-25bfcc4479f0 · inbound

Rethinking Human Preference Evaluation of LLM Rationales cites this paper.

Rethinking Human Preference Evaluation of LLM Rationales Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:27.898978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:27.898978Z digest=sha256:ac33bf3a9ac6cbf7d1f0ed00e0e38d6de5f7ba22b29d3ea1d5304c25dd0fea9a

Observation 228824bd-89a2-410c-b3e9-c9a035ef7786 · inbound

Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability cites this paper.

Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:42.817913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:42.817913Z digest=sha256:8c1335357062043ecd44e3c1738e71b15d281d1ef1c8e8a9181470a523052fb4

Observation 012c0adc-9f11-47b1-93ed-1f892e9ce70a · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T17:01:39.969925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:3fc9776b69faaaffefa60a792fd4dc5025517f1f484ec5b9c733d207d2494ec7

Observation 009d1fb1-6d35-46d2-92bf-eef58b50ac02 · inbound

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning cites this paper.

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:56.345107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:56.345107Z digest=sha256:4aab5176ded87fb178fde8d1d01625246ce64130e82d9f700f34204f2b3b3ae8

Observation fb320e90-b39c-420a-97c4-507f3a07e8b1 · inbound

Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning cites this paper.

Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T07:45:28.889091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-25T07:44:05.582003Z digest=sha256:2302052a2bda4af8e8442d9c3254f79af216f24d5dc87ab2860f45015703540a

Observation 5c8ed9b6-98ce-49dd-a0d8-ef30cdfba3ce · inbound

CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling cites this paper.

CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:26:50.860897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:26:50.860897Z digest=sha256:261fed1fae2f30cc70910ff3705fb83a85abf86dbc3795ae976dfef90c7122b0

Observation e6c14c3a-3b70-40cb-95c9-78c013e4aab9 · inbound

QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild cites this paper.

QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T11:01:17.190998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:56:28.973065Z digest=sha256:59b3f41b7b5492bc957eed4f04da80a818a6ab70739944e7995f26e5d593abfc

Observation b7456044-d828-481a-ab2c-7a238b0c30d3 · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:06:13.702902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:07a68c66f63ae317731bd19a064d372d4785d50ca70e404e651badd1d97cb7f6

Observation fe30e200-780e-4307-b31f-fa45eebefbc3 · inbound

Automated Alignment between Elicitation Interviews and Requirements cites this paper.

Automated Alignment between Elicitation Interviews and Requirements Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T11:09:41.074310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:09:41.074310Z digest=sha256:40d34429b498c7c1f1145477c260e86743e227367fa59aee6f2b674ca78aa569

Observation adc5df69-5d9b-4765-8e74-6c377448726c · inbound

Aligning Deep Implicit Preferences by Learning to Reason Defensively cites this paper.

Aligning Deep Implicit Preferences by Learning to Reason Defensively Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:11:06.730794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T08:11:03.152989Z digest=sha256:ccdbe3da619345b7f9812b51d59e1cbc919ad39178b9f527026a6d74a143e0e8

Observation c8ae59d6-fcb6-49d6-ade9-2f0e63befebf · inbound

Aligning Deep Implicit Preferences by Learning to Reason Defensively cites this paper.

Aligning Deep Implicit Preferences by Learning to Reason Defensively Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T10:15:53.984725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:15:53.984725Z digest=sha256:dbfa57351a82604f5312440c203155ed85486f1b619f1bd67a5e7b8aea90bb7a

Observation 8908a78f-6480-4ed7-8c6a-940a89cf5e6c · inbound

LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization cites this paper.

LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T07:02:26.585565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T07:01:43.510924Z digest=sha256:a694d6377807d4f778ef437924bb3f105cf67137d58dc035b28e40ef955b379e

Observation 1e412267-22f4-4a34-ae49-cb608c54c2aa · inbound

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL cites this paper.

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:32:17.346126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T01:31:40.920567Z digest=sha256:cffe71938d4f70c0abb2ae3922d24de2975aac953eb6a55cc17962a26b774f6e

Observation 40d205e2-e6c5-4a37-b3a2-f9ed15946c10 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:55:38.035924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:1f825aac72cab715e48ecf7760ad291da1f826da03694252d59d6d97a2704465

Observation b4391a83-ad5a-4391-8bc0-acb6d360c5b2 · inbound

Reading Between the Lines: The One-Sided Conversation Problem cites this paper.

Reading Between the Lines: The One-Sided Conversation Problem Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:40:33.655532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:39:22.598660Z digest=sha256:d86e65beb25718ed7ca152fde8b4332faaf8e66d7adb5d966ba4e1fd4e9c634e

Observation 4d09e5d7-c31a-411a-aced-dedf4d0eb6bb · inbound

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models cites this paper.

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:25:28.447604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T23:24:40.372674Z digest=sha256:3f3feaf20f8d94f2db651260be974c7e9b6f1dfd534721551898870dbb5dfd58

Observation 1ec81eda-673f-4822-9042-aa18cec22a17 · inbound

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models cites this paper.

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:25:28.521910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T23:24:40.372674Z digest=sha256:ab81e137affd8f0cf2e7e3c70fc0879dbc94da79db9d04628d09051a835d35aa

Observation 1469f486-4021-4d8b-be53-a1ccd05ff185 · inbound

Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go cites this paper.

Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T22:23:30.982151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:23:30.982151Z digest=sha256:32a2f07c10a433ddf4ba50fadb313335240f4a876dfd90219d878d09ec5eed77

Observation 497b1f07-4f75-4ba9-b1b0-8afabefe53b2 · inbound

Pessimistic Verification for Open Ended Math Questions cites this paper.

Pessimistic Verification for Open Ended Math Questions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T20:01:29.973988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:01:29.973988Z digest=sha256:5a62df75b1db763949fcc36584cc626c1d12b9285a222d50eaa11bfccb5c97b8

Observation c00e22b8-cfd9-4563-be13-59d5f11344b9 · inbound

PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models cites this paper.

PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T02:38:53.851741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T02:38:11.118057Z digest=sha256:067cc601bf0cef7fdeb54d271e323f1d187b957412e3c699089e4bfe926da548

Observation 32d189e4-7b3f-46f3-b09b-1801d253ebbf · inbound

InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation cites this paper.

InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T17:15:48.221901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:15:48.221901Z digest=sha256:d73f3ea841f38b32c5301aada7080c74f2da1c07eb9124abc2a26a42d8967671

Observation 5bd1853e-0f8b-4e3f-979b-c5a6cc8abc97 · inbound

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models cites this paper.

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:38:34.182199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T21:36:24.376401Z digest=sha256:0334d4ad905acd20a40fdab633d22af80dc0b2b1d44986148ae3fd0c54f61fcd

Observation d67378b5-9685-4d9d-995c-1501f041aa9e · inbound

ClinicalReTrial: Clinical Trial Redesign with Self-Evolving Agents cites this paper.

ClinicalReTrial: Clinical Trial Redesign with Self-Evolving Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T18:11:09.389053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T18:09:30.291505Z digest=sha256:ccb8cf73d956812d1ddf49fcd472ed0f9c5f1b71be680046a55ae0261411b5d5

Observation 2cefaacd-a7aa-457a-a0a3-e5ec2a3270a2 · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:37:59.991191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:54db68fccef3e9c6df4f76bdb58ed2562a60926ed0c2613787c523b792fe7361

Observation 93e96a40-a3d7-4f6e-ad10-243df9e1d880 · inbound

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems cites this paper.

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 173

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:53.746742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:47:28.248540Z digest=sha256:02593eee7b2748d23ac18837064715889ddb14657a681b3d83b861fc39da64fb

Observation d60f9d61-758e-48e0-9c71-f23dc7a8c5d1 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 277

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:37.982801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:37.982801Z digest=sha256:bbf3f3b6968e54a94eee343a999e752d410378c4da319de2a602d02d09396b9b

Observation e028793b-e67c-49a0-aeb8-c55dc563c994 · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.366420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.366420Z digest=sha256:8f4a9d63c6ca1bcb748ec61528ff20a8f55a38de83609e5b906cc04e5ed61877

Observation c33a4aee-cdf4-4ea5-9312-67612fb98833 · inbound

StepShield: When, Not Whether to Intervene on Rogue Agents cites this paper.

StepShield: When, Not Whether to Intervene on Rogue Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:33.095817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:33.095817Z digest=sha256:c77ce8da63d5dbf4bb66c1df558a1f3c7807f2947711eeb1546ba1e7c44256d0

Observation b482c4bc-182e-4f88-ac1d-f7eb0691517b · inbound

A clinically validated framework for auditing AI chatbot behavior in mental health interactions cites this paper.

A clinically validated framework for auditing AI chatbot behavior in mental health interactions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T06:18:36.946972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:18:36.946972Z digest=sha256:99001a0e680903d7bb8170e8426417f8a276c5fa150b6f470263afb526e62d98

Observation b2e287a0-cf16-40a7-a001-24d90415dbda · inbound

SAGE: Scalable AI Governance & Evaluation cites this paper.

SAGE: Scalable AI Governance & Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T03:31:57.159932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:31:57.159932Z digest=sha256:7339d43a81a30075bbff10bc43af70102733d32ce34aee738753cc11ccd8cad6

Observation bdcf0897-057f-4d0e-8a6c-358ac17044e2 · inbound

Bayesian Preference Learning for Test-Time Steerable Reward Models cites this paper.

Bayesian Preference Learning for Test-Time Steerable Reward Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:20:10.373530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:1f5746f1f7f278177d32ef89b5e2641fd62aa3a2117943bb485e64c9cc452e56

Observation fc65662b-f03e-46ea-a57f-27068119551c · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:28:48.779663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:1235db2a36102b18c1176a61407b4137f566d26d43677bfec4ab4c3e3b6b3d79

Observation 5999a959-c4eb-4ff2-9082-9569ed568d9e · inbound

Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments cites this paper.

Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T22:43:03.629247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:43:03.629247Z digest=sha256:397291e6e17d0e45d8ca4643788e3cb3c64ede63b926bf06788e0089d1c75f21

Observation 367dc6c0-0192-4d2e-983b-0a93212f9910 · inbound

Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments cites this paper.

Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T20:27:49.650865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:27:49.650865Z digest=sha256:d6d685b4ccc5580c9136be0a40bf837f1f1347fd4ee26131647355b32fcab663

Observation ed96ed93-6686-4047-ad8d-e2295fafd973 · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T06:45:25.636294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:0a0e29db72d87785456f2b2bb3071ea1c26f8d85039f9edbd7b62ff7d54262f2

Observation 5064c39e-4149-440d-ae01-72fc7c1b281c · inbound

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use cites this paper.

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T05:54:53.441547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:54:53.441547Z digest=sha256:b2910923170a2bf65951f82894d43bddc9dea94097cb77318b254aef9e2cff3c

Observation 3aaee713-1b19-4cdd-bdad-d1e9b66717db · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:54.419502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:54.419502Z digest=sha256:b96535db84eef08d63d43839998fbd5a3b587d190174ce454956fe6be9961000

Observation d889f58f-e4d3-4d8f-a0e0-0e0f5390f830 · inbound

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models cites this paper.

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:58.971096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:58.971096Z digest=sha256:b0909e5e3ff2df32c9799a3cd3787fd8a6ed31e134b089238d5104226a744fdd

Observation 58a92dfe-273a-48f3-b703-0bea38e7b0ab · inbound

Scalable and Personalized Oral Assessments Using Voice AI cites this paper.

Scalable and Personalized Oral Assessments Using Voice AI Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T10:09:59.559097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T10:05:30.015226Z digest=sha256:cd4ebc281d18990e705ea8abfe3c11c185ef8b38ceef0adc79ad907ff63c479f

Observation d0a9f836-ec31-4153-9cbc-329abce6a5d1 · inbound

Scalable and Personalized Oral Assessments Using Voice AI cites this paper.

Scalable and Personalized Oral Assessments Using Voice AI Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:00:41.587603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:00:41.587603Z digest=sha256:223deaac90f2a155d6084258727e6f733262d0a353947bbf3873774513e53e6d

Observation 17103961-9333-406d-bad5-4f3a69ef99b8 · inbound

Synthetic Data Generation for Training Diversified Commonsense Reasoning Models cites this paper.

Synthetic Data Generation for Training Diversified Commonsense Reasoning Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T08:15:16.934625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:11:22.897371Z digest=sha256:79ac301dfa9a978c3005190068a873c7679648eb6412a54180999b89312414ea

Observation eb9902cf-92fd-48da-bec6-939bd008ded6 · inbound

Agentic Business Process Management: A Research Manifesto cites this paper.

Agentic Business Process Management: A Research Manifesto Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T08:25:18.154337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T08:24:56.763438Z digest=sha256:b9b6cd3f5c9e6a6907396d62fa4c53c493110af10865759c39b4a8fdc76fe1e9

Observation fbef79b8-acef-4017-8f48-719f9938cd32 · inbound

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines cites this paper.

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T17:52:54.820035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:52:54.820035Z digest=sha256:049745d0c1645014cfb5f07ccaa0abecccc5d7bd18f3b4ff3ac879ce21e80102

Observation bdfec1b5-24c0-4f2e-8e50-81691499ed1c · inbound

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs cites this paper.

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T05:42:17.786502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:42:17.786502Z digest=sha256:8071858211a7044efbd85bbc2b5872bb6201cbc6a07a83fc3e9adcaa1309240c

Observation 8474c254-e454-4893-aec2-e0185d31bd06 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 133

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T21:03:20.042338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:c588859877633681bd5fb164a85ec3fd4cadebb59ceade0441ecbf674d2fefe5

Observation 31db940e-5b6c-4a30-95d8-e0713fadf8a6 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 133

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T10:40:00.850281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:7fb4ac06c23abbb435e46816eeca0a6de92ba562861fd6647b12b983c74dac54