Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:52:59.033645Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 100 inbound Pith citation observations for arXiv:2306.05685.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:52:59.033645Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T05:34:40.410309Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
52 of 52 outbound references displayed
External citation measurements
8605
pith, observed 2026-08-05T02:28:24.338817Z
Observation c3f9150a-0bf5-48f7-88b3-e3e822ae231e · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena PaLM 2 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7c33d917-af12-4745-be3b-a0054a21a150 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f56df5af-05f3-4af8-bf15-aeff97581859 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Position bias in multiple-choice questions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a87b7471-459c-422a-9790-db258d8c2609 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evaluations of self and others: Self-enhancement biases in social judgments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09baee0a-afe2-4161-8531-0431dc543526 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4505964-88d7-4bcd-8a49-358f75ddea94 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 814939cf-b7f6-4969-91c5-9e1bec2f57ff · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0d93a60b-5759-4767-8d3a-d98958c459d9 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ab99aab-cafc-481c-bd8d-885eafaa8def · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9f2f3665-ceb6-4f9b-8182-c9f2431e286e · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b6674a3-fecb-4a1a-ac7d-254840df5b89 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Flashattention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems, 35:16344–16359
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 60135306-7518-4a01-a77c-02b859d36201 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena QLoRA: Efficient Finetuning of Quantized LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e41114d-4857-432e-9fe4-75e731cf2093 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41c776e8-d9d2-4e35-a3a2-e329d193b2fa · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a702fc8b-546c-4aac-a4d7-b38999c4ffe1 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bf0492e9-4aa1-481f-b6ab-0344a36a23cc · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Koala: A dialogue model for academic research
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 882919d0-3361-467d-8025-cc4e9587a3b4 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d19eb9a3-a7a6-4d82-b093-2fd66a707a2c · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena The False Promise of Imitating Proprietary LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 92c0b87e-7c8c-425d-adda-f0ffa9fccc7b · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Measuring Massive Multitask Language Understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 276d5d61-adce-4dee-a3d2-4613ec4588f7 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Is ChatGPT better than Human Annotators? Potential and Limitations of ChatGPT in Explaining Implicit Hate Speech
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ef3ec40-fe00-42b7-8282-502c13017f18 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Dynabench: Rethinking benchmarking in nlp
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de397f39-9f19-4e9f-bcd2-2b16d88d5019 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Look at the First Sentence: Position Bias in Question Answering
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0b41416-d58c-4176-9f98-fb9772bc5685 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena OpenAssistant Conversations -- Democratizing Large Language Model Alignment
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c4b2457-6247-403d-a0b9-40662455db54 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Holistic Evaluation of Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 56ddae28-18e5-4c78-ad64-cccf5a2c424d · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Rouge: A package for automatic evaluation of summaries
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6420962c-9527-41b1-a684-e98a96cd9d4c · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5de40e0e-d688-4fbc-9ca9-36e30c1e56b6 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c6f0fdbc-bd59-490a-826f-8bad7e6cd33d · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Cross-task general- ization via natural language crowdsourcing instructions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 023175b4-86cf-413d-8ee2-5c0fdef23333 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Evals is a framework for evaluating llms and llm systems, and an open-source registry of benchmarks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1b604ac3-fa49-490e-80bc-e49732fbd078 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Gpt-4 technical report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de3513db-3987-4583-b27a-f726addd9802 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Training language models to follow instructions with human feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1cf3c607-128d-4481-a6ee-b71b6901bd4c · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Bleu: a method for automatic evaluation of machine translation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec220355-10e0-493f-a74f-774e3da1878d · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Instruction Tuning with GPT-4
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef89e6ef-394c-495e-be24-673beebaf236 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Center-of-inattention: Position biases in decision-making
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7dd84eb-7ec4-4080-bb57-e394a44036df · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Coqa: A conversational question answering challenge
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54fc39ce-0e59-4560-b018-97fb962b959f · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Winogrande: An adversarial winograd schema challenge at scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e1adb2e-beef-4bda-b5ac-906e746478fb · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc722651-daa7-4fec-b78f-d007ae6f3224 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Hashimoto
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b379bdfb-6423-4660-acf0-23794153ae71 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LLaMA: Open and Efficient Foundation Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58e12cdb-2880-4458-a142-0261ec8e81d6 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Large Language Models are not Fair Evaluators
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 08fa3ac1-8d8b-4e67-95c4-57a59c7ebdb9 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Position bias estimation for unbiased learning to rank in personal search
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 99088c72-458b-43cf-9221-c81e0fe11a8c · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4054626-5488-4272-9046-d0d7255c4bf2 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 042d7b39-f694-44d2-9384-80ecd97ac6f6 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8da3b96-f217-4222-963f-415b8e76bddb · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Super-naturalinstructions:generalization via declarative instructions on 1600+ tasks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f529359-2d9e-40ff-b5f2-39e47f9b7ab0 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Finetuned Language Models Are Zero-Shot Learners
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6685e195-3cf2-4ad8-b95a-195380b67c1e · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0ae0f4c0-46e2-422c-96b8-525f61e08af9 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ac1331a5-b308-4285-9989-f21ed41ea6ae · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena SkyPilot: An intercloud broker for sky computing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 26a16eed-6774-4b15-80bd-9a99e39b7af6 · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c1b92dad-805a-47d2-943c-dc989a98b5ff · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb658417-ddde-4e1e-85b2-65ebdd99d0de · outbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena LIMA: Less Is More for Alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f72bf196-212b-4c85-b3a9-b24f04886e8f · inbound
WizardLM: Empowering large pre-trained language models to follow complex instructions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1367731e-1131-4f97-ba09-7c5979034c87 · inbound
Large Language Models are not Fair Evaluators Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cebe60e5-147d-4845-8fa1-b8dd4a3d0d07 · inbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f90ff1ce-b535-486d-8492-d0cfeb3e2748 · inbound
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 052242c2-f41a-4fe0-bc60-810a6b579599 · inbound
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc8959be-4c8c-47f3-a70c-3a5cda64fcdb · inbound
Textbooks Are All You Need II: phi-1.5 technical report Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0c63962-c1af-4c2e-bb0e-710a0372b3d7 · inbound
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3130017a-da2d-49d2-bf03-33ddd3fade96 · inbound
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e2d0391-945c-4c1d-a43e-124be978a2cb · inbound
Studying Lobby Influence in the European Parliament Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d6224d4d-b0d8-40d9-9432-72ac74903836 · inbound
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd725577-172a-4ad1-8c24-ae36a7cc700c · inbound
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc576694-5a01-4e56-becc-416091abdec1 · inbound
MemGPT: Towards LLMs as Operating Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cdf31227-10da-489d-9165-4cfaf1e39177 · inbound
Zephyr: Direct Distillation of LM Alignment Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b12c085f-27ff-4366-8209-095acbc4c3e2 · inbound
The Falcon Series of Open Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 088c9e89-3ce7-410e-8745-03961d61da6b · inbound
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 196a7187-8552-4dc3-bc31-ac0d35994c4b · inbound
AppAgent: Multimodal Agents as Smartphone Users Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c706940-5785-426b-8f50-4b8b46b1b738 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 27a7750c-c8d6-48be-99e3-9f2371b8a217 · inbound
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f2032401-9541-4524-aaa3-25603ee12505 · inbound
Mixtral of Experts Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 39d1a7cb-912f-439d-9acd-bd67ad725d2d · inbound
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c8235e4-cde4-4e2e-9794-3495b8e7a912 · inbound
KTO: Model Alignment as Prospect Theoretic Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation feda9130-605b-45b9-9ae9-89e7d424b270 · inbound
World Model on Million-Length Video And Language With Blockwise RingAttention Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d468840c-4596-427c-ad01-6ce08cc3c3bc · inbound
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0326ac5c-d92f-4a0b-aafa-749f459f01c5 · inbound
ORPO: Monolithic Preference Optimization without Reference Model Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 528933f3-c0ec-4b08-849a-5ea1f4c09bc4 · inbound
RouterBench: A Benchmark for Multi-LLM Routing System Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd42278d-1978-49f0-81df-f7c053c8745c · inbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95b35164-7faa-4f42-a8a7-fac2463f6dc5 · inbound
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71f63490-e1d6-4585-b818-c0693bd39714 · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cb9f199-ad53-4cca-ae65-d2774ebd1ba3 · inbound
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 273
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 737e8445-6ce5-4896-920a-334a477100e6 · inbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e5d07f4-d36f-483e-9df5-acce550ee1ba · inbound
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13dd085b-36e2-49f1-b5c6-9e8b5f9fdc4b · inbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9315eda1-2d6a-4f23-9543-8901b8611b92 · inbound
Why Do Multi-Agent LLM Systems Fail? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43e780d3-cafa-43ff-98be-628f5a5392bd · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 240eb6cb-23e8-4411-b9be-1ccc02f26069 · inbound
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 82c6491b-073d-48cd-9408-004033b0667d · inbound
Seed1.5-VL Technical Report Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation db66ab1b-95a1-4555-9462-5ab75a143a43 · inbound
Tuning Language Models for Robust Prediction of Diverse User Behaviors Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c52f11c7-9882-45bf-87f5-182aa1952a75 · inbound
Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation feb105f9-173b-4e0e-9f53-991e3a26dde0 · inbound
Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1041e806-0c87-4455-a0d6-e4349fdc1a2f · inbound
Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 464ecab5-6583-4867-a904-ba5f6ed32fac · inbound
LLMs Judging LLMs: A Simplex Perspective Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9c4b1a50-88ea-4461-944b-2f7ed61a1caa · inbound
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0f794dd2-4430-4638-b7e0-a2e529df5015 · inbound
Listener-Rewarded Thinking in VLMs for Image Preferences Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 314c7ca6-8ba2-4e65-85ea-e43f7ebea27e · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b4852fc-c0e4-496a-bfd9-8c4fe01178ae · inbound
Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 929cbaa5-cc58-4f21-b1da-598553038c49 · inbound
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ab01ebc0-fa6f-48e0-b1f0-92babacb6e70 · inbound
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bf28a27-842f-47d5-be4f-26a3a1b8ddb3 · inbound
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8de947b-8b25-4cb3-871d-ecce4390806c · inbound
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc29ab5-a9a5-4367-a3c0-04ee523c00e2 · inbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005b982f-aff8-4b26-92a0-123ac3976a19 · inbound
Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 963681bd-0677-4864-9bc4-8c316c747d66 · inbound
PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d55773f-5807-4f6a-b8e0-4bb875b9df02 · inbound
How Small Transformation Expose the Weakness of Semantic Similarity Measures Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b02f4cf-8c24-475a-8010-0af7e54faa1b · inbound
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a58348b-4a58-4cc3-ad89-25bfcc4479f0 · inbound
Rethinking Human Preference Evaluation of LLM Rationales Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228824bd-89a2-410c-b3e9-c9a035ef7786 · inbound
Tractable Asymmetric Verification for Large Language Models via Deterministic Replicability Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012c0adc-9f11-47b1-93ed-1f892e9ce70a · inbound
Evalet: Evaluating Large Language Models through Functional Fragmentation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 009d1fb1-6d35-46d2-92bf-eef58b50ac02 · inbound
InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb320e90-b39c-420a-97c4-507f3a07e8b1 · inbound
Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5c8ed9b6-98ce-49dd-a0d8-ef30cdfba3ce · inbound
CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c14c3a-3b70-40cb-95c9-78c013e4aab9 · inbound
QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b7456044-d828-481a-ab2c-7a238b0c30d3 · inbound
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe30e200-780e-4307-b31f-fa45eebefbc3 · inbound
Automated Alignment between Elicitation Interviews and Requirements Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc5df69-5d9b-4765-8e74-6c377448726c · inbound
Aligning Deep Implicit Preferences by Learning to Reason Defensively Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c8ae59d6-fcb6-49d6-ade9-2f0e63befebf · inbound
Aligning Deep Implicit Preferences by Learning to Reason Defensively Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8908a78f-6480-4ed7-8c6a-940a89cf5e6c · inbound
LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1e412267-22f4-4a34-ae49-cb608c54c2aa · inbound
MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40d205e2-e6c5-4a37-b3a2-f9ed15946c10 · inbound
ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4391a83-ad5a-4391-8bc0-acb6d360c5b2 · inbound
Reading Between the Lines: The One-Sided Conversation Problem Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4d09e5d7-c31a-411a-aced-dedf4d0eb6bb · inbound
Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1ec81eda-673f-4822-9042-aa18cec22a17 · inbound
Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1469f486-4021-4d8b-be53-a1ccd05ff185 · inbound
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497b1f07-4f75-4ba9-b1b0-8afabefe53b2 · inbound
Pessimistic Verification for Open Ended Math Questions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c00e22b8-cfd9-4563-be13-59d5f11344b9 · inbound
PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32d189e4-7b3f-46f3-b09b-1801d253ebbf · inbound
InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bd1853e-0f8b-4e3f-979b-c5a6cc8abc97 · inbound
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d67378b5-9685-4d9d-995c-1501f041aa9e · inbound
ClinicalReTrial: Clinical Trial Redesign with Self-Evolving Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cefaacd-a7aa-457a-a0a3-e5ec2a3270a2 · inbound
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93e96a40-a3d7-4f6e-ad10-243df9e1d880 · inbound
LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d60f9d61-758e-48e0-9c71-f23dc7a8c5d1 · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 277
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e028793b-e67c-49a0-aeb8-c55dc563c994 · inbound
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c33a4aee-cdf4-4ea5-9312-67612fb98833 · inbound
StepShield: When, Not Whether to Intervene on Rogue Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b482c4bc-182e-4f88-ac1d-f7eb0691517b · inbound
A clinically validated framework for auditing AI chatbot behavior in mental health interactions Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e287a0-cf16-40a7-a001-24d90415dbda · inbound
SAGE: Scalable AI Governance & Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdcf0897-057f-4d0e-8a6c-358ac17044e2 · inbound
Bayesian Preference Learning for Test-Time Steerable Reward Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc65662b-f03e-46ea-a57f-27068119551c · inbound
Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5999a959-c4eb-4ff2-9082-9569ed568d9e · inbound
Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367dc6c0-0192-4d2e-983b-0a93212f9910 · inbound
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed96ed93-6686-4047-ad8d-e2295fafd973 · inbound
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5064c39e-4149-440d-ae01-72fc7c1b281c · inbound
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aaee713-1b19-4cdd-bdad-d1e9b66717db · inbound
AI Can Learn Scientific Taste Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d889f58f-e4d3-4d8f-a0e0-0e0f5390f830 · inbound
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a92dfe-273a-48f3-b703-0bea38e7b0ab · inbound
Scalable and Personalized Oral Assessments Using Voice AI Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0a9f836-ec31-4153-9cbc-329abce6a5d1 · inbound
Scalable and Personalized Oral Assessments Using Voice AI Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17103961-9333-406d-bad5-4f3a69ef99b8 · inbound
Synthetic Data Generation for Training Diversified Commonsense Reasoning Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb9902cf-92fd-48da-bec6-939bd008ded6 · inbound
Agentic Business Process Management: A Research Manifesto Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fbef79b8-acef-4017-8f48-719f9938cd32 · inbound
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdfec1b5-24c0-4f2e-8e50-81691499ed1c · inbound
XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8474c254-e454-4893-aec2-e0185d31bd06 · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31db940e-5b6c-4a30-95d8-e0713fadf8a6 · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.