Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T08:50:44.523468Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2605.21748.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T08:50:44.523468Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c397a752-f00f-463b-81e8-9c74ad5378f6 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Judg- ing LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b8b526f-84d0-44ef-904d-c9566cb7bf0c · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Gonzalez, and Ion Stoica
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2dbe8f62-9060-477b-a248-2d40c0304930 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Gonzalez, and Ion Stoica
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bbb4ce5c-909c-4a56-b2c6-0e15503beb4b · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator LLMs Get Lost In Multi-Turn Conversation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 19550c42-0882-4a97-a7e3-2aa784fb7ec7 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Measuring Massive Multitask Language Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77a09f96-5efa-42dc-a74e-bf6429047934 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Pervasive label errors in test sets destabilize machine learning benchmarks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc71affc-4384-4226-9cc2-6a61c34c550b · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4797f7bc-e324-429a-a53d-548359216bb5 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator EIP: Weighted ranking of LLMs by quantifying question difficulty
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8388305f-03e0-4594-86aa-69862c853af4 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 41850b87-4dfb-4506-86ee-82b20beb4ebc · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d643d0f-9c26-4fd2-a221-3ca604943af9 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c6d17e7-c342-42a6-8b72-14dcfcec858c · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MT-Eval: A multi-turn capabilities evaluation benchmark for large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation feaf168f-dd20-4bc9-bfc6-145b91e19e42 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Hal- luHard: A Hard Multi-Turn Hallucination Benchmark
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d9eb932-1397-4d66-96b2-5bd407f1352a · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6fec7f69-89a1-40dd-9c9d-864d82518f8e · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e1e2d7f-8f8c-47be-abde-f937c9b5a805 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Training language models to follow instructions with human feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf3b7733-0e33-46d6-aedd-4d2fefc8b75d · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Deep reinforcement learning from human preferences
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c2d4c3a-4079-4c54-87cb-98deaa7e474e · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Learning to summarize with human feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b05c606-d779-4baf-9e46-a6356531a9e8 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Ask a Strong LLM Judge when Your Reward Model is Uncertain
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f41e1909-b7c1-4630-8ac6-fd8c9a964595 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Is ChatGPT a Good NLG Evaluator? A Preliminary Study
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6bcf78d-c174-4a8c-bf50-a055da0767f5 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator G- Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4443390c-581f-44a1-b74b-4d1c0b3c8ab9 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator RLAIF vs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a06a47f-9ec7-488a-a628-7c8396bfe37b · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Prometheus: Inducing fine-grained evaluation capability in language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8dc8fe2b-9429-45fa-9399-44eed1ffe47f · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b386ce00-2394-4ca6-a25b-82d851a64f3c · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e5767bd-8062-46ec-a391-e05dca7cdd8c · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c554d2b5-d24d-4f14-b2cf-97bc3b522b19 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b0cf3f7-beb6-41af-8351-99b6a3089e43 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator J1: Incentivizing thinking in LLM-as-a-judge via reinforcement learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f11318c-4357-42f5-9019-f685a6176449 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator RM-R1: Reward Modeling as Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5090c3c7-0887-4b8b-9010-ec258418ef6e · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Evaluat- ing large language models at evaluating instruction following
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd157f75-b655-4451-b48a-a27f23f7eeec · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator JudgeBench: A Benchmark for Evaluating LLM- Based Judges
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ef7233e-79cf-4dcd-a599-be05daa0628f · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Does context matter? ContextualJudgeBench for evaluating LLM-based judges in contextual settings
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6be2ca06-9dfa-41cd-a413-b8ce2d723752 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator DHP benchmark: Are LLMs good NLG evaluators? InFindings of the Association for Computational Linguistics: NAACL 2025
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 930b188d-1083-4cfb-94d0-b740ef599e7c · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator ReIFE: Re-evaluating instruction-following evaluation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b945eefa-c929-4d80-a575-679be178343f · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator JuStRank: Benchmarking LLM judges for system ranking
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fec24612-ca4b-49ed-8c28-522874057890 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Judge Arena: Benchmarking LLMs as Evaluators
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 41eab83e-6521-4db8-a67a-75b282aea111 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba92f790-1a9c-4d40-8b39-598a5081d137 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 61ff39ff-1311-4d2c-ab9b-a12dae368ce2 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd2a0700-e44f-4903-9643-9dede2d6fcab · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be4c97ff-f6fe-4a64-9e0a-f2ae63ab1b14 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Judging the judges: A systematic study of position bias in LLM-as-a-judge
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d88b8b3-83c2-4d2a-a323-e729364e90a7 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Wider and Deeper LLM Networks are Fairer LLM Evaluators
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 172b778a-1a33-40fa-b5ae-2f65f75578ae · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Reasoning Gets Harder for LLMs Inside A Dialogue
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 683b3e46-d2b5-4043-aa98-94d0428e0b27 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Gemini 3.1 pro
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ccc7a5d7-1e24-412e-8693-b9fa25f3ff09 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Cresswell
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6de7d784-c2fc-4677-8d3b-4db9d9541aba · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator HaluEval: A Large- Scale Hallucination Evaluation Benchmark for Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca6b4567-6dfe-46b2-9661-1e97a9db6637 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Aegis: Automated error generation and attribution for multi-agent systems
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed875923-2118-4878-8ba8-135a5e9cebf5 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Self-critiquing models for assisting human evaluators
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6cb3bf21-fac4-44e9-9d8d-92424ee6e417 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Let’s Verify Step by Step
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 579a2790-fb87-46d6-89b6-1909f593d008 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Variable Selection using MM Algorithms.Annals of Statistics, 33(4):1617
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a450c63-f270-44cb-895e-2869749cfd77 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Qwen3.5: Accelerating productivity with native multimodal agents, February 2026
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 420aa937-2190-4f6a-8bd9-1c546bed339c · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator LoRA: Low-Rank Adaptation of Large Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a80995b6-849b-4966-ab97-4c25de1c87d7 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71b6ff74-255d-47af-b9a5-d1a5af028ee6 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fe92fe42-b33e-4bf5-b91f-f4a716ce4311 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator SP500-EDGAR-10K
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b70199df-1dba-4c23-ac8d-45b9595ab487 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator OpenRouter
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fb722dc-c9ec-4b27-9f66-a60cb9d2502b · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Think-J: Learning to Think for Generative LLM-as-a-Judge
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d2c64139-9ed3-4a90-81bd-fc1b503b645d · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Claude Opus 4.7 System Card
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2abb1c9f-cbf9-4aee-83e3-94a9ed4362f2 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71cfcdf7-1810-4818-84f1-0eacfc9b5dc3 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Language models with conformal factuality guarantees
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4c325267-5ef6-4c53-b287-7aff4030d433 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Cresswell
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0bead259-eb68-4fbb-90c0-0be214c92082 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Cresswell
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a4557e47-a7b4-4e1a-9a67-f7edcf2c5ce1 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Round 2 is mostly clear and foregrounds the main correction, so it does not actually exhibit the required disorganized flaw
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bf41c687-2e54-4d73-b386-18108eaf8236 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 668fe133-2b7c-4869-b66c-8a697b3f768e · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator The plan is a single narrative paragraph rather than a turn-by-turn script, but it must be specific enough to determine which turn realises the flaw
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa53d657-1330-411f-a386-383bd90f5f44 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator If convo_b exhibits a different assistant_behavior_type than the one declared, or if no turn realises the planned flaw, labelnoise
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7dd63a76-9ca5-484e-a893-3fad2ae72ad8 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 547edbe0-0634-40b1-a7a3-cbcd019d2d44 · outbound
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator those little indicator things
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.