Pith. sign in

Paper Citation Record · LEDGER

MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2501.17399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17399 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:58:32.983533Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.198096Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation af5b04b5-9afb-4daa-9b21-9b3b41903e3f · inbound

Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities cites this paper.

Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-09T18:58:32.983533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:58:32.983533Z digest=sha256:5d8e392b14d2472b8b48e97411124ac18547f8292838f1b0a0be5001f75cd5d6

Observation 140c103e-9697-47e7-b78b-0a93bb3b179d · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.335166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:ceb33249582ea02b1e5126d65bc288220c7a5e841089623735cb5aa5392f6e12

Observation 6e46e7b3-0481-483d-8932-7b2ab6cebf05 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.420401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.420401Z digest=sha256:cef5d1545ac1ceb906b4931d30db907747ebd884d55821cd893e1cd552624caa

Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.110876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.110876Z digest=sha256:f8d1150b226668823abc99ce6c124093ae88ce127d17b6b9fdaf4bd4ba3e0116

Observation 583020f1-6aaa-4059-a310-44715a23cdb8 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:17:45.350173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:f4e51cfee58d333fbc4b1ccfe05916dc455a63af407a9f057736c2e59732b78a

Observation 0e6f16b8-a732-4ca0-90e5-adf7914b8582 · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.518006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:eca229a7a3ee9cd6f30345b96487d341d1dd9e2fa2d0b74a6aa21f8cad162ae8

Observation c76d5d24-86a7-48e5-a1ca-4b014cc39af5 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.856959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.856959Z digest=sha256:ca569efc1defa9d7a0ec97aefef0ffe305265b0fb7b11890a124b0bd588bf52b

Observation b2aa6ed7-d66e-413e-b8f1-d92c97ad80f3 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.845701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:574fd37f5b39c2d3523c0089fb371e970cf7ff42a7b775daa00ef23746cc3471

Observation 0a513bc0-9837-42b3-8f72-4faf142f9ccd · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.669131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.669131Z digest=sha256:cf2d73b887f64ce49783a15e63b975296de040a184f9270f54dddc21717de215

Observation 57df649d-5060-47bc-9a32-08469a0b455a · inbound

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries cites this paper.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.990246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.990246Z digest=sha256:99452c5e2d0b7c1ce895d749af37711aecdc0d51bf450642da2f210b0a9fddb1

Observation 4cfd9f18-3805-4e8e-867d-263f705e7254 · inbound

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting cites this paper.

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:12:10.089168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:12:10.089168Z digest=sha256:9962f5bd1d94cb78a0ce37f674584bfba1bf4f60e95ae8b32ec1665a62d46e1c

Observation 5c091189-e56c-49de-bc03-2054b4e04159 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:39.919893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:39.919893Z digest=sha256:25c4569dfc8506d1d5e0af393961f9460607d64da8ac606534c08808140561e9

Observation dc68545d-532a-47d9-900e-a2d6861c90f4 · inbound

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks cites this paper.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:46.562258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:46.562258Z digest=sha256:b66d8f15ec7593eb47266e045c76a00ece8be8c8d0b28b62dbb2d0b635bf5a5e

Observation 99759f4e-bdf1-4db9-bfec-fcf4a21a6ead · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:16.105119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:ca98254ac11a2bc9603f736d0847aab5b91ef802c9b02d88cde65115b149ecea

Observation 6d2bcf41-ade4-4328-9738-7dbb80d785dc · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.469751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:126bde57830b9648a0a84b4b282a6c5b487fa9a0a820d7ed81b5ec8a5aab8e42

Observation f6d4d5ec-bc87-4f69-a129-da250de9fe45 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.445640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.445640Z digest=sha256:420347f17fc411f4653f87cf08c22c8022d0c421467c0646ab8b7a4694f5b018

Observation f49b607a-92f2-41a3-8729-aab61af32184 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.360793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.360793Z digest=sha256:a62c3c0926b30b2e12a8c6357260a0c86439239d59e57d78b13539e13a98d875

Observation c5fb9061-7d35-49af-9232-0e9cad0e51a1 · inbound

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation cites this paper.

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T09:17:40.301640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T09:13:55.204404Z digest=sha256:476e6f10428fb44699ad5b9e3cb5f50cb3e8337b95f861a86ed2898b08983dcd

Observation 2983f954-59cb-4294-81c6-ac074e7e8e25 · inbound

Evolving and Detecting Multi-Turn Deception using Geometric Signatures cites this paper.

Evolving and Detecting Multi-Turn Deception using Geometric Signatures MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T15:23:32.697954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T15:17:58.804973Z digest=sha256:396e858acae1d98a1de83e803e8ffbad5f0ca5574f241ddcd950a7fb2cfe7598

Observation 9b3ea34e-c4b4-41ff-bf02-b439edf5abf8 · inbound

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses cites this paper.

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.740957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T07:16:30.725358Z digest=sha256:e92f537f46474891f971576bad71b640e2fc48bcb953221a9c403abaa2becbb7

Observation d3050268-e08e-479a-b75d-258cccea1f49 · inbound

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios cites this paper.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.351978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T03:39:36.657903Z digest=sha256:01f4db3fd03fc7809063a8f4192e1532bc050204d8bd1c4e66ed3c469655870d

Observation 36c02bfb-a6e5-40e5-8841-42a8188303c7 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.199387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:0e82aee3ae7a5dec2587a0a3768c7c02659a2b620afe6cc291c3865ba9c4052a