Pith. sign in

Paper Citation Record · LEDGER

MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2401.16745.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.16745 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:53:23.265499Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e09085b0-df98-48c3-b9dc-4cc05a50a673 · inbound

LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation cites this paper.

LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T18:46:13.271391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:46:13.271391Z digest=sha256:b932cdff231ca4f62e696e321dde23a67c2ddbdb71cb8af33c8166f27f3e8e40

Observation 83aea4ad-7624-47e5-8d3b-e0ca6f8fcdae · inbound

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems cites this paper.

MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:22:56.715566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:22:56.715566Z digest=sha256:7498bae39b479bb3119675eee298e2ad7e5d8a4651e1ac0d4ae8342fd4b31d99

Observation dd26b6f7-8aad-4975-9346-9dd74318f6ac · inbound

Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types cites this paper.

Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:11.193825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:11.193825Z digest=sha256:95a617fec746e8c64ea7c32a070ae356a13eac0e72d48ae2b06d7e9e7a5ff922

Observation db642df7-edfb-4da8-bde4-4a837a70e2f1 · inbound

PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides cites this paper.

PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:26.569174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:26.569174Z digest=sha256:fd11d8a3aebda2ac09501dcb6bf5db61e6633ebf869c8204137ba188ef263000

Observation 570ba48a-6af6-47d8-85df-0e53f3d8e0f4 · inbound

MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models cites this paper.

MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:47:38.147969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:47:38.147969Z digest=sha256:ec88039c5122bcd1375896b9a55d291daaed1051b5d0b8900705d3242d6b9472

Observation 372e1cc8-c5a1-4a4b-b150-799272fb5cfa · inbound

Towards Efficient and Effective Alignment of Large Language Models cites this paper.

Towards Efficient and Effective Alignment of Large Language Models MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:42.704868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:42.704868Z digest=sha256:414e4e59e1cbae529b8fabb7735941e75bd12c911c082882bc0da8c45aa593ac

Observation 6371c8b8-2223-47bb-ae69-5fa74cfed51e · inbound

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities cites this paper.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.874023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.874023Z digest=sha256:d8b52b96cefc63e1f652d92e142103c77f8d23031f7ede6317c68ae223025fef

Observation 4950a1a8-6d54-4e7b-9100-741d92b2e643 · inbound

ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning cites this paper.

ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:53:23.265499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:53:23.265499Z digest=sha256:7ea2108782882978e648d2a58e26b67123375fdc65a9bea3ca6a1a29ae1b8f88

Observation 4f84cbc8-05ef-4b6a-be83-5eaf92c4356d · inbound

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing cites this paper.

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:55:33.840163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:55:33.840163Z digest=sha256:0f231617fa9583bf55b68c0e84ea43ddc36aa0eb3b92e490ce008dc3ee9f8a0a

Observation 5a0fa028-b01f-4103-82d8-16698ef36b85 · inbound

A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher? cites this paper.

A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher? MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:50:31.919847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T23:45:54.490984Z digest=sha256:65a4982483165c6a6d23f5f837bd350afa8867f8c7672357cefe243835549564

Observation 90453220-c5b7-46d2-a3d6-291a51294cbe · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:26:02.416543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:729c74123396077a74574cce1749d6d8100e8f290c5425b2cb8db2af25aa0dd6

Observation adbfc25d-0dd6-496e-91d5-c9389cf933ec · inbound

EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts cites this paper.

EditPropBench: Measuring Factual Edit Propagation in Scientific Manuscripts MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:00:37.470180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T19:07:57.029024Z digest=sha256:82d702cbe0083592a6efc93d763f6cdd9ceb65bdc4892a9174fe8c2f55a5a5d9

Observation bea8f247-7084-4e4f-9d98-954c2932bc65 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.085404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:b425087cf0441d04b76cbf9a011fa77d474ca22ff11cc33b2434db797f5e9f3a

Observation c12f2aa1-818a-48a9-82aa-19ae31de65da · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.666168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:5f1ebe2d663db2f122cad2e6ae89f3703daccb06348569718c427954f604fc69

Observation 92bfd61c-b505-45b1-9217-309de1853398 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.727354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:d26a28430a9debc4284a9d1502d4a606df8ccfd5afd57a6bf899607cf70edcb9