Pith. sign in

Paper Citation Record · LEDGER

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

As of 20 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2607.07985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07985 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:13:50.737253Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact4
  • verified fuzzy5
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd31e340-2dc2-43fe-8ee5-679705374780 · outbound

This paper cites AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:17:10.083198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:594026756347314da5fa577e7007e35805d84f099af96d2ebf424e49d7949946

Observation cecdee03-1df9-4d8c-b611-48ebb03e34da · outbound

This paper cites Audio large language models can be descriptive speech quality evaluators,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Audio large language models can be descriptive speech quality evaluators,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.365750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:fb4f9e5518b3e5647fd217e0562a265770dad26b4d89a6c7f6d70d6954e93a26

Observation 0d40c82e-b9f3-4ec0-a22d-b143d45943ad · outbound

This paper cites SpeechQualityLLM: Llm-based multimodal assessment of speech quality.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents SpeechQualityLLM: Llm-based multimodal assessment of speech quality

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:17:10.073760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:748472593b1b53b192c6ec0a14b64164eaa7af3ee7b6002c4ba780d8f3ecfa6a

Observation 981201bd-73b8-472f-927d-f5b3bc4db1b7 · outbound

This paper cites Interval estimation for the difference between independent proportions: comparison of eleven methods,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Interval estimation for the difference between independent proportions: comparison of eleven methods,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.367838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:da4282f77edc756a21e76f5a83adf2422a82cdbba67dc3a3e8976bcb62880dba

Observation f1998f89-b570-4f77-aeae-9620e3a19786 · outbound

This paper cites ITU-T Recommendation P.808: Subjective evalua- tion of speech quality with a crowdsourcing approach,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents ITU-T Recommendation P.808: Subjective evalua- tion of speech quality with a crowdsourcing approach,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.372691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:242bb56127eb7c8d7a054023af5b704618c96632fa316e9fcaa268304036301e

Observation cf5cbe06-84c3-43e3-a639-b15fb57af392 · outbound

This paper cites Krippendorff,Content Analysis: An Introduction to its Methodology, 4th ed., Thousand Oaks, CA: SAGE Publications, 2018.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Krippendorff,Content Analysis: An Introduction to its Methodology, 4th ed., Thousand Oaks, CA: SAGE Publications, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.359811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:32aafaf8ded63efcb07c28e89daabce581367226847a428dc2fc581af9554445

Observation 994dd9ee-9fca-4aef-aae6-982ac4586c12 · outbound

This paper cites LALM Judge Validation on Full-Duplex Voice Agents,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents LALM Judge Validation on Full-Duplex Voice Agents,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.370527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:81c1c9552372417d53bf6ed1c4c045e15073dee23a50acdfd6cfd56397244702

Observation 25971a8a-8c26-46ea-87ba-1e2cf16b58d8 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T14:17:10.076707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:f3b27ba24c609a11cf3535fe225223128e531b7039bb99107afdb4e82a794b0f

Observation e1514b4b-bb24-402a-9d18-b0c2a5cf5632 · outbound

This paper cites Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:17:10.079665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:dea70907a878ef0575f28584fcef7409833ca91f62b0327e4d47f940bae262b4

Observation 94cd768e-0a03-4f15-8b22-c35b40ebdf73 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:17:10.085901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:1ece89414e712113dee8345235abe8709970f5f77246cd6e2b34324f234774d3

Observation 0006f4aa-2bad-4c1e-824a-dfad92273668 · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Non-Determinism of "Deterministic" LLM Settings

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T14:17:10.070543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:f36fe7e0f638d44e8deb531c9ff61de460e747ecb954e3de239ed727b418fc48

Pith citing papers

No inbound Pith citation observations are available.