Pith. sign in

Paper Citation Record · LEDGER

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

As of 20 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2607.07985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07985 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:13:50.737253Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact4
  • verified fuzzy5
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bd31e340-2dc2-43fe-8ee5-679705374780 · outbound

This paper cites AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:17:10.083198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:bb02cbb194a2dc4d0fd536114ebb68dafdbf06757e67735eafafdbed1049e8d9

Observation cecdee03-1df9-4d8c-b611-48ebb03e34da · outbound

This paper cites Audio large language models can be descriptive speech quality evaluators,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Audio large language models can be descriptive speech quality evaluators,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.365750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:e93a899b9cd32f66f58f9925af0f903e180c9351e745114714239928a3f8575f

Observation 0d40c82e-b9f3-4ec0-a22d-b143d45943ad · outbound

This paper cites SpeechQualityLLM: Llm-based multimodal assessment of speech quality.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents SpeechQualityLLM: Llm-based multimodal assessment of speech quality

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:17:10.073760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:daa5de94fa212fd02d4c5119f5a40db287c0d29f01950ca1c37b392415bbea29

Observation 981201bd-73b8-472f-927d-f5b3bc4db1b7 · outbound

This paper cites Interval estimation for the difference between independent proportions: comparison of eleven methods,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Interval estimation for the difference between independent proportions: comparison of eleven methods,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.367838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:bd70c3ff05bde4ffebcd1e649c208f0348468a0800fcd136e2599cd0c2fa11f8

Observation f1998f89-b570-4f77-aeae-9620e3a19786 · outbound

This paper cites ITU-T Recommendation P.808: Subjective evalua- tion of speech quality with a crowdsourcing approach,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents ITU-T Recommendation P.808: Subjective evalua- tion of speech quality with a crowdsourcing approach,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.372691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:867a64dd26be4f8452b817bb273924a63a5c8c3b7cb23a632a8a39c0f55ea407

Observation cf5cbe06-84c3-43e3-a639-b15fb57af392 · outbound

This paper cites Krippendorff,Content Analysis: An Introduction to its Methodology, 4th ed., Thousand Oaks, CA: SAGE Publications, 2018.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Krippendorff,Content Analysis: An Introduction to its Methodology, 4th ed., Thousand Oaks, CA: SAGE Publications, 2018

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.359811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:29428761cec97db746482f3535ac561487160a87539fd10bfae406a5450e4bad

Observation 994dd9ee-9fca-4aef-aae6-982ac4586c12 · outbound

This paper cites LALM Judge Validation on Full-Duplex Voice Agents,.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents LALM Judge Validation on Full-Duplex Voice Agents,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:17:10.370527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:ac0befda83a6cbde4da44c4de8ba3cf6d3d3d48d74225461ac645d9acbb48943

Observation 25971a8a-8c26-46ea-87ba-1e2cf16b58d8 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T14:17:10.076707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:26f8b800fd4cb2f0f54e2ec1fb010f0a814fcae3d4df1feff02257c84b8e0f50

Observation e1514b4b-bb24-402a-9d18-b0c2a5cf5632 · outbound

This paper cites Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:17:10.079665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:f494fb84dcd503061a51ceeb5c3ac6ff63814073988672a00258c3ad4111b6e4

Observation 94cd768e-0a03-4f15-8b22-c35b40ebdf73 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:17:10.085901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:b0e28316329e0075fcc1270403a204f8caf35bf41e0d8be17af5b6ed380d1015

Observation 0006f4aa-2bad-4c1e-824a-dfad92273668 · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents Non-Determinism of "Deterministic" LLM Settings

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T14:17:10.070543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T14:13:50.737253Z digest=sha256:9adc6d39112f643bbd4dbb22fd58559f576e6386d24669de18492d35abcfe62f

Pith citing papers

No inbound Pith citation observations are available.