Pith. sign in

Paper Citation Record · LEDGER

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

As of 15 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2507.16835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16835 v2

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:05:43.261112Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T17:15:02.257046Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:46:07.347517Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7571889b-130c-4c48-bbc3-5c6ff5517ea7 · outbound

This paper cites Better together: Quantifying the benefits of ai-assisted recruitment, 2025.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Better together: Quantifying the benefits of ai-assisted recruitment, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.688806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:42.587667Z digest=sha256:95ed1ef116bac545bfd369c2f7098b9470961fa98ba9a27dd1a243705187daa4

Observation bd615ccb-ebe8-421b-ac83-bf323d9b7eb4 · outbound

This paper cites ESPnet-SDS: A unified all-in-one speech-to-dialogue system.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems ESPnet-SDS: A unified all-in-one speech-to-dialogue system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.669263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:42.667250Z digest=sha256:fc09235b6af0e95342fb8e5d3629b340f51f608e0526495ec9eac83d46eb1a4c

Observation c9f17111-0357-4308-bf92-4635b573e950 · outbound

This paper cites Machine learning and information theory concepts towards an AI Mathematician.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Machine learning and information theory concepts towards an AI Mathematician

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:05:43.551975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:42.785348Z digest=sha256:43313c8749a638ae4ba1fea12ac77a4463b83a03c8be8ac6afc709b8ef7d9e88

Observation 7307b25c-70b8-4d2c-ab19-297a497847ed · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.644956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:42.881789Z digest=sha256:dc07212bdd04caae87c5874e12fdbbd3bf480828cec79163a1ef00ce2acd08c7

Observation f4e00927-7a6a-4e1e-9d73-082027da7e44 · outbound

This paper cites Lslm: A listening while speaking language model for real-time full-duplex dialogue.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Lslm: A listening while speaking language model for real-time full-duplex dialogue

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-06T17:05:43.517081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:43.009263Z digest=sha256:be09899b6a86de75bc435b4341a88a93207a8b49f7a74b5ef8b7d8d485ca283c

Observation eb50e7f1-d3be-4d55-bc0c-324f1c2e4024 · outbound

This paper cites The Dirac equation on metrics of Eguchi-Hanson type.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems The Dirac equation on metrics of Eguchi-Hanson type

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:05:43.323684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:43.130881Z digest=sha256:3c640dd6c62c0b3ad3fe0a6f5a7917e8aa01b743faf83c9db526a582b367367e

Observation 282dd8fd-7200-4784-812b-1fcfd8307590 · outbound

This paper cites {index}" - LLM model: {model} - interview_transcript:.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems {index}" - LLM model: {model} - interview_transcript:

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.621529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:43.199595Z digest=sha256:87483fb98a77b9dc354205a2d0f6fa90b0ad956f89f79b5674101cdb54a7a0a0

Observation 51a0b8ee-b2e6-4a57-ae10-0130b839df4c · outbound

This paper cites an unresolved cited work.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:05:43.602377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:43.254652Z digest=sha256:0e32e00bcf684a75f548725db19c39b592b0117d37c0f7ff9d2d23d5a13e5c01

Observation 20b4e3a4-03e6-4eaf-a0c1-7ceb94421629 · outbound

This paper cites Mid-level.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Mid-level

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.575860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T17:05:43.261112Z digest=sha256:0fd387f339c99be4f49d42a5003d03fe597ce83792ad3395a1979c109474ad32

Pith citing papers

Observation dd08e919-6386-4dd1-b338-b2a8393b6e5b · inbound

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB) cites this paper.

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB) Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:07.362073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:15:02.257046Z digest=sha256:9b1bd94a637447ac2e2fda0372d2cad3903a7fff2f71f746886a614fd979b696