Pith. sign in

Paper Citation Record · LEDGER

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

As of 14 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2507.16835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16835 v2

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:05:43.261112Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T17:15:02.257046Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T17:46:07.347517Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7571889b-130c-4c48-bbc3-5c6ff5517ea7 · outbound

This paper cites Better together: Quantifying the benefits of ai-assisted recruitment, 2025.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Better together: Quantifying the benefits of ai-assisted recruitment, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.688806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:42.587667Z digest=sha256:69794b52d2e0b26eb4ca650c98d41aecc9cc7975d77e928719c6ef00cccc3678

Observation bd615ccb-ebe8-421b-ac83-bf323d9b7eb4 · outbound

This paper cites ESPnet-SDS: A unified all-in-one speech-to-dialogue system.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems ESPnet-SDS: A unified all-in-one speech-to-dialogue system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.669263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:42.667250Z digest=sha256:d26f8d22d3aae8a62c12eff161f20450ad4d372a7820f4ae97f06a4fcf55799a

Observation c9f17111-0357-4308-bf92-4635b573e950 · outbound

This paper cites Machine learning and information theory concepts towards an AI Mathematician.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Machine learning and information theory concepts towards an AI Mathematician

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:05:43.551975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:42.785348Z digest=sha256:c1577c8940fd0e0903969fd8149deec881baa7b376d114c42e29b79b0c32232d

Observation 7307b25c-70b8-4d2c-ab19-297a497847ed · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.644956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:42.881789Z digest=sha256:3022ce9b874a94f51a21977369956f0d527e8d30e79f9adc2a7317912aca1f7d

Observation f4e00927-7a6a-4e1e-9d73-082027da7e44 · outbound

This paper cites Lslm: A listening while speaking language model for real-time full-duplex dialogue.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Lslm: A listening while speaking language model for real-time full-duplex dialogue

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-06T17:05:43.517081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:43.009263Z digest=sha256:7d1f6e9cefdba4ef91a6ee9effe574c19e109df9d08bec773e8afb86815115d3

Observation eb50e7f1-d3be-4d55-bc0c-324f1c2e4024 · outbound

This paper cites The Dirac equation on metrics of Eguchi-Hanson type.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems The Dirac equation on metrics of Eguchi-Hanson type

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T17:05:43.323684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:43.130881Z digest=sha256:4e7355e5bc054e9bd9551d541c7544e1ad9db7f8fae9fcca366be66412ed8755

Observation 282dd8fd-7200-4784-812b-1fcfd8307590 · outbound

This paper cites {index}" - LLM model: {model} - interview_transcript:.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems {index}" - LLM model: {model} - interview_transcript:

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.621529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:43.199595Z digest=sha256:5d400f803c029fcd655acd1eaf623dedc6fe619ae1676439c592ddd29424840f

Observation 51a0b8ee-b2e6-4a57-ae10-0130b839df4c · outbound

This paper cites an unresolved cited work.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:05:43.602377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:43.254652Z digest=sha256:529a9afcd09a63c1c327be49698a76e4a62cf7bbabbf630fdc22d7afc96362a8

Observation 20b4e3a4-03e6-4eaf-a0c1-7ceb94421629 · outbound

This paper cites Mid-level.

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems Mid-level

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:05:43.575860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:05:43.261112Z digest=sha256:979ee6f951735802cd9b86f2d76fb31ad57a7b5f9dc73fd8b706e55b1a9c207b

Pith citing papers

Observation dd08e919-6386-4dd1-b338-b2a8393b6e5b · inbound

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB) cites this paper.

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB) Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:07.362073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T17:15:02.257046Z digest=sha256:6dc2fec27c00914cef03db834f88d26810ff7d8a27d4d7e29f2dedddb9c049b0