Pith. sign in

Paper Citation Record · LEDGER

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2506.02457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02457 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:23.320594Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:22.389818Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T07:39:48.680240Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24a79a7e-e7f3-46b6-8e29-f2eabedc2326 · outbound

This paper cites SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.389818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.389818Z digest=sha256:8946c1327029bf96aedf0f313adfc5a8f4b51f147b8cd16f13d283be13741748

Observation 6935127e-fa43-4f0d-b77e-26f4d1a3a331 · outbound

This paper cites Speech LLM Speech LLM extends the understanding capability to speech flow, performing modality alignment between speech and text via an encoder with adaptors.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Speech LLM Speech LLM extends the understanding capability to speech flow, performing modality alignment between speech and text via an encoder with adaptors

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:26.828263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:22.451786Z digest=sha256:2240bc309856ab5284a9f3c42dfdefc2ca6cff65d2210667c938b2f59d8472c1

Observation 335ffbc3-ee3e-4578-a789-24643a2af936 · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:26.608938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:22.535835Z digest=sha256:ab0a9f94648a62c5ca57820b4d72601357ef25b7e577f0e0a8724db9e6e198cf

Observation 3007e468-ce9c-436f-8f31-d8c14bf5e103 · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:26:23.863004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:22.625987Z digest=sha256:ac3a5e956c65bf5417e60e07133a493e4bffecb4a3126a10720709b051a39311

Observation 741100b8-5df6-4d9e-8de8-ec39610b01af · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:26.436653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:22.693809Z digest=sha256:b4caaba8014be83f1edca0220ea351dd544a132db06c8e9686aa6c4c5581bb82

Observation eed943fb-c74a-42e8-86bb-cac7be6255bb · outbound

This paper cites an unresolved cited work.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:26.295669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:22.770624Z digest=sha256:f81892001837b631e7d666a45d6ac9ae710c0cb37f3d75fc8ec9d6f3da23bdcd

Observation 04f99b43-0c47-40cd-b373-185506acfab5 · outbound

This paper cites GPT-4o System Card.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant GPT-4o System Card

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.870915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.870915Z digest=sha256:87e63f5e9b7a97aacaa37f585f9d7399bc229aabb22b52fba5266e9e0487c195

Observation b1ef39e5-7134-495b-aee8-c300409e6154 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.926836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.926836Z digest=sha256:4eb923c1b24266e6b31aee39cd5e2e2dad95566bdf7f6d5320a7b9636fed8b5d

Observation 3a38c437-6768-4ac8-b6b1-defd92aa8332 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.008363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.008363Z digest=sha256:0d6ac555e47ea736f8e3834d9a6fe38c859df66d3bc7ec27ad04ccf0137b555f

Observation f821b026-4326-4f47-b148-bb854874ed7f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Moshi: a speech-text foundation model for real-time dialogue

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.050878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.050878Z digest=sha256:47294e4127033844288701bbc4e89dfd2299a5027e98c541d6c476eb462fcff3

Observation 657ac16c-0126-4684-af44-6e56729f7170 · outbound

This paper cites Dynamic- SUPERB: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Dynamic- SUPERB: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:26.163525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.138843Z digest=sha256:1dbf68a5365b96e7c25dae3cb96cacd9576edea91a73011f79e0a3bdf5f5bdd2

Observation 5aad3ced-54b1-4a76-8309-61beb49207e2 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.197926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.197926Z digest=sha256:7d63e4497098ddfed1b4312932c8dd15ef07c271dc7eb397a1bdfec161bb0074

Observation 0db694d5-841a-41cf-99e1-1c9efe1a8c85 · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.203188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.203188Z digest=sha256:f871f096302410d039d81ff97fd42e1dd2b6e1145bff7bacda4a9266cae0f7cf

Observation 1dae2d60-7300-4edf-ad27-3acb7069dff2 · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.208362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.208362Z digest=sha256:45ab51131a01d869c98996daf408db47b99f56a1849019367873833edbb8100c

Observation 277382e1-99a2-4fc9-ab1b-e2432272bb66 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.213004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.213004Z digest=sha256:9bde6de724e2523efcc1390e5c3d0b3a6661873cd5a6407466f7cf971bc1a4c1

Observation 3208b10f-2a13-4b05-85bf-ef46656c5d91 · outbound

This paper cites SALMONN:Towards generic hearing abilities for large language models,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SALMONN:Towards generic hearing abilities for large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:26.018205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.217381Z digest=sha256:f46f65eed3ea63d6fd720b301d30d2429dc8fd4b9665cef8dcb18f144dd5c264

Observation c243da26-5a10-41da-9c47-b0fc12b019fb · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.916923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.222170Z digest=sha256:aef345882f51a396e54fb0ac881c2724404d04dcb747ea158b6343aba8d3bee7

Observation 2f99f1ab-e977-4b39-b9f2-dc13e6840bf7 · outbound

This paper cites Qwen2-Audio Technical Report.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Qwen2-Audio Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.226286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.226286Z digest=sha256:52d6b321a255714b0ec062f8f96d4bbf391115a03d4d5033bb329c4a6c0db443

Observation d72b60f0-b1c0-4ad4-86ed-f944745a5b04 · outbound

This paper cites SNAC: Multi- scale neural audio codec,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SNAC: Multi- scale neural audio codec,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.756587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.230921Z digest=sha256:e0d84845eda3ead41fce258d3d33a97d14d2fc64184c3ce45abca83e450effbd

Observation 78bf5a99-e0a4-4aac-ba3e-64c54fd2a8ab · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.234959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.234959Z digest=sha256:548d47b6f01305ced45ef56abdcb9b7cdd393e508563ac20970c9fa7fd7ada3f

Observation 25d6d93a-a35b-43e3-a76c-698b211e4008 · outbound

This paper cites HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.610697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.239575Z digest=sha256:690cad50bd543718c00da9c29c746162335faeeafe97a0e3464d8839f16b3014

Observation 81babd28-831b-43ef-be8d-cef09381c961 · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Speech resynthesis from discrete disentangled self-supervised representations,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.430775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.244104Z digest=sha256:8117e9b6cba774ebd0d07a8dbe6604c1471b400909671ae4accebbd507980bea

Observation 37d13974-9f50-4609-bf99-3a907d1da2c4 · outbound

This paper cites Westlake-Omni,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Westlake-Omni,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.265736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.248523Z digest=sha256:52506912405766e1163f868ffb71dbe06c3a3c5c4103d9f681a1befaf3aab964

Observation fd4d79c4-b7f4-466a-bb6d-a80d9ec59e96 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.253202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.253202Z digest=sha256:19a384b8a07dd8bd3a09d3c541012698abdb5e35cbab688b8c4e3b4efd8a15b4

Observation f725328a-7a76-49dc-a579-cb24b7ecc36c · outbound

This paper cites Be- yond turn-based interfaces: Synchronous llms as full-duplex dia- logue agents,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Be- yond turn-based interfaces: Synchronous llms as full-duplex dia- logue agents,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:25.096745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.257541Z digest=sha256:c1d40968b4ee46d0b5a5c0dd6bd7f1e46e1bdcffd8d5ec6602a29d1453a4ddca

Observation 101c2dd5-5392-419c-820a-30b3c3d19b08 · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.261871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.261871Z digest=sha256:51a4df470936f024f0dae117865e9b297ccba25ad051007c905555cd90be1c30

Observation b577a389-a753-4102-adbf-4c98624ca753 · outbound

This paper cites Baichuan-Omni-1.5 technical report,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Baichuan-Omni-1.5 technical report,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.266830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.266830Z digest=sha256:5a3f0b19e4fd6e1df81186dc37f2c41422fd24bff8a19edf54ab3892d7d8357c

Observation f9735e47-eb7d-479c-a8ac-03dc3d0f8cf8 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.898400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.271445Z digest=sha256:35b0821e2a2cb275f9a7c214bc396cfc8163ded26f806db8eafa7bf5a2143b9b

Observation 2733f0ee-bbe7-44dd-b717-497f1ce9d094 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Lib- rispeech: an asr corpus based on public domain audio books,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.276005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.276005Z digest=sha256:36ef61a78d373420d66731934a3217a084018736f7a4a6f953c41c6c409841d4

Observation ffa586ec-7787-4748-bc32-a3b4aa3794d3 · outbound

This paper cites LibriSQA: A novel dataset and framework for spoken question answering with large language models,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LibriSQA: A novel dataset and framework for spoken question answering with large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.706381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.280776Z digest=sha256:f0498bb37726212ed657b1b0a0bc4fff21036e016aa430828a2cae6efb3c9f0f

Observation 4ad395e0-aac3-43e9-820e-feafba8e888d · outbound

This paper cites Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.560476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.285411Z digest=sha256:1966e23182059dbb8c5c1d55259cd0a3ab14e9683e394a2a324b550f506d70c5

Observation 6d061c8f-99cc-4c34-97b1-977e6766fafb · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant IEMOCAP: Interactive emotional dyadic motion capture database,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.289433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.289433Z digest=sha256:8c29289263313dc3592db5d003d2e861ff7f87ee08efd59063ee41871b1f5261

Observation 3f915456-202c-420d-b2bd-e2b094f68824 · outbound

This paper cites Common V oice: A massively-multilingual speech corpus,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Common V oice: A massively-multilingual speech corpus,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.374900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.294124Z digest=sha256:3542c9ec60f4a92422ec0682db1a0a40a530b67fd40fc4f114927be23e7f9828

Observation 16d4b01d-6eed-45ce-aa65-308b97d668b2 · outbound

This paper cites Stanford alpaca: an instruction- following llama model (2023),.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Stanford alpaca: an instruction- following llama model (2023),

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:24.140936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:23.298170Z digest=sha256:edf3a45376fc6cc2fec9f90a6cf53d612c398ac3a6ffad01b60028b0697fc5b9

Observation 3ca2bfd7-9187-4d2a-945d-c5e9210fa9d0 · outbound

This paper cites The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.302596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.302596Z digest=sha256:7b8e9293c57fd4e7d05a1f620a5ade765f9b49022b2daba8ffee705525b6e8b2

Observation f82a96c8-3bec-43ff-899d-1ad5b1167973 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.307131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.307131Z digest=sha256:c597fb62c29eb848bf164930d4cc332794fbd7b3f92fb90d34fdaf31d9d3c09c

Observation 3141e860-23e4-4cdc-b2df-98ae5a89919b · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Robust speech recognition via large-scale weak supervision,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.311743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.311743Z digest=sha256:0d55415c91a2a343660a565400a329229e7dcd266bb1019aef70ab4fe6fc9179

Observation 4b7c8061-2bde-426b-8c21-ff56e53f4e71 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.315974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.315974Z digest=sha256:d4f87e0acaff8d260304f42658fa38973269e2cd49be6fd9c738ee85db45ab2f

Observation 8b15efbd-c7c1-496d-b92e-9470a0a2afa2 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.320594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.320594Z digest=sha256:cc7fad1fee9039d2dc0388d39a77100b9607edef7e2b81f32aa91049da3933f5

Pith citing papers

Observation 24a79a7e-e7f3-46b6-8e29-f2eabedc2326 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:22.389818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:22.389818Z digest=sha256:8946c1327029bf96aedf0f313adfc5a8f4b51f147b8cd16f13d283be13741748

Observation 77213481-1a76-4b50-ab3d-5c2d1777249c · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.681853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:5f787aaedaec6ab84e4aeebc878dd1af224aa3b99fce086df2accb627f098b58