Pith. sign in

Paper Citation Record · LEDGER

CASPER: A Large Scale Spontaneous Speech Dataset

As of 8 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2506.00267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00267 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:48.666663Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20d2571b-8058-4693-9321-3ff0d70de7f0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CASPER: A Large Scale Spontaneous Speech Dataset Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.217794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.217794Z digest=sha256:9875b77950e0079244242791a547271180122350b243d787fda1469a660473a2

Observation a10d61bd-a6b4-4a14-93bd-853225f79f54 · outbound

This paper cites The Llama 3 Herd of Models.

CASPER: A Large Scale Spontaneous Speech Dataset The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.264067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.264067Z digest=sha256:f7aeeb9f89f94a3814b19e57c4059eb78062d2d82201d08e8c66af3fddbbfa48

Observation a5aa4fb5-fcd4-47e0-b9a3-ca9464af18d7 · outbound

This paper cites Prompted llms as chatbot modules for long open-domain conversation,.

CASPER: A Large Scale Spontaneous Speech Dataset Prompted llms as chatbot modules for long open-domain conversation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.412743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:46.337326Z digest=sha256:9dc8586237d3f1c0f36cdf2c8ca74f76928e7090d5ed1d72bda0595a266c9b55

Observation 68702bf0-4fe6-483d-8b5a-193ef597904d · outbound

This paper cites Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations,.

CASPER: A Large Scale Spontaneous Speech Dataset Soulchat: Improving llms’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.256684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:46.400788Z digest=sha256:3295291ebb5199166e27c391e87170207e2b84d5effefa8a8bf1b1018daa0136

Observation 01af7eaa-2dc5-47ca-9f3f-43e76e4097e7 · outbound

This paper cites When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection.

CASPER: A Large Scale Spontaneous Speech Dataset When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:49.271166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:46.473021Z digest=sha256:53d93487b9616a56ca87c6b3552505b288e6931bc3c348eddbb45d1881eda64e

Observation 8f8ffbc0-0994-4e5b-8f1d-117c528aabbc · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

CASPER: A Large Scale Spontaneous Speech Dataset Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.554818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.554818Z digest=sha256:708fd7505cc1406c30f631c9409cd92ccacf0baead3e89609fac39b63d7db8af

Observation 12be38bd-e5b4-4212-8fc6-b77c2f511ef0 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CASPER: A Large Scale Spontaneous Speech Dataset CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.607323Z digest=sha256:ea67aac2500c26a879be03b63a66f7d67b8f8294c28e716da75bad7b4d2b1f59

Observation 585a613a-78b9-4a1a-a340-eb7e9c2aa412 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CASPER: A Large Scale Spontaneous Speech Dataset CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.665649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.665649Z digest=sha256:d600b0c1f9becec431ebcaf330f944ea002342329f5689f2cc19627b5b89f778

Observation b25e6881-bafd-4cab-9324-aa2c04069ef6 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

CASPER: A Large Scale Spontaneous Speech Dataset Librispeech: an asr corpus based on public domain audio books,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.738062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.738062Z digest=sha256:5fc33c47b22cb9c838e8d293e6181538cd2364921dbb9bd147ddb86518bc7726

Observation b9409cea-226f-4a71-8934-78f224063019 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

CASPER: A Large Scale Spontaneous Speech Dataset GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.821940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.821940Z digest=sha256:d2a942f7c56174ac17b2c098c13d7120835734af5e78bea4d1c8cce69c1333c3

Observation def633c0-3a31-4809-b760-03aabdc36b7c · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

CASPER: A Large Scale Spontaneous Speech Dataset Switchboard: Telephone speech corpus for research and development,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:51.104579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:46.894202Z digest=sha256:cc0391faa09323679de93d1cc4fc9fbb6db7ee8aa865fc48f7f790e5f02636d8

Observation ebbbb0b5-1a56-43d1-8e08-a6798db13e8f · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,.

CASPER: A Large Scale Spontaneous Speech Dataset Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:46.962247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:46.962247Z digest=sha256:3d64363c29d5402ddb69ad5d7ce86b548a3fe8f0b8b16d65ff7533005f4703a4

Observation ee7048ed-e1d4-4086-98f8-c0cb7c5203c9 · outbound

This paper cites Generative spoken dialogue language modeling,.

CASPER: A Large Scale Spontaneous Speech Dataset Generative spoken dialogue language modeling,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.946083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.029630Z digest=sha256:649d68a07326f9a45f5c9affa2c6db5c1c7164261bfa527c7578a5bdb8e76504

Observation b0be04b9-5e03-4012-bf34-8f2746fd085a · outbound

This paper cites Audi- olm: a language modeling approach to audio generation,.

CASPER: A Large Scale Spontaneous Speech Dataset Audi- olm: a language modeling approach to audio generation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.095650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.095650Z digest=sha256:8630b813076e1eee70685cbf95f55e0649998b3c74564b232c8aef3813c7109f

Observation 6a76a570-96a6-487f-bebe-f838a94c974f · outbound

This paper cites On gener- ative spoken language modeling from raw audio,.

CASPER: A Large Scale Spontaneous Speech Dataset On gener- ative spoken language modeling from raw audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.794228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.191571Z digest=sha256:f375ff9c57c10abe4a7dd7f9977772d5711442b27ea87ec5c1aecde0bacf745f

Observation 98b52e29-d426-4844-84cf-f880bbf562d9 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CASPER: A Large Scale Spontaneous Speech Dataset Moshi: a speech-text foundation model for real-time dialogue

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.225110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.225110Z digest=sha256:c496e8adf5364adbeeaf4e79c835e533b61eef399c059346900416196f7f4f02

Observation 1c4e6eae-9825-44c4-9fa3-70505062d41b · outbound

This paper cites Callhome american english transcripts,.

CASPER: A Large Scale Spontaneous Speech Dataset Callhome american english transcripts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.639637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.243341Z digest=sha256:81fc10c17d78e2561ab990e917930d7b98a7b6f5212a723d41b44af79a950cd3

Observation 21adb551-327c-41e2-8348-237c749781a7 · outbound

This paper cites Santa barbara corpus of spoken american english,.

CASPER: A Large Scale Spontaneous Speech Dataset Santa barbara corpus of spoken american english,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.464850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.322984Z digest=sha256:b1968713006ce9bf8be049734c5a29368a67433a2a1619c3aff7a5a2e0d98506

Observation f58fb5a0-5858-49e1-8d2f-4cc00fd0c9fe · outbound

This paper cites The hcrc map task corpus,.

CASPER: A Large Scale Spontaneous Speech Dataset The hcrc map task corpus,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.263251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.375416Z digest=sha256:bc788ca1d50d7ce1819ae2933519adcd547b587e005456cbf0d3fdd10262e435

Observation eac548f5-d474-48af-8544-dd4a6551dcd3 · outbound

This paper cites The casual conversations v2 dataset,.

CASPER: A Large Scale Spontaneous Speech Dataset The casual conversations v2 dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:50.101839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.439037Z digest=sha256:6fdb72099d87091b03aeed79ae81cb36dc0ff829b469b975d58d37db6e1911bc

Observation a4d735f0-0d94-4c9b-b4b3-7b003aef0fac · outbound

This paper cites Scalable spontaneous speech dataset (SSSD): Crowdsourcing data collection to promote dialogue research,.

CASPER: A Large Scale Spontaneous Speech Dataset Scalable spontaneous speech dataset (SSSD): Crowdsourcing data collection to promote dialogue research,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.961062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.530317Z digest=sha256:fb4e3f4da571700d3263dd93cd2c6f756529327ef419fa0bd19fb238c7526acc

Observation 93663259-1ebd-45fb-abfb-c9c63adfee40 · outbound

This paper cites Silero models: pre-trained enterprise-grade stt / tts models and benchmarks,.

CASPER: A Large Scale Spontaneous Speech Dataset Silero models: pre-trained enterprise-grade stt / tts models and benchmarks,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.595510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.595510Z digest=sha256:023c5882c39b0e27d9eb6472852aa26efe7c7762c28dc41a1b4574886809620c

Observation fe169053-4210-4de0-8198-e5079bc077c8 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CASPER: A Large Scale Spontaneous Speech Dataset Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.662234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.662234Z digest=sha256:74bf7014453511a08e2d2ae5a33f63d177b742c6ba78cd0f1d3c1e8499f9b9aa

Observation c6d6d965-b0d1-4d6d-a4aa-b65e52633e82 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

CASPER: A Large Scale Spontaneous Speech Dataset Whisperx: Time-accurate speech transcription of long-form audio,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.803665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.716069Z digest=sha256:d3551479b708cbfde880ed1855105395cea0735692a405e718442b8679f49c00

Observation f5e0fc51-7c8e-4be7-92e2-c25b8c34c624 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

CASPER: A Large Scale Spontaneous Speech Dataset SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.792349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.792349Z digest=sha256:fb8085808d29e626c190b38aa2c24a6f88849347db47ae600104711bad015b9b

Observation 2a408a6f-7572-4d92-9c5c-cfc9a5a95e59 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

CASPER: A Large Scale Spontaneous Speech Dataset wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.894977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.894977Z digest=sha256:d4605a6da7eca79795f3035ec1a203f4a8588542bdeaaac37862746d5c09de9e

Observation f8027b9e-d990-41c2-9b85-4cd0b71ccdc4 · outbound

This paper cites pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,.

CASPER: A Large Scale Spontaneous Speech Dataset pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.643117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:47.969722Z digest=sha256:958b47eac8f35760c32c8b1c7db951d41a36f9831e091509f12a935d1663747b

Observation dd6ac9b8-915a-453e-a378-98b7b56aca07 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization,.

CASPER: A Large Scale Spontaneous Speech Dataset Powerset multi-class cross entropy loss for neural speaker diarization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.053460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.053460Z digest=sha256:c13e6541df5effc8be1baeb4014ff6b6e731ee7290dde06077a365f2e2d99b8a

Observation 93064c8b-ec03-4c01-b10b-2f179de41d11 · outbound

This paper cites End-to-End Neural Speaker Diarization with Permutation-free Objec- tives,.

CASPER: A Large Scale Spontaneous Speech Dataset End-to-End Neural Speaker Diarization with Permutation-free Objec- tives,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:11:49.495203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:48.390178Z digest=sha256:98054b962ff131550afd9659665116bbbc89223c1e1ebc2c42a32379e7c9997c

Observation 4d4c8745-b6ff-4949-98cc-6fd0a86d2b12 · outbound

This paper cites TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context.

CASPER: A Large Scale Spontaneous Speech Dataset TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.497846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.497846Z digest=sha256:ec42cdfb7424a2eaa4a00ca5fe4976cfa4b923a8a01e8954b2be7de4c77d7819

Observation b51f5b80-4fd2-4e5b-aff6-46b514b89da9 · outbound

This paper cites MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection.

CASPER: A Large Scale Spontaneous Speech Dataset MarbleNet: Deep 1D Time-Channel Separable Convolutional Neural Network for Voice Activity Detection

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:49.008646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:11:48.666663Z digest=sha256:8e73212b4ea267369c899130fee17e02646ff3354b58d7f2ca2648b30b58c5bd

Observation e119618a-71ac-450a-aeb4-c03b60e10a78 · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

CASPER: A Large Scale Spontaneous Speech Dataset NeMo: a toolkit for building AI applications using Neural Modules

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:48.308009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:48.308009Z digest=sha256:32eee2b661bd5caf3bbe01e1135a398d783f48f34870ed4f849ef2c0f83cfd86

Pith citing papers

No inbound Pith citation observations are available.