Pith. sign in

Paper Citation Record · LEDGER

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2608.11587.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11587 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:58.527198Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:58.415460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:37:58.604758Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f19e0b3-41b9-4b0d-966f-c8e930a8b1d8 · outbound

This paper cites Depending on the label taxonomy, the task can be formulated as speaker diarization, vocalization classifi- cation, or a combination of both.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Depending on the label taxonomy, the task can be formulated as speaker diarization, vocalization classifi- cation, or a combination of both

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.861913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.410982Z digest=sha256:b0978cb61a3cc0589bb91592e8bf640ee328ce7350999c23913f4705f89bc3ee

Observation 85931b77-fc6b-41e1-b17f-d272322f871b · outbound

This paper cites Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:37:58.608986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.415460Z digest=sha256:533cd228df6df114e1fb327197275e0763b5e258a42a47386130bb85dc58e201

Observation 5613fc63-21f2-4f34-86c2-ab752e73caca · outbound

This paper cites an unresolved cited work.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:37:58.851409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.419450Z digest=sha256:c7f6a28757f8f0a1af6fc1fe311654134e06473398b24a55e7379b8133207aa1

Observation f893d402-5bae-476b-a70f-272e3268979d · outbound

This paper cites an unresolved cited work.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:37:58.840046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.423648Z digest=sha256:745e927a315b5a23a53e4d4d7cbdca65af42658a556f3408513a18e7cd8bdf5c

Observation e6dcbee3-c214-474a-a2db-b41e1dee46a5 · outbound

This paper cites an unresolved cited work.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:37:58.830443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.427581Z digest=sha256:e8353867c9ef8c1fda6a1f0f4d6de7d6668894cc463b42069574d8ead6aae8a7

Observation 340c13f6-98fb-4592-882b-7fa2a8ae33e1 · outbound

This paper cites w/o LoRA.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning w/o LoRA

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.819694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.432051Z digest=sha256:58c704b9c31bd0111c197cf96cfc22684e02cde03d6a75d0b2077d3812743cd5

Observation 9ef79544-396e-4c91-b6a9-22f6905adf36 · outbound

This paper cites In the current design, offsets are learned only for training families, and inference on new families relies solely on the shared tier tokens.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning In the current design, offsets are learned only for training families, and inference on new families relies solely on the shared tier tokens

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.808612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.436080Z digest=sha256:a731a7ddf2d8135c7e1920d7acf3040547df18c74320a4a0bec9b289a58e47c5

Observation be981798-466b-42d9-9a9d-cf5eda55c8a2 · outbound

This paper cites By combining a LoRA-finetuned Whisper encoder with structured speaker conditioning and tier-specific heads, the model supports overlapping speakers and framewise prediction.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning By combining a LoRA-finetuned Whisper encoder with structured speaker conditioning and tier-specific heads, the model supports overlapping speakers and framewise prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.797765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.440202Z digest=sha256:657ff7511acacea237001b01c521b9d931bf66d75adb777460c600a3458a9b07

Observation 066d5771-9987-4c7c-b5d0-0cfb4e7f08e6 · outbound

This paper cites All technical content, experimental design, analysis, and scientific contribu- tions are entirely the work of the authors.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning All technical content, experimental design, analysis, and scientific contribu- tions are entirely the work of the authors

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.786849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.443926Z digest=sha256:86a1dc516574ead7d89d81248c9a9f2124f287f620e91a3ebf105009a6e64ccb

Observation d39187cc-5b69-49fa-986e-d7bed726bb28 · outbound

This paper cites For the experiments presented here, we used the Delta System at the National Center for Supercomputing Applications through AC- CESS allocations CIS240417 and CIS250040.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning For the experiments presented here, we used the Delta System at the National Center for Supercomputing Applications through AC- CESS allocations CIS240417 and CIS250040

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.775753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.447533Z digest=sha256:7f8d6cfe7335b4a6c2de33c1e9849439222ae37ef183ff1b99fef2f8d6b5ba83

Observation d43b58f2-4d25-4582-9e8b-d163416162ab · outbound

This paper cites An open-source voice type classifier for child-centered daylong recordings,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning An open-source voice type classifier for child-centered daylong recordings,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.765244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.452612Z digest=sha256:088a1d2686f0b734cc1fc70d38d7910314e9558fe987aee72d213c08870b6c74

Observation 1f08f8ed-e446-422d-9de2-b712e32e509c · outbound

This paper cites Analysis of acoustic and voice quality features for the classification of infant and mother vocalizations,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Analysis of acoustic and voice quality features for the classification of infant and mother vocalizations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.753786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.456733Z digest=sha256:5c0231b9661cb277a92239523108e1ab7dbc161abad9b0380427f9f7f802f6ae

Observation 83c5e8d7-f678-4bad-9126-b99b602b2ce8 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.460817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.460817Z digest=sha256:191d4ad5a97b0beb8bc93d815a506e575718cecc362d1cf887c82597c04efd49

Observation 81970ea0-3654-40c7-ab9e-089bdb55826d · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.464621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.464621Z digest=sha256:7fb24f2ecf9b33de374e6b78ff9f9ab195c9539e407553228069f295cc2bddc3

Observation 7f723355-a9d6-45ef-a69c-6f5b0f6b1486 · outbound

This paper cites Ssast: Self- supervised audio spectrogram transformer,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Ssast: Self- supervised audio spectrogram transformer,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.730392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.468293Z digest=sha256:32c714e69da1dbffef40476dae905a7c40f632a865bfc7d1b68f726ff1d98955

Observation ca0306a4-6514-4e7f-bd3d-fe7c070f8be9 · outbound

This paper cites Towards ro- bust family-infant audio analysis based on unsupervised pretrain- ing of wav2vec 2.0 on large-scale unlabeled family audio,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Towards ro- bust family-infant audio analysis based on unsupervised pretrain- ing of wav2vec 2.0 on large-scale unlabeled family audio,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.719171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.472056Z digest=sha256:53b353d3cdd870c4426ea57cee233fda9965bcb0090c8c0ce496e64254089b71

Observation d1fc9b1f-8512-4744-945f-4c449bdaf826 · outbound

This paper cites Band- split self-supervised mamba for infant-centered audio analysis,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Band- split self-supervised mamba for infant-centered audio analysis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.707606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.475453Z digest=sha256:dc820c6d4a60f52fd6b72e8ba0ef7e88bcb037b416b8ed6463e378a56860f1eb

Observation fcab2698-f3fd-4943-b7ad-08bba330d823 · outbound

This paper cites Robust self supervised speech embeddings for child-adult classification in interactions involving children with autism,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Robust self supervised speech embeddings for child-adult classification in interactions involving children with autism,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.697393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.479443Z digest=sha256:de461581431fd02e418de87b0443b600dde3732a58a396daf7bfb24b3d5a021b

Observation 10861f42-de29-4d87-9f7b-df53877a8d01 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Robust speech recognition via large-scale weak supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.482950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.482950Z digest=sha256:2390da6bff48dd9aa1cdc4bdbdf0c853c2fb7d18198eb6fa2450f46aac46e33d

Observation 44225a6c-1bc8-40df-aeea-ee68f5961914 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.486408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.486408Z digest=sha256:dfab9c978fec4381b02e6ff17c40d72078e19b88be096d5756cdcbb795d8d5d3

Observation 9f4cbe6a-f1ec-437d-9ef6-05e396f4b027 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Lora: Low-rank adaptation of large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.489633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.489633Z digest=sha256:571645569814e4ee43569196b4e8523142d6fe718e0b7d1d8fc77cafee969af6

Observation 3699ebe0-0dce-4014-9e6c-6b81f38bd462 · outbound

This paper cites LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.493362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.493362Z digest=sha256:1a5249739931de5f1960f88baa640efc04630819123371361ee0f37115ca979c

Observation a4e021d2-ec4e-49c0-9e39-b22a74218af4 · outbound

This paper cites Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.497599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.497599Z digest=sha256:711d4dba2b9c6d1217a8768c00235940bdd51558547cab9eb3746743dc3f10d6

Observation fbbf2178-691b-4d7d-8502-b2cfebee8cde · outbound

This paper cites Sparsely shared lora on whisper for child speech recognition,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Sparsely shared lora on whisper for child speech recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.667830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.501666Z digest=sha256:27a0178b6af7f7d0b5a4fbb16e68aaaf2b7f0f03c9dca98cea85c57d213d17f7

Observation f7800dad-1039-4625-9f92-e59fc25be48c · outbound

This paper cites Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Whisper-at: Noise-robust automatic speech recognizers are also strong audio event taggers,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.657405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.504948Z digest=sha256:7cdf6dab6736cfeb097785c3043572868df30f6601d07dba397e6b66e2401d30

Observation 03a70b6f-972e-482f-b23b-f5fd29698aa4 · outbound

This paper cites Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:37:58.574313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.508242Z digest=sha256:99d7629c28733e680b07d28604547db9f9923ef1568dd6df8d6e94f52d69e556

Observation 262b03b5-6715-41b1-b411-9ad2b52a825e · outbound

This paper cites Data efficient child-adult speaker diarization with simulated conversations,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Data efficient child-adult speaker diarization with simulated conversations,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.512279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.512279Z digest=sha256:27aa07ac2c6fddeba2be625e7196f7670af594476a70a6e1da3e63fe084f1800

Observation 03d92929-91e8-4f51-8300-d15577639681 · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:58.515987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:58.515987Z digest=sha256:a56e2b0dabc48f82f58aef242e89b9a12932cb2b78aca82f981083ce19747a8c

Observation 2f633839-6a74-45d0-9798-85f6edf2048b · outbound

This paper cites Preliminary technical validation of LittleBeats™: A multimodal sensing platform to capture cardiac physiology, mo- tion, and vocalizations,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Preliminary technical validation of LittleBeats™: A multimodal sensing platform to capture cardiac physiology, mo- tion, and vocalizations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.641316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.519766Z digest=sha256:4042d513c16c0dbc657f85be43ceb0ce70ec675b6cffec7fd631d90a4d7f3fbd

Observation 173baa84-40d6-484a-aede-d591008a6fb9 · outbound

This paper cites Praat: doing phonetics by computer [computer pro- gram],.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Praat: doing phonetics by computer [computer pro- gram],

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.631037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.523640Z digest=sha256:de0b5c51ff5a1396bd6367b1a007b062a3c84fb16abb272ecc5b9e54fb64314c

Observation 8038aa1b-a8bd-48ef-8431-969e39e68630 · outbound

This paper cites Listen, adapt, better wer: Source-free single-utterance test-time adaptation for automatic speech recognition,.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Listen, adapt, better wer: Source-free single-utterance test-time adaptation for automatic speech recognition,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:58.620077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.527198Z digest=sha256:75dadd66f23b777666ae5a520e7449a467978f56a4e0629942b41461d3476e39

Pith citing papers

Observation 85931b77-fc6b-41e1-b17f-d272322f871b · inbound

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning cites this paper.

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:37:58.608986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:58.415460Z digest=sha256:533cd228df6df114e1fb327197275e0763b5e258a42a47386130bb85dc58e201