Pith. sign in

Paper Citation Record · LEDGER

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2506.01365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01365 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:11.832886Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:51:09.479147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:51:12.675415Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy26
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef798b6f-3a92-45cd-b13b-cb6e7ab9bd21 · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.275042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.419307Z digest=sha256:045006b1ea7a276a55392b2ecc2087214db75d21ad243c6789fc9cc8258df656

Observation 939a4c94-c33e-41bb-a9eb-1d76b740d270 · outbound

This paper cites Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:51:12.791241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.479147Z digest=sha256:12348be450bf42b1919e78d94af278c1c224715fe50c3d2d6b2cb30df4c625c2

Observation 7d51e26a-34d5-40f0-8512-f376cac3e2b3 · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.248660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.538212Z digest=sha256:b7275432d279b0e7a2e55a6fc814b7b9e28a7b5ecec741a50582c8160d4c6ccc

Observation 37bfde91-3d9f-43eb-8622-27d6fb12c8ff · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.220498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.594732Z digest=sha256:16f047b16d464a487e0fed52a99478ef9bac006f63a9948395898954e93b5600

Observation 12d97d49-01a5-4d69-842c-ece9c0f60556 · outbound

This paper cites an unresolved cited work.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:51:16.192287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.640441Z digest=sha256:f4e72955557897d9ef070233fd7f528e993c1bc4609f4b213cf9f7b3a6bae8f1

Observation 833c88b7-5636-4adb-96dd-659c7764811d · outbound

This paper cites MFCC vs PTM Features Pre-trained model based features have proven effective for various speech tasks, including V AD.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion MFCC vs PTM Features Pre-trained model based features have proven effective for various speech tasks, including V AD

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:16.162184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.706274Z digest=sha256:3a7464fe6614a870ccf2ba2751e4f28e953ddbf0c6fc0cfabb70ba8c0c8d5381

Observation 49fb3993-86dd-45c4-a184-64e4fb714be2 · outbound

This paper cites Dataset and Evaluation Metrics We conducted all our experiments on three publicly available datasets, i.e., AMI, Callhome, and V oxConverse, to ensure do- main diversity.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Dataset and Evaluation Metrics We conducted all our experiments on three publicly available datasets, i.e., AMI, Callhome, and V oxConverse, to ensure do- main diversity

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:16.136028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.750716Z digest=sha256:709bf8f57dc2df8c872fb7fc22e131f4642e5284c8ed49a81c33c0e591237427

Observation 3ef24849-27fb-4d4d-8e30-9ecf004c37ac · outbound

This paper cites A survey of convo- lutional neural networks: analysis, applications, and prospects,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A survey of convo- lutional neural networks: analysis, applications, and prospects,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.879696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.213444Z digest=sha256:07c3e06075d4ae528e0febc4765149578f60293e464da09284f8c44a8c655688

Observation d398ae63-8724-4b6a-9f6b-eb80523e5a3e · outbound

This paper cites MFCC vs PTM Features First 3 columns in Table 1 shows the performance of V AD with individual features in terms of DER, FAR and MR.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion MFCC vs PTM Features First 3 columns in Table 1 shows the performance of V AD with individual features in terms of DER, FAR and MR

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.789229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.860129Z digest=sha256:9a4a464fdad54e8f0e456879e2935e54dfea79fa8b45a6112702b4cebbffef42

Observation eec9b035-31f0-4487-b22d-453ef6f20906 · outbound

This paper cites Our experiments show that simple fusion methods like addition and concatenation consistently outperform the more complex cross-attention mechanism.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Our experiments show that simple fusion methods like addition and concatenation consistently outperform the more complex cross-attention mechanism

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.646275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.900221Z digest=sha256:157ca3a1d3a421296ca5aecdfa9cd8682eccf0acc367829dfa5827d70807220f

Observation b7cf4512-d6fe-4bc2-90b6-bc82a25d59b4 · outbound

This paper cites Temporal modeling using di- lated convolution and gating for voice-activity-detection,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Temporal modeling using di- lated convolution and gating for voice-activity-detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.569898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.938166Z digest=sha256:dabb9e230fb716ef52a798d4d3973be0a10a10ed13b4516a1c3b082cef74aa83

Observation fd8a25d1-d16d-4e32-bb0e-4e3e79db3956 · outbound

This paper cites Wavoice: An mmwave-assisted noise-resistant speech recogni- tion system,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Wavoice: An mmwave-assisted noise-resistant speech recogni- tion system,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.488071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.973317Z digest=sha256:01e5127fbb5462632b720f2f29a9edc59ab4c9764d2a726a775c70b2b197f7a5

Observation 8222dd9d-4393-4bb4-a8cb-5ce5f7d8b7f2 · outbound

This paper cites Profile-error-tolerant target-speaker voice activity detection,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Profile-error-tolerant target-speaker voice activity detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.321383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.016834Z digest=sha256:23040a39ed021be18998670d1d332e1d0b1d21269b3af0d99878f9e29236bf8d

Observation 679a867e-043f-416a-8b01-a477f1d0aee5 · outbound

This paper cites Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:12.535791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.058899Z digest=sha256:5b5be56b1c72191b3e0a0087566239ab86baaa7d6fb4060dc00172444fab42a8

Observation e10244b7-7dfe-4a9b-b0ec-2a8203743f11 · outbound

This paper cites Unveiling the state-of-the-art: A com- prehensive survey on voice activity detection techniques,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unveiling the state-of-the-art: A com- prehensive survey on voice activity detection techniques,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.185183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.100937Z digest=sha256:eeb9ad352bbabea5c0db0b210eb588b71e0a5df5e49e66d1091d630b4cf11d90

Observation 783b5cb6-c1da-48ec-9972-8751fff5b9c9 · outbound

This paper cites V oice activity detection: Fusion of time and frequency domain features with a svm classifier,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion V oice activity detection: Fusion of time and frequency domain features with a svm classifier,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.036588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.141767Z digest=sha256:ad011c0d7c70d5ee106eff81bcd76612aa0b4e81ac35764a8adaca924d063c44

Observation 5cc4d5a4-f4aa-4fa3-be5d-d4802a5ec1f3 · outbound

This paper cites Analy- sis of derivative of instantaneous frequency and its application to voice activity detection,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Analy- sis of derivative of instantaneous frequency and its application to voice activity detection,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.979766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.179305Z digest=sha256:a9b2c468cb888b5b4078cc1b9332f8d0222783ce9c9f21a52262f7ba3400933c

Observation 65f42d73-cb2d-4211-8c63-256f6bbaf333 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Robust speech recognition via large-scale weak supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.788432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.788432Z digest=sha256:c6322293ff17747e2d104bff295c8787619a93e05bc69eb63442aeb78730312a

Observation f763f485-49cf-42f7-932f-47987f604b98 · outbound

This paper cites A review of recurrent neural networks: Lstm cells and network architectures,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A review of recurrent neural networks: Lstm cells and network architectures,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.767898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.265665Z digest=sha256:22687e83c98bbf56c9342de1942204156f2215d2d70cc5054b3661996e2613c5

Observation 4266d00f-cbff-4121-94cd-08e422b1343e · outbound

This paper cites A hybrid cnn-bilstm voice activity detector,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A hybrid cnn-bilstm voice activity detector,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.698211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.304986Z digest=sha256:c99cfbc963b782251f43c0840d1454a31a94ea3dfe66b009cfec4bab8f2a16df

Observation aa3820dd-fd68-4e94-998e-3ed5b946acf6 · outbound

This paper cites Feature learn- ing with raw-waveform cldnns for voice activity detection.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Feature learn- ing with raw-waveform cldnns for voice activity detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.505567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.357047Z digest=sha256:71c0c95be3a45cc6699ff64fa512d2b135fa026839e58597a98d8ef214a37299

Observation e16e1dc1-7fb9-485b-80d5-66fac811bdc7 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.407683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.407683Z digest=sha256:dc02b3492da9e793f39c716fa3aac433995f45049ed1f30df26d8c7941991074

Observation 867f4b01-b5f4-4431-988b-1c5c344019ba · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.456958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.456958Z digest=sha256:d701c505da44ad8dbf62b5691fc07bd6fbf22473cd7fac7e809ea0357a39f3e5

Observation 2bec1e50-8035-4bad-a19a-a50b21c65200 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.503910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.503910Z digest=sha256:331f6e4ddf4b3727d2f0486437dbfd090c57bea8e53aaf90527b4cb895df4252

Observation e161ca1d-c447-43eb-a71c-267d0ee7f257 · outbound

This paper cites A closer look at wav2vec2 embeddings for on-device single-channel speech en- hancement,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A closer look at wav2vec2 embeddings for on-device single-channel speech en- hancement,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.316717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.585637Z digest=sha256:2b048cb074d74b76d0c5d3c25d0504e90bc9f7dd40b8f33dc0b3057ed34100bd

Observation 13345372-e7a2-4b62-a873-1fdfa45bfc5e · outbound

This paper cites Unispeech-sat: Universal speech rep- resentation learning with speaker aware pre-training,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unispeech-sat: Universal speech rep- resentation learning with speaker aware pre-training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.190460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:10.643292Z digest=sha256:72ba9224b2af8e1076e0a59c9bbc64ea6b665a12b64a6137f88f6e9048e43c8f

Observation 775471c7-0d58-4ada-adc7-fc7114dd8980 · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Scaling speech technology to 1,000+ languages,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.710355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.710355Z digest=sha256:1b5433cea65a431d664ad48858e3a5effe45607c6946690b5c33a31bb3abe5ba

Observation b353c641-69c0-46cc-baa6-6a632f777a94 · outbound

This paper cites Spot the conversation: speaker diarisation in the wild.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Spot the conversation: speaker diarisation in the wild

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:12.046250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.586294Z digest=sha256:bc6ceac452c93c2df812d4c4ab0edc46370097e0c14e2995bd051399d66d2840

Observation 70e2a298-5d4d-4288-8bcd-cfa4935a237b · outbound

This paper cites Base version checkpoints are considered for wav2vec 2.01, Hu- 1https://huggingface.co/facebook/ wav2vec2-base BERT2, WavLM3, UniSpeech 4and Whisper5.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Base version checkpoints are considered for wav2vec 2.01, Hu- 1https://huggingface.co/facebook/ wav2vec2-base BERT2, WavLM3, UniSpeech 4and Whisper5

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:15.960782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.817109Z digest=sha256:af02859c338b89990898a8441842ead8fc160ee9c9632c5ff916f937a13ae0dd

Observation 53a273fc-365b-4118-93fa-4e73e37f9a8d · outbound

This paper cites Enhancing whisper’s accu- racy and speed for indian languages through prompt-tuning and tokenization,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Enhancing whisper’s accu- racy and speed for indian languages through prompt-tuning and tokenization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.873019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.873019Z digest=sha256:5b8c620631e5180adde14b81fdd2431d8fa8b5034d501a2abf57c10536281f89

Observation 06ccb6e3-9ae4-49cb-a111-d5754034baee · outbound

This paper cites Self-supervised speech representation learning: A review,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Self-supervised speech representation learning: A review,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:10.961151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:10.961151Z digest=sha256:2e600c919d49a63a44666beb4bad378fb2dbadce4e771f5bd25441ba30ba4c9e

Observation 620f9e5c-8bd2-4289-90f8-8bc9f2f8eaeb · outbound

This paper cites Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:11.040230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:11.040230Z digest=sha256:d19264e0d32026001c77381eea4d70b2e63ffadba44414342a9b03de96dbc3f3

Observation 6c1ee7fe-557f-4c6b-880a-fb3f624a8a3d · outbound

This paper cites Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:51:12.273028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.101007Z digest=sha256:067f424cd74f9dcb116d967476a6fbd8808f35b37908f9f0afa91e8b87fcd008

Observation b3150781-14ea-40fe-8a97-74a8b0ef0d2f · outbound

This paper cites A transformer-based voice activity detector,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion A transformer-based voice activity detector,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:14.040176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.169108Z digest=sha256:c85b9a16043b36365ea155e300714ac5521def59c50f527e382a4e9916cd3f8c

Observation b382a52b-bbd2-47fe-b586-6774fcc0acad · outbound

This paper cites Multitask detection of speaker changes, overlapping speech and voice activity using wav2vec 2.0,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Multitask detection of speaker changes, overlapping speech and voice activity using wav2vec 2.0,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.742840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.271319Z digest=sha256:3daca22bf1aa47302f75971973f2bd31d9139f830ee4be05d45016932936a1b7

Observation b5687662-d01c-41fe-a795-98a53b41442f · outbound

This paper cites Feature extrac- tion using mfcc,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Feature extrac- tion using mfcc,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.589638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.374716Z digest=sha256:94161899bd0de8dea82ab532c9e448fcb95580426e770ab4488b3e99d35516bf

Observation 4b3137c4-f96b-4e44-ad1f-7748b7e574d7 · outbound

This paper cites Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.459115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.439124Z digest=sha256:aa96caa2fea2f9d5fe80801ee8e9250df0f25807e455924b75e3bd4923b2aa12

Observation d0a65c65-e5e5-4e6c-87d4-a2e1b5cc3b28 · outbound

This paper cites 2000 nist speaker recognition evaluation,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion 2000 nist speaker recognition evaluation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.335994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.507324Z digest=sha256:51592f423e1881221178d46e754fe24171b6bcbb322c147f1b741dd37fd77f57

Observation 46a94862-248c-4110-ad59-a7f343668c0a · outbound

This paper cites Pyannote. audio: neural building blocks for speaker diarization,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Pyannote. audio: neural building blocks for speaker diarization,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.227135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.666189Z digest=sha256:c7dd9d5f9da2dce15bd08d476a7e2b19427d7c8a6348d187dbf927fbf8157966

Observation 98c46501-02bf-4a18-b6c4-d86355140ca6 · outbound

This paper cites End-to-end speaker segmentation for overlap-aware resegmentation.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion End-to-end speaker segmentation for overlap-aware resegmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:11.727872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:11.727872Z digest=sha256:b921029dc652d71d05f87ddf845981e1f14297ee573286393ff6918a6f1bccb7

Observation 8372bbce-579c-4296-9b7d-c874a795c96a · outbound

This paper cites Speaker recognition from raw wave- form with sincnet,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Speaker recognition from raw wave- form with sincnet,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.170204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.759953Z digest=sha256:d571b7ebae414153b9ff9cd1304e63e35f0b0140adaaa68e321eda703ba3e18a

Observation a3c4611a-7a41-452d-b56e-3aa0e059262c · outbound

This paper cites rvad: An unsupervised segment- based robust voice activity detection method,.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion rvad: An unsupervised segment- based robust voice activity detection method,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:13.021646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.789628Z digest=sha256:1151ca1df67f14df71a576243603591ddab65dc3831fc94a2a5d5062ee2b4938

Observation 49dbc3c4-1c47-4149-b9ef-c03473418d92 · outbound

This paper cites Boosted deep neural networks and multi-resolution cochleagram features for voice activity detec- tion.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Boosted deep neural networks and multi-resolution cochleagram features for voice activity detec- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:51:12.933562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:11.832886Z digest=sha256:9a07e7b427df708142b18962db90b8a1d10506e7b18f589098011ffcfe235f1b

Pith citing papers

Observation 939a4c94-c33e-41bb-a9eb-1d76b740d270 · inbound

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion cites this paper.

Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:51:12.791241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:51:09.479147Z digest=sha256:12348be450bf42b1919e78d94af278c1c224715fe50c3d2d6b2cb30df4c625c2