Pith. sign in

Paper Citation Record · LEDGER

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.09792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09792 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.812385Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.622698Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:46:04.891505Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc2de33d-730c-49d9-8dc4-cc305f4e20d3 · outbound

This paper cites Most existing studies focus on improving the audio-visual fusion mechanisms [1, 2, 3, 4, 5] or addressing visual cue-impaired scenarios [6, 7].

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Most existing studies focus on improving the audio-visual fusion mechanisms [1, 2, 3, 4, 5] or addressing visual cue-impaired scenarios [6, 7]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.469713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.617264Z digest=sha256:13069b767cc475838b0920102ac254c41ee2e5f3cd7af4d5b5ecbc7e605fefd4

Observation bd51abf4-a74b-46fe-be2e-d77959f7cc4a · outbound

This paper cites Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:04.897605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.622698Z digest=sha256:06b23d21ab2ca6b20eb8106a9e3513a0f5b90808156b1baf29953bb5be929810

Observation bdec9dcb-8ec5-42ea-9e00-963626b2f9bd · outbound

This paper cites Dataset In this study, several experimental settings are considered: •Training Set:A two-speaker mixture training set is simu- lated following previous work [1, 6, 2, 3].

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Dataset In this study, several experimental settings are considered: •Training Set:A two-speaker mixture training set is simu- lated following previous work [1, 6, 2, 3]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.453563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.627656Z digest=sha256:1d6c453fb3143c4631054ab90645c816d8c1b060c676d6589076b2b25f4d8a12

Observation 6dfff320-741e-4807-8bea-7aaa4718b772 · outbound

This paper cites an unresolved cited work.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:46:05.438722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.633169Z digest=sha256:c558e624688e3a8d152b1ea6df8583518b8504c961933568a751dccd6ce625f7

Observation 715afddc-d228-491b-bd38-b22495bee4f9 · outbound

This paper cites Full occ.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Full occ

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:46:05.423908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.638070Z digest=sha256:6af5be7200fb21ed1fed524d36be5d82c92158bcf632d52b87ebcdc910d00863

Observation 554b79fa-87ed-4956-9e48-178c2cbd62e7 · outbound

This paper cites an unresolved cited work.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:46:05.408576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.643723Z digest=sha256:52ad2c04d5d9c96bd7dad7fd8ed2cea446f7f2439af78583a5abab3a67968324

Observation f77eb3b3-5e62-42a7-a8a1-15d631fae36a · outbound

This paper cites 62401377, Shenzhen Sci- ence and Technology Program (Shenzhen Key Laboratory, Grant No.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction 62401377, Shenzhen Sci- ence and Technology Program (Shenzhen Key Laboratory, Grant No

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.393041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.648600Z digest=sha256:006c7f9db2476c97285a2c00e305209d05d5aa61136ea0399655ebf209c2dbbb

Observation 277ee2e9-90dc-421c-8b38-bbdb74a5141e · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Muse: Multi-modal target speaker extraction with visual cues,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.377153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.653741Z digest=sha256:3bd221e555550eb2e4fe77aed48efb3283a1907ce75765f358b7fa02bfe2a490

Observation 5ecfee67-d124-4cc9-b4a4-b30340e60df6 · outbound

This paper cites Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.361888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.658898Z digest=sha256:947b06085dc4d86501ccf59bf9305fcd2ceddb6aafb414c0cd37a84c452e058b

Observation dc1f0523-48b4-4580-b3cb-863bb54eea93 · outbound

This paper cites Avhumar: Audio- visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Avhumar: Audio- visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.347080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.663935Z digest=sha256:cf49dbee7ff2f24bb77a9c5f9611d55e858dd3ebed9cee897f09d87c9e0deb55

Observation e9d1b61d-adb8-4864-84af-404c35a7c337 · outbound

This paper cites Target speech extraction with pre-trained av-hubert and mask-and-recover strat- egy,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Target speech extraction with pre-trained av-hubert and mask-and-recover strat- egy,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.331829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.669226Z digest=sha256:66e90fb86a7291dc5b2ba01dcc2f1b39a11ed39ada001965a3036cb18ad6d7df

Observation 9f2fce59-8df6-46fc-a264-defe611a2a0e · outbound

This paper cites c 2av-tse: Context and confidence-aware audio visual target speaker extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction c 2av-tse: Context and confidence-aware audio visual target speaker extraction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.317657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.674131Z digest=sha256:af3b2dbd0b4876df9af056325915ecfc615957cf0ab7451e4408dfd2e63b1881

Observation fb5b5b21-c426-4760-ba87-16c0ddf158b4 · outbound

This paper cites Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.303072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.679148Z digest=sha256:d467ebe48022bde273a47166135207e92433103004d365722cf96286809bb35d

Observation e0ad3710-a4cd-4595-8c2f-a4944678ed28 · outbound

This paper cites Restoring speaking lips from occlusion for audio-visual speech recognition,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Restoring speaking lips from occlusion for audio-visual speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.287472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.684022Z digest=sha256:d82f5d53e6b15b69af815f462ea75eae74f335f4f3a302f04a66456cb03585b9

Observation 1d9e7c30-6caf-405d-8f3f-df8ed062ab1f · outbound

This paper cites Semantic en- coding during language comprehension at single-cell resolution,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Semantic en- coding during language comprehension at single-cell resolution,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.271997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.689743Z digest=sha256:9f62fe0b28ac7dd7ff208911ce4be16b404fb4b690580bfe527f2c82548e8eaf

Observation 747ba67d-bf1f-43dd-bba5-3db1cc01ab66 · outbound

This paper cites Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.256584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.694126Z digest=sha256:2a3df84c4ba35ba854dbde19b46e4e84d6c4eacbd9fcf041185bd32cae74cad9

Observation ce3d5104-6aa5-4de9-8092-a6289ee7cd5f · outbound

This paper cites Large language model can transcribe speech in multi-talker scenarios with versatile instructions,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Large language model can transcribe speech in multi-talker scenarios with versatile instructions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.240980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.698793Z digest=sha256:d040701c94b1c30483d552f168d77672f6777bb4a3c814f9e4812b4c329adec8

Observation e133f23a-cfdd-4793-945f-bc8822d23e76 · outbound

This paper cites Target speech extraction with pre-trained self-supervised learning models,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Target speech extraction with pre-trained self-supervised learning models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.225837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.704344Z digest=sha256:583e14cd75df35a7b60b8a8f8743609018e63c497ebc2a7ab481ce6530379eb8

Observation e15dd3bf-0e42-49e5-aeb5-e76d5bcad2a6 · outbound

This paper cites Probing self-supervised learning models with target speech extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Probing self-supervised learning models with target speech extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.210381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.708724Z digest=sha256:01824a9f8b9dbbd46f9682fb33c8da4ad8e823080903218ab6151add025baff9

Observation 78c595a9-fa23-41e9-b7e4-e1ed7ad327c5 · outbound

This paper cites A large-scale evaluation of speech foundation models,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction A large-scale evaluation of speech foundation models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.194831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.713437Z digest=sha256:213875abd7e6120a3078ed50b5ea277babc90d0de23613ab3113027af9803ad4

Observation a571c4ce-40a4-4f71-b03d-16c5d536b767 · outbound

This paper cites Transferring knowledge from large foundation models to small downstream models,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Transferring knowledge from large foundation models to small downstream models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.179377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.718269Z digest=sha256:b683b7fccdb302608cb9b19b1637fe1dc926665b08cad7fe785a2fad98729810

Observation 62790262-403f-4144-b32a-f3b86abd677e · outbound

This paper cites Knowledge transfer from pre-trained language models to cif-based speech recognizers via hierarchical distillation,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Knowledge transfer from pre-trained language models to cif-based speech recognizers via hierarchical distillation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.163405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.723757Z digest=sha256:3429fb5c6a0551df49e3a25c3edb5b2b054604e7523f945a7eeae7b805590cb2

Observation b915031b-a18a-458b-90d7-4e2dd441ec03 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.146935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.728231Z digest=sha256:f03342bfd0448ec34bd895c87d88705bf834ab4003ffc479292290919647cea8

Observation 97115ffb-4df0-4dc0-9975-36526792fae7 · outbound

This paper cites LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.733824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.733824Z digest=sha256:6e0ae177347b8cf5a58e8cabd070473def29a1ec6beccb08d0d03b1f01f61076

Observation 2ddd5d4f-bc6b-4752-958a-3c196629d5fd · outbound

This paper cites ALMTokenizer: A Low- bitrate and Semantic-rich Audio Codec Tokenizer for Audio Lan- guage Modeling,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction ALMTokenizer: A Low- bitrate and Semantic-rich Audio Codec Tokenizer for Audio Lan- guage Modeling,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.131868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.738953Z digest=sha256:ce93451a49d67b938870d7159a2e6eba1c47e16591cec7584c921e0971ea89cf

Observation 4a37c920-5726-402b-9940-466c6bb4effb · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.116365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.744419Z digest=sha256:5778119079559c9294fd092e6065f9e6a7fd2de85af86b112dc77d31e8db1ea9

Observation c2378fd2-e3ac-4f81-96a9-249ff3ea9aac · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Roberta: A robustly optimized bert pretraining approach,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.100660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.749038Z digest=sha256:327d005dc1efe71e94dcf39300d06b1bff6bd0ededac6cfa8dc68e3f5869c2c8

Observation cccba63c-671b-4ce4-a72c-1fa36577a663 · outbound

This paper cites Separate in the speech chain: cross-modal conditional audio-visual target speech extraction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Separate in the speech chain: cross-modal conditional audio-visual target speech extraction,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.084425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.754727Z digest=sha256:f38f87998b7c05f2035426999520299f92ea0e16a25cb926dea14e93e67f3278

Observation 9e63d60d-5eac-43ff-b8f5-113ff6fe942a · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction V oxceleb2: Deep speaker recognition,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.759101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.759101Z digest=sha256:6e4510586e7d1dfd78511e7f886179ebe57a6fd60051486f0a09480b8c6b1e4b

Observation cf384b2e-344f-4da7-a016-6670a52231d2 · outbound

This paper cites Watch or listen: Ro- bust audio-visual speech recognition with visual corruption mod- eling and reliability scoring,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Watch or listen: Ro- bust audio-visual speech recognition with visual corruption mod- eling and reliability scoring,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.058171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.763651Z digest=sha256:9adfe01bee223b2b600805bd2fd62bcd756d19a95342cf486b92656c8a3e72cf

Observation 8f3733b2-e3a8-4f34-9193-24faa3e99dcd · outbound

This paper cites Lrs3-ted: a large- scale dataset for visual speech recognition,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Lrs3-ted: a large- scale dataset for visual speech recognition,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.042463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.768090Z digest=sha256:d2be8941aab8259653b95f3741f81f67b9e2f94d0411af310b8e9178247c749b

Observation df70e485-3872-4cd2-b00f-2584726a6d38 · outbound

This paper cites Sdr – half-baked or well done?.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Sdr – half-baked or well done?

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.027505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.773088Z digest=sha256:1ba68af64b972e02925618ab543f7ec808011131bbb417291ea13056593f67ac

Observation d35c370d-25fb-4aa2-92c0-6f0d5cbb1174 · outbound

This paper cites Single-sided Real-time PESQ Score Estimation.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Single-sided Real-time PESQ Score Estimation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:04.857164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.777432Z digest=sha256:17ec07687ea26f04a0b8e4a4dbc08e1a2ad491359d7432b9187845ef6d35a936

Observation d0040916-1427-4b42-8027-a37154c8e2ae · outbound

This paper cites An al- gorithm for intelligibility prediction of time–frequency weighted noisy speech,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction An al- gorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:05.011753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.782208Z digest=sha256:3a31ea44676fe3d156e67975768b9a06ecca4d75c1f3ee22a3f7b2a28a71ea0a

Observation 2feebede-1b54-43ca-a246-0f0eb293d75e · outbound

This paper cites SpeechBERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction SpeechBERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.995934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.787637Z digest=sha256:32a1cda47d625bbe351b493df00bff1687388bb144fb39c022dc6c8f61ffe072

Observation e82f49b8-00ac-4725-ad93-a1c3d3503b1e · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction How should we extract discrete audio tokens from self-supervised models?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.979381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.792441Z digest=sha256:0982b4191a0dea925ebce64004b851b17515e3064bdf6b891280b918bf2d1646

Observation 673af5ad-5b81-4606-a91d-8a5f8194060c · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.962900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.797897Z digest=sha256:3ca1d74ce784be38932f13d6ddb42e73ce27074fa502b1291c5693dba7603680

Observation 5426972f-cbe5-4462-b88d-d586d4b961b1 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.946664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.802382Z digest=sha256:e68461965ddea11c5f140abd952be1c0a48aba65b7cd6eb7a501a1766f44935e

Observation c5af5730-58c1-48f1-870e-a77dab4ef6d3 · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.929645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.806966Z digest=sha256:edb580c63edc7f9e4fdd746b911368efcee15daa56d8c8c93321ee414a33062d

Observation 4ac4cd68-4834-42c4-8b31-c4753ab3a989 · outbound

This paper cites Intuitive multilingual audio- visual speech recognition with a single-trained model,.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Intuitive multilingual audio- visual speech recognition with a single-trained model,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:46:04.913816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.812385Z digest=sha256:f9218cc8f27011efc0ce17eb244c5b7aa87fc88395693eddeced01293a495ed0

Pith citing papers

Observation bd51abf4-a74b-46fe-be2e-d77959f7cc4a · inbound

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction cites this paper.

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:46:04.897605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:46:04.622698Z digest=sha256:06b23d21ab2ca6b20eb8106a9e3513a0f5b90808156b1baf29953bb5be929810