Pith. sign in

Paper Citation Record · LEDGER

Optimizing Speech Multi-View Feature Fusion through Conditional Computation

As of 12 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2501.08057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08057 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:13.120441Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88c42a83-1785-410c-8ab9-fb90274451a7 · outbound

This paper cites Introduction to digital speech processing,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Introduction to digital speech processing,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.676357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:12.952741Z digest=sha256:70d2d6d85847fe4e3ce93529f26e5af52acc80014364819082c84888f63b4002

Observation 1937764e-8328-45c1-90e1-9d91f823195f · outbound

This paper cites Learning robust features using deep learning for automatic seizure detection,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Learning robust features using deep learning for automatic seizure detection,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.664268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:12.957265Z digest=sha256:5b178832d7cf230142f782872577636ba15d3ff3234d3725d412d7c251532373

Observation de5af010-494a-43ef-a812-c3d97c274977 · outbound

This paper cites Listen, Attend and Spell.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Listen, Attend and Spell

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.962138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.962138Z digest=sha256:7b00f26427ed217aae5f9e422e013de0522235372fcee6239a34ecbfa487c69e

Observation dcc3f86b-bba6-49b1-bb64-2770c4ced7f2 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.967080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.967080Z digest=sha256:12377f838edcfb33ae8f780f9e171da0d736a595a58ee3759dc37b72b6782b4e

Observation 1be994b9-d8e9-43f4-af09-481ad06a238a · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.652416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:12.972165Z digest=sha256:0baeb0f1432e571226741d1743eb34c49f01b02ce109b8a95bbce457fd4dcc39

Observation 134a1727-a30e-4c74-9574-32c86366f7ce · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.640340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:12.978136Z digest=sha256:0deea3a78df450e148c219ab9109212da2b22c0c3e93df684e1b7383034ea55a

Observation 809aaec6-045b-47cc-a99c-3980a4e25ae2 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.628277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:12.983242Z digest=sha256:c40684086411f6a675892edd7ed00c947d8843e681e2e49a20796123ca6b1a6c

Observation b4f68caa-92d2-4541-a491-63382220b722 · outbound

This paper cites Exploration on HuBERT with Multiple Resolu- tions,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Exploration on HuBERT with Multiple Resolu- tions,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.616277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:12.988042Z digest=sha256:8631cc1a0f6799d11827e87ffcf1bc06cdf207659dd80bb898eff171dcc9a17e

Observation e9db1f10-40da-4977-beff-caa965cdb38f · outbound

This paper cites Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.992523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.992523Z digest=sha256:f301d5a999b6c20c5018718e6f6319242a6c937b1891fad80977e4abd35da06d

Observation 7b03d0e0-e681-4756-a5ba-ad0797c0c66a · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.996937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.996937Z digest=sha256:b37c242ce2ed0a5e337fbee0f20e2e1c01018344f402a7a029660f52035951b9

Observation 79750105-313c-41cc-81e7-5be56738d061 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.001744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.001744Z digest=sha256:a1d5fc81959db26705bece54d4c7d1d2e9cfa8f7098878f9355c86437fb99fc4

Observation 4ee0998a-0bc6-43dc-9c7c-0a006c6fc662 · outbound

This paper cites Multi- view information-bottleneck representation learning,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Multi- view information-bottleneck representation learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.604733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.007235Z digest=sha256:f2bc78118e1869b153447f5ef30e1358567000c77937fd493baf70e2d284a649

Observation 97adc650-762e-429e-830a-ce6f1a778a6f · outbound

This paper cites Deep multi-view learning methods: A review,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep multi-view learning methods: A review,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.589761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.011623Z digest=sha256:3684e9a6cac0872ca78da60840e4d962665036e63bd55c1268048285631b1d3d

Observation 3cf63879-175c-4448-8c8d-431e7a867dbe · outbound

This paper cites The Platonic Representation Hypothesis.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation The Platonic Representation Hypothesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.016293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.016293Z digest=sha256:041da615a407b529d9f8883dec2cc6e487cde4a9410711778af1de7f2cd3b5ee

Observation 80e901f7-d97f-4080-a2ec-78a4f9ec1370 · outbound

This paper cites Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.574820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.023101Z digest=sha256:8913e86f848aae38127eb60ee09dd108615e7563713f2a71bf01a1d352a39f69

Observation 42892fa8-b830-4431-b653-a619d8cf2d38 · outbound

This paper cites Improving speech emotion recognition by fusing self-supervised learn- ing and spectral features via mixture of experts,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Improving speech emotion recognition by fusing self-supervised learn- ing and spectral features via mixture of experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.559020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.028995Z digest=sha256:f5d3adaa84ac6c7346204838c020d3cbc95c2a919d33e1055c7483698e22605a

Observation 7db3994b-4a62-4b3c-a152-892cfc5cc83b · outbound

This paper cites Deep learning of representations: Looking forward,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep learning of representations: Looking forward,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.544333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.035113Z digest=sha256:60cd350d4b847bbbeba5b8f8de6adc919cceaac352b7706a61ca6ed431967499

Observation d1f92b65-ee65-4a44-b930-b9bf5021c6f9 · outbound

This paper cites Depth-Adaptive Transformer.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Depth-Adaptive Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.040015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.040015Z digest=sha256:3a287219717d42aa070ffbcc8a1d1b467544532f1d21f2b55b3994f4a5e35b98

Observation 5f92795d-8300-4ff3-a32a-b55de0f7c3a0 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.048303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.048303Z digest=sha256:a817f5c85b849d8d9a62035ce7df946f084535ab80159a39989423e5e6aea7f4

Observation 1de2a624-7275-406c-99f7-68507eac3d18 · outbound

This paper cites Modeling task relationships in multi-task learning with multi- gate mixture-of-experts,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Modeling task relationships in multi-task learning with multi- gate mixture-of-experts,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.525266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.054186Z digest=sha256:f7225090999202a7c7a815954b7a35b1df6557d5880ee6acb1ebf281ea908c86

Observation 263ec31d-b428-4805-8b60-79b344d77f15 · outbound

This paper cites Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.059024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.059024Z digest=sha256:8b35bbcd8f38587837bd6994585ba81d70680062a46c4680c2ffe13faa75dedf

Observation fa28073f-17e4-43b5-b200-a7138ab9261d · outbound

This paper cites Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.064469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.064469Z digest=sha256:079de85eb0a4b0f7a1720c82860d1071bbd41cf6e008f3e8d123f6c95b7c3339

Observation 83c81a68-fcf4-4208-8a2d-4c9dba5f3425 · outbound

This paper cites Soft Alignment of Modality Space for End-to- End Speech Translation,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Soft Alignment of Modality Space for End-to- End Speech Translation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.505664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.069007Z digest=sha256:d3341fdd7ac8cda70392deeff30be8e2019122c5ed757e3977c2ff0b6cf2da2d

Observation 0c0b2c8b-78c2-4d45-8c40-a2ae4576bdef · outbound

This paper cites Deep residual learning for image recognition,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Deep residual learning for image recognition,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.073686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.073686Z digest=sha256:e57abaedd57f8bfc896e145753d0800071ba793703682fb2ecbbf62b58045e5c

Observation c91f3c40-8959-479e-8e60-55284b580e68 · outbound

This paper cites Progres- sive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Progres- sive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.472636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.078447Z digest=sha256:b2d09f93ec97708d855c8f148a899218937e01f111a278fb0093d33201ada451

Observation ae696b85-5208-455c-aa0f-73fe035f49be · outbound

This paper cites Gradient surgery for multi-task learning,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Gradient surgery for multi-task learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.459433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.085525Z digest=sha256:2ade74973c120955cfc88dad33b33457385fe88579bb379dc88efaa7e1568ec3

Observation 713c1147-2693-4839-b9f0-74d40a04d96d · outbound

This paper cites Must-c: a multilingual speech translation corpus,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Must-c: a multilingual speech translation corpus,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.445597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.095399Z digest=sha256:57ae83de8b132a720e65072ab0f36c38c6a62725f5e51d925fcbeed12595a016

Observation 71a5f744-1695-4b27-8df8-d5357121a6c1 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Librispeech: an asr corpus based on public domain audio books,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:13.427050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.101348Z digest=sha256:43033303ca3e4941d8ccce894add7a103b3c6263adc80bf323b085458d414da4

Observation 59eb1700-894c-4115-9694-734437d47877 · outbound

This paper cites Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Modality Adaption or Regularization? A Case Study on End-to-End Speech Translation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:35:13.209587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:13.106571Z digest=sha256:bc49bb6a348a809e4c76c5aa40370d89e370d258f21f3c6fc3edf023f0015b02

Observation 0b0fa77c-ad3e-4681-a185-8a1832ed8f8f · outbound

This paper cites Attention Is All You Need.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.113877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.113877Z digest=sha256:02380492583959c1800e2c5e82dc624921d07cd7f49abd52b6a786136bc9fedd

Observation 1d095de8-3b52-489a-81f0-adfc5bb1981d · outbound

This paper cites A Call for Clarity in Reporting BLEU Scores.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation A Call for Clarity in Reporting BLEU Scores

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:13.120441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:13.120441Z digest=sha256:7eb293d6c17ae1b162ef0855c18df0aef65d392ed1988f8d49a5c512a9778a2e

Pith citing papers

No inbound Pith citation observations are available.