Pith. sign in

Paper Citation Record · LEDGER

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2506.02258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02258 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:54.084398Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc71a14e-b9b5-4e86-b62e-1eccbc0621ee · outbound

This paper cites Speech emotion recognition based on hmm and svm,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition based on hmm and svm,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.216794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.141074Z digest=sha256:cb04fa15f7cf86914c61df36693ecf53b6a5ae71afb6b496008cfa6f70ebe4a2

Observation c9b6b090-75e2-4ef8-97e9-15441a19ba65 · outbound

This paper cites Emotion recognition in speech using mfcc and wavelet features,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Emotion recognition in speech using mfcc and wavelet features,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.203011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.180343Z digest=sha256:25bf4621ef8217f7e3d1d773d3f83d5b330680a83c463307e67f07a628c8c13d

Observation 1b912b5f-c264-4fae-b6e5-828260b43db7 · outbound

This paper cites Speech emotion recognition based on feature selection and extreme learning machine decision tree,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition based on feature selection and extreme learning machine decision tree,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.190003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.236085Z digest=sha256:632d02c216501748194b689c80d5299378654c580491da61c71e5527e0762160

Observation 05cb934a-48bf-4892-b403-97ca842b7208 · outbound

This paper cites Speech emotion recognition with dual-sequence lstm architecture,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition with dual-sequence lstm architecture,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.173025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.283010Z digest=sha256:b90ca17f141a79ef1340e8fd273ed996ecd6beb09a5d3431367003bec54a8082

Observation 1f062ee6-db11-4b59-8db9-c7498806e1c2 · outbound

This paper cites Convolution neural network based automatic speech emotion recognition using mel-frequency cepstrum coefficients,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Convolution neural network based automatic speech emotion recognition using mel-frequency cepstrum coefficients,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.158322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.344930Z digest=sha256:2c6525d9461c1776b63450e07e37406fe4dc384b20dae604ee35d8db3b4a616c

Observation 8b4d47a9-4339-41fe-9960-235e3d22a499 · outbound

This paper cites Ctnet: Conversational transformer network for emotion recognition,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Ctnet: Conversational transformer network for emotion recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.146277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.398994Z digest=sha256:c62bfb61fab003632ba063a92af1b8e07362c7b5dd6153f09a5b495c4324f5ae

Observation f15749ac-0c24-496d-aae4-a5f7766a748f · outbound

This paper cites Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.459509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.459509Z digest=sha256:745196c30531acd166f467f469ee87081d7b985ba83343c797b783a27a5454a8

Observation 1228f4ed-326d-4082-b7cf-063425fb9bbd · outbound

This paper cites Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.540223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.540223Z digest=sha256:b51cc0bf84d098e208a28cc9603a21315ff0215dd1dab034a89ac48fc65348fd

Observation c2055c74-ac05-47b6-b767-cc3b7bec6d69 · outbound

This paper cites Adapting wavlm for speech emotion recognition,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Adapting wavlm for speech emotion recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.135147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.604361Z digest=sha256:b97d370585e26b0eaa223614caa50d60a3ccc4676fd37de37241cc050336d033

Observation dcf272e4-ab17-47b4-bd5d-66874f4c8801 · outbound

This paper cites Audio mamba: Selective state spaces for self-supervised audio representations,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Audio mamba: Selective state spaces for self-supervised audio representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:55.104324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.691205Z digest=sha256:965eb2cf75d74ed5f4ed2fab3ecb41d5345578df6a79de1180fa55b1e19282a5

Observation a954a949-eb11-4c71-91e0-39eab3549663 · outbound

This paper cites Speech emotion recognition considering nonverbal vocalization in affective conver- sations,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Speech emotion recognition considering nonverbal vocalization in affective conver- sations,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.977226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.700086Z digest=sha256:0bd4d36838ac1ae53ad4fa38c1cdb93489b1e983d017b81e365e995dc41c9de8

Observation 3d6631ce-bc9b-4570-868c-032edd60a33a · outbound

This paper cites Jvnv: A corpus of japanese emotional speech with verbal content and nonverbal expressions,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Jvnv: A corpus of japanese emotional speech with verbal content and nonverbal expressions,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:52.780585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:52.780585Z digest=sha256:ff03627913bdf4c364358ddce5f7ac76cee2df130e9bb403a0c2cc08ac914dbe

Observation 16b4b1d8-020d-4e38-8091-32195789b2fc · outbound

This paper cites Large-scale nonverbal vocalization detection using transformers,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Large-scale nonverbal vocalization detection using transformers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.797553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:52.917984Z digest=sha256:c1917d94cac1db3e5f50b307e4657728a408694cdc67d945243aeef62d6587dc

Observation f3947231-5db5-4fc1-a000-36366dc4d0f7 · outbound

This paper cites Investigation of ensemble of self-supervised models for speech emotion recognition,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Investigation of ensemble of self-supervised models for speech emotion recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.641425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:53.110719Z digest=sha256:86c4d52b243c02f7ed159b95a025c26f1ff743c91c0c184d1eb0a5c010e67af4

Observation ca4873d6-0c0d-4e6b-a43a-e65f13184559 · outbound

This paper cites Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Heterogeneity over homogeneity: Investigating multilingual speech pre-trained models for detecting audio deepfake,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.588877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:53.269332Z digest=sha256:537db77f905562b46a8d6f990a2487e4260591dd0abdd40db9c80e5820528c1b

Observation acfaa700-c80b-48a7-8cf7-14352455a022 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:53.455978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:53.455978Z digest=sha256:b7c45890dd1a13e836fa06b742c745cf41306c9b77e618b77a53933b27da4bf0

Observation cb1e393d-8591-4161-be9a-ea9244a82828 · outbound

This paper cites Unispeech-sat: Universal speech representation learning with speaker aware pre-training,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Unispeech-sat: Universal speech representation learning with speaker aware pre-training,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.455841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:53.619048Z digest=sha256:4f505160e833f831f9bbbc4385e2693af116be3977ae15d7769b3458801cd9d6

Observation f5cf2178-ceb3-4c9e-9b77-b447c2c7174a · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:53.770157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:53.770157Z digest=sha256:b295508b2ceb85ad1b67954cf2a74ca5f375ecd9e6dfa7bf076315b8f9aeef26

Observation a3d5bfc7-3246-4ca6-9c49-d8b336f159fc · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:53.913426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:53.913426Z digest=sha256:a7bad4bfef910a9e872abc4945ea139494a572beb87315cbb8c0ab81573d6512

Observation 36458d9a-1854-4b5c-8ab7-553d023f5c4e · outbound

This paper cites R ´enyi divergence and kullback-leibler divergence,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? R ´enyi divergence and kullback-leibler divergence,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:54.057247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:54.057247Z digest=sha256:542f67e0d5ae405efb2f993a36211b1b38486fc05c07e6f454dbb6983228636d

Observation 5ff0ecd1-5f3b-40b7-9a2d-3b9b7d77c9a7 · outbound

This paper cites Asvp-esd: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Asvp-esd: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:54.077393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:54.077393Z digest=sha256:b7102ef330fb54a4d0b78ce0929c9d72f1d672efc4bf75818d2232739446a0cb

Observation 82c93a67-9c89-4e6a-995f-fb5e3a48f7f6 · outbound

This paper cites Jnv corpus: A corpus of japanese nonverbal vocalizations with diverse phrases and emotions,.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? Jnv corpus: A corpus of japanese nonverbal vocalizations with diverse phrases and emotions,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.319362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:54.080965Z digest=sha256:8460bed3a4e34a9ba31339ef011cb4c7804d0811c5e7cc2147d2b64feb9ccb46

Observation 1b27c255-fff4-428e-a2a8-86c7ea9f4a45 · outbound

This paper cites The variably intense vocalizations of affect and emotion (vivae) corpus prompts new per- spective on nonspeech perception.

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition? The variably intense vocalizations of affect and emotion (vivae) corpus prompts new per- spective on nonspeech perception

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:54.187393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:54.084398Z digest=sha256:1c8ca39c3344e137a00d3822b05f46cacaf030f8d99ad94714982ddf92b8fdd3

Pith citing papers

No inbound Pith citation observations are available.