Pith. sign in

Paper Citation Record · LEDGER

Multi-modal Attention for Speech Emotion Recognition

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2009.04107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2009.04107 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:33:41.059452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T16:07:53.138866Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c4c9938-94d3-491a-9d76-9c366aceef82 · inbound

OminiControl: Minimal and Universal Control for Diffusion Transformer cites this paper.

OminiControl: Minimal and Universal Control for Diffusion Transformer Multi-modal Attention for Speech Emotion Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:41.059452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:41.059452Z digest=sha256:4507cccb0950ddabadae7bcb3591ef08b82c2bfcbe948a2959f6a17668d4d55e

Observation 54dca83d-8cfc-4143-a6b7-3633de53c5ce · inbound

LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer cites this paper.

LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer Multi-modal Attention for Speech Emotion Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T16:37:25.837141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:37:25.837141Z digest=sha256:6386a36329063e35a00954deb9c05334c56864e07714be87185767c666d3dd3c

Observation 7c6409e1-6aa2-4a5b-be23-50389245a262 · inbound

MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation cites this paper.

MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation Multi-modal Attention for Speech Emotion Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:57:58.069148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:57:58.069148Z digest=sha256:421932dc713b0d1bb303c4c3f2cc4f44d331dcf999ffb4e7e232b5baad3c99d6

Observation ddc34ba7-9a96-4450-a29a-63af3615c266 · inbound

In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer cites this paper.

In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer Multi-modal Attention for Speech Emotion Recognition

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:07:53.140578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T16:07:53.054355Z digest=sha256:67c4ad01c9e92ba76a1cc4f2e780cdf88a760f6f7962e40ded734f1999728fe2

Observation de3e0fe3-0290-4195-b903-e405c77dad86 · inbound

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data cites this paper.

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data Multi-modal Attention for Speech Emotion Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:08.695045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:08.695045Z digest=sha256:62ec218f4724519548a6940dfdcee6e8a6ffb1d9d9711b5191254b00a57dc35d

Observation d99b507b-246d-4ca3-ba17-1fc6b982fe04 · inbound

Learning Annotation Consensus for Continuous Emotion Recognition cites this paper.

Learning Annotation Consensus for Continuous Emotion Recognition Multi-modal Attention for Speech Emotion Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.880411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.880411Z digest=sha256:941d35d55757421e7a1189c749dfe1b7d2ac7ad1c715d66b28c18d8784251905

Observation 70218a74-fba3-4563-817a-34557ad3eb31 · inbound

RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers cites this paper.

RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers Multi-modal Attention for Speech Emotion Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:28.052890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:28.052890Z digest=sha256:b0db076d62ce1b7d7dc3acb9cf7efd8e82b508294374242de291f96ba3ba225a

Observation 9ced870f-7f9a-4704-bda6-22122cf3619b · inbound

Image Editing As Programs with Diffusion Models cites this paper.

Image Editing As Programs with Diffusion Models Multi-modal Attention for Speech Emotion Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:38.472327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:38.472327Z digest=sha256:7d8d24430e86d40e1253337c03ae98ac3af69b7fcbca123b5acae3f3bfc45f3f

Observation e91dbf3d-efc3-490d-bece-3e8589f93713 · inbound

WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending cites this paper.

WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending Multi-modal Attention for Speech Emotion Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:57:32.728477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:57:32.728477Z digest=sha256:eee4f3a148cb7dfe24eaee5f9d9929b025e8c17b03e19104740a8eda8c51b703