Pith. sign in

Paper Citation Record · LEDGER

EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.07162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07162 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:27:30.784224Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T04:50:56.285704Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7aed92f4-38a5-4bba-bcf9-32263290b115 · inbound

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing cites this paper.

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:27:30.784224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:27:30.784224Z digest=sha256:1917bf510d86a2bd00f835b75fb5c6d6103dca9e6b7a520c705e032c4ef36fa3

Observation f7f15213-8484-48b1-95af-70344a6384e4 · inbound

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch cites this paper.

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:35.619947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:35.619947Z digest=sha256:b0e1f4830de66f7daed691eeec0b0ba27ba1860a5597a2133f2f9099de5aca73

Observation 9b40a672-56f6-44db-bbc9-79e3aa93cda2 · inbound

Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition cites this paper.

Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:46.216697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:03:46.216697Z digest=sha256:e9a01b1ec6255f175b5d1a80068e6f74ef28473300f030246a689187c625fb0d

Observation 111aef6d-dfbe-4c16-ad97-46e77b42d933 · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.685540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.685540Z digest=sha256:c60827296ebc051301a95f20983f099fc2ca5586d468c1feb5bb3ca6f55bb86a

Observation a913a405-6aa8-4709-99b5-ba60796d5be2 · inbound

Differentiable Reward Optimization for LLM based TTS system cites this paper.

Differentiable Reward Optimization for LLM based TTS system EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:03.992535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:03.992535Z digest=sha256:7503832195c7a90f2e8e7b4bef9c0ef430a1cfd3d5c8483cd86011b7901d7f8e

Observation 5128c4bc-6773-4381-bdbf-52782b84470e · inbound

Marco-Voice Technical Report cites this paper.

Marco-Voice Technical Report EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:17:06.215755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:17:06.215755Z digest=sha256:d47ef43ea35e6f5e92efc5cc8e45168fe4c562e50b0db415984e82b51e948b48

Observation cb258ebd-d063-4120-b459-a5dddc0c6ee4 · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:59.835425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:59.835425Z digest=sha256:a62734f0b1dca7c4135f34acc7ab2e47e9d29267a727ea0d91ff97e98f09e04a

Observation 8469d56e-a787-47ba-a552-3ddcd28e3b26 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:56.288746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:c8e5742b6f49822e64a8a3aa2389ac4b7ad01c5a1b0b47f8e531c90ae793a8bc

Observation e0a38af5-a145-4534-8a0e-75a1f567826c · inbound

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models cites this paper.

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:43.989920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:43.989920Z digest=sha256:a70ad794312e0b17b758eb60bb65a56ef491ed2f606f1335453c529b9fccbcc1