Pith. sign in

Paper Citation Record · LEDGER

Exploring State-Space-Model based Language Model in Music Generation

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2507.06674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06674 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:02:24.400382Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:02:24.319848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:02:24.469579Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52ca52a8-ca71-4afe-bf4a-16e4c5b02055 · outbound

This paper cites Exploring State-Space-Model based Language Model in Music Generation.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.654861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.315970Z digest=sha256:73d0f72585e9ba0ea85cd84b5f40c8ce1979642d7539844e07a7e5b49d18d26d

Observation 00fb6b57-fae5-4d8d-8ec0-16a53675b20c · outbound

This paper cites Exploring State-Space-Model based Language Model in Music Generation.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:02:24.474768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.319848Z digest=sha256:0aadbbcb347d1021810bd1151f896d9ddcf09ec36fc7f754cdbe97daa1f5102d

Observation 933e7d64-8e41-4ebb-b62c-3c85af3ace30 · outbound

This paper cites We re-sample all the audio into 44.1kHz and convert them into mono audio, splitting the tracks into non-overlapping 30s clips with vo- cals removed by HTDemucs [3].

Exploring State-Space-Model based Language Model in Music Generation We re-sample all the audio into 44.1kHz and convert them into mono audio, splitting the tracks into non-overlapping 30s clips with vo- cals removed by HTDemucs [3]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.646549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.323872Z digest=sha256:c371938ee7a98441a928825fa69586970d6d07f1cefd9b9e63e041a94470ac4d

Observation 92e4a2a2-8d09-43d2-842a-bfbf43b85354 · outbound

This paper cites Both FAD and KLD are lower the better, while CLAP is higher the better.

Exploring State-Space-Model based Language Model in Music Generation Both FAD and KLD are lower the better, while CLAP is higher the better

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.638280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.327623Z digest=sha256:c373c348beddfaa21ef1a0cd25e0ca3fe516a9877598dcb2b0da7e30c9bfa5af

Observation 67e548ea-689c-482e-8389-041b800fafdb · outbound

This paper cites Decoupled weight decay regularization,.

Exploring State-Space-Model based Language Model in Music Generation Decoupled weight decay regularization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.629458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.330769Z digest=sha256:cc4a5244b7e2263813870c6915d7161d7a532518f595eac8d8725e88bbd418b8

Observation c631407a-b575-4016-9edd-38cd09ccdec3 · outbound

This paper cites The MTG-Jamendo Dataset for Automatic Music Tagging,.

Exploring State-Space-Model based Language Model in Music Generation The MTG-Jamendo Dataset for Automatic Music Tagging,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.620895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.334460Z digest=sha256:75c06b4c8128493ce1669f4bdf2cf94c58bafc0054719d09211568fa111b2799

Observation 00e27017-3956-4034-9a79-9e6083dc83e5 · outbound

This paper cites Hybrid Trans- formers for music source separation,.

Exploring State-Space-Model based Language Model in Music Generation Hybrid Trans- formers for music source separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.612619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.337857Z digest=sha256:3ce842825f93553ed9a0ee025dc1651e35a2f65ad280c84e4ec493171161cc38

Observation 11f9d51f-5607-4a3d-842a-24024fa52584 · outbound

This paper cites LP-MusicCaps: LLM-based pseudo music captioning,.

Exploring State-Space-Model based Language Model in Music Generation LP-MusicCaps: LLM-based pseudo music captioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.603686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.340676Z digest=sha256:3f5a7815a02c674cecee1cd55d66a118b77c628ed264ac1b95b0f8830b436dd9

Observation 598b522d-7970-4ceb-9cde-8c1f1bf82d13 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring State-Space-Model based Language Model in Music Generation The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.343931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.343931Z digest=sha256:c9183d6ecdfb604c3bb4f82b75d38a73c03b79fbaa869939eb6bca7590228f07

Observation b85630d5-8619-4ffd-9772-07fe6ea42ae6 · outbound

This paper cites The Song De- scriber Dataset: a corpus of audio captions for music- and-language evaluation,.

Exploring State-Space-Model based Language Model in Music Generation The Song De- scriber Dataset: a corpus of audio captions for music- and-language evaluation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.594043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.347731Z digest=sha256:2f36d518b4c677d63ada2fc2fb3b2b3dca52de756a0fd46024938d9e0aaa2144

Observation 000d75ce-3637-4739-b292-b75a829d140c · outbound

This paper cites Scaling instruction-finetuned language mod- els,.

Exploring State-Space-Model based Language Model in Music Generation Scaling instruction-finetuned language mod- els,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.585432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.351332Z digest=sha256:cf83eff66aa326fee067c22b3f9dfb6bcba824db95e9989cd8ae86ed9ef145ba

Observation 7277841f-9344-40b2-a704-6fb3e73b614d · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Exploring State-Space-Model based Language Model in Music Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.354155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.354155Z digest=sha256:6279f81bdabfeb8f6c822a0c13cb3a7c4da69cb76936237ed54f6bc5a269c90d

Observation 30aec164-c4aa-49f2-b7c3-cb9b7ef8d4fc · outbound

This paper cites Transformers are SSMs: General- ized models and efficient algorithms through structured state space duality,.

Exploring State-Space-Model based Language Model in Music Generation Transformers are SSMs: General- ized models and efficient algorithms through structured state space duality,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.576707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.357348Z digest=sha256:b5323bfd6ec866adc93d6c2a99c095511c30194d9f9f8c386a4357fe41a4a5d4

Observation 2e692a93-aa63-4c43-8af3-e2c6e037539c · outbound

This paper cites SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series.

Exploring State-Space-Model based Language Model in Music Generation SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.360110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.360110Z digest=sha256:0c741b51c94e84d4168d3a6a43acd65e574b2e95bde340904122d6cd5a132d70

Observation 8fb1814f-e275-41fa-a96a-c12fce8d4b72 · outbound

This paper cites High-fidelity audio compression with im- proved RVQGAN,.

Exploring State-Space-Model based Language Model in Music Generation High-fidelity audio compression with im- proved RVQGAN,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.568184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.363148Z digest=sha256:77fd62f99ab9f8b49a7a21905114e2c24c5ed18fa006acf33d8a3558a8c16e98

Observation 756330a9-bde3-40ae-bc9f-83851b512d1a · outbound

This paper cites Coarse-to-fine text-to- music latent diffusion,.

Exploring State-Space-Model based Language Model in Music Generation Coarse-to-fine text-to- music latent diffusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.559894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.365909Z digest=sha256:a10f7251ded79068dd8e8aa29b8f44db710296cec0e4049887218b90372fdc33

Observation f1861585-e6c7-4b67-8a64-3bd05e108397 · outbound

This paper cites MusicLM: Generating Music From Text.

Exploring State-Space-Model based Language Model in Music Generation MusicLM: Generating Music From Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.368575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.368575Z digest=sha256:a9f104637f7537fc47b03f9924e809bd3dc5a402f2a93be697e6d8b653671365

Observation 712c6245-8d1c-458f-a56a-62f7ad753d21 · outbound

This paper cites Simple and control- lable music generation,.

Exploring State-Space-Model based Language Model in Music Generation Simple and control- lable music generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.551716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.371630Z digest=sha256:8f5609fa552093149a5668c07d26b48d87aecdbb749e1bfeff8c40ff90b8da96

Observation 4c80ae5d-f2e2-4948-9501-035de94da30d · outbound

This paper cites Stable audio open,.

Exploring State-Space-Model based Language Model in Music Generation Stable audio open,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.543061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.374380Z digest=sha256:53a24188788f181190586a36a26b260af44d697874d1615f5b85c399f21d1911

Observation eb504c78-366e-48ec-9b61-4c94020b752b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining,.

Exploring State-Space-Model based Language Model in Music Generation AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.533682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.377638Z digest=sha256:ff52cb438ac4e9f8158aca00b1cd795b68067909d950cadf4fd4e0f05e2835b7

Observation 6052973c-9722-4fbe-9ef5-59f9fd2b497e · outbound

This paper cites Mustango: Toward controllable text-to-music generation,.

Exploring State-Space-Model based Language Model in Music Generation Mustango: Toward controllable text-to-music generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.523416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.380443Z digest=sha256:4b7e0dabc62602b75d1570c2f3c716b3e9232b6549b69b6df749bfc2ec89e7d0

Observation 18188d7c-b941-47ea-abf2-a948124e77d0 · outbound

This paper cites JEN-1: Text-guided universal music gener- ation with omnidirectional diffusion models,.

Exploring State-Space-Model based Language Model in Music Generation JEN-1: Text-guided universal music gener- ation with omnidirectional diffusion models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.514119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.384619Z digest=sha256:88879d021d4d45b8e6359f74f17fe233fef95adb613160ca46e0bd0791990fcf

Observation 981006d2-d1a6-4647-8332-b80cd2f36eda · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Exploring State-Space-Model based Language Model in Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.387244Z digest=sha256:8d5180d81d8ee318c2af04371db93ee528297a454fc27067229717915812a4fc

Observation 826ee30f-3694-43f6-8d8d-1672cfda6509 · outbound

This paper cites Music ControlNet: Multiple time-varying controls for music generation,.

Exploring State-Space-Model based Language Model in Music Generation Music ControlNet: Multiple time-varying controls for music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.503756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.390173Z digest=sha256:bab319efafb7ffbbf31cf5a9e46d5f9d7e7ecca8f91346336ec55fcb743c318f

Observation 24806345-b78f-48b7-bf08-ab5630f53a63 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Exploring State-Space-Model based Language Model in Music Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.393797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.393797Z digest=sha256:c75582ad3505cffc51d151d1cd8ad5abb8f1fac8b771ccc2500b239dd6782b8b

Observation 3892877c-48bd-44d7-9e39-bd73ceda8384 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

Exploring State-Space-Model based Language Model in Music Generation CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.493484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.396775Z digest=sha256:eb80fd0d1069d2481250ed4ab6cdcb29cd26f4e910e31edb3b93c4b6212d93e4

Observation 90bba355-2469-4f7f-b1ed-0d607a9d9f5e · outbound

This paper cites MuseControlLite: Multifunctional music generation with lightweight conditioners,.

Exploring State-Space-Model based Language Model in Music Generation MuseControlLite: Multifunctional music generation with lightweight conditioners,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.484848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.400382Z digest=sha256:65bf3e2d4d440fc1b4dbab058e3b219a73468ecdc185d642532fe2789f1efa94

Pith citing papers

Observation 00fb6b57-fae5-4d8d-8ec0-16a53675b20c · inbound

Exploring State-Space-Model based Language Model in Music Generation cites this paper.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:02:24.474768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:02:24.319848Z digest=sha256:0aadbbcb347d1021810bd1151f896d9ddcf09ec36fc7f754cdbe97daa1f5102d