Pith. sign in

Paper Citation Record · LEDGER

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2507.11096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11096 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.637571Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.422666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.492700Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy36
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1836cb7-56b4-40b5-8a95-1dd7398a3814 · outbound

This paper cites EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.371790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.410643Z digest=sha256:6db4c66b307881de183e135abf797d74b4178f22a64825ddc4ecaf090ae94ee3

Observation 39fb741f-b162-410a-847e-7ee8294c203f · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.362693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.418748Z digest=sha256:793168d7b5137c969bab8165ba9c29b15a8ba84bc29fe8cb20ae4d0519e3e8c9

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · outbound

This paper cites EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:2667cc41be133a822258569ba3062d1ca92c1fa27f12cef2feca2b98b0cb0105

Observation 62c952b9-1fb4-4883-81ca-5f8d0c33861e · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.352291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.427018Z digest=sha256:174a4c89187e0dfed74c73898ece26503dcc265ed6b7bd25fbf742900d5d1fdf

Observation cd8834d9-09b7-4c01-aebb-3e43448982ef · outbound

This paper cites Yang et al.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Yang et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.342500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.438096Z digest=sha256:0f15592df8468a1356e52df18f09d320b6e3ea932e1853b3bc17e18ba565f2da

Observation 14ee1a2e-467a-4ef6-b21b-050c29796121 · outbound

This paper cites We began with Auffusion, leveraging its existing capabilities for prompt- based editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing We began with Auffusion, leveraging its existing capabilities for prompt- based editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.269463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.468785Z digest=sha256:6c3b7821ae0e1ed718409576264db3642549735d6e67cf6192d1e47e01ebd94b

Observation 74e85710-6cef-4bc2-8fa4-66c6473c4f52 · outbound

This paper cites Masked autoencoders that listen,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Masked autoencoders that listen,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.197286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.499750Z digest=sha256:b0fed38e0cbeae67ab761f8896309fdf2580ba68475ba8119694b463785a05f4

Observation 8f6c65de-b84c-44fb-be5b-9f5ef05bc44e · outbound

This paper cites The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs).

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.309637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.455289Z digest=sha256:3eb4a41eea7ccfdf05aa0dbb4ec7f45fac7041388265397c94c4116e30ba3f3b

Observation d2e3bb99-8b24-4daa-9251-504a96eea1a2 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.289769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.459793Z digest=sha256:698aff68f6d1c63162905af79c5659f8bfbc551609fee93f4a6dc91dd9bfd404

Observation 7ffbaf05-ae58-446b-8e7f-cfb9a91a1885 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.279470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.463949Z digest=sha256:79b311a2a94ff7be60ea7991660c8541f06e0d770825d5ac5acd61cea2ef697a

Observation 5ea41673-d7c3-4a12-95d4-bf4abbf531ea · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.175498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.520183Z digest=sha256:70c2e1412cc9b1d225c2456691414ab2e2bc3e47cb0e3e4797c2cee0f98a079a

Observation dfbc75aa-9c63-4334-925d-661392fa34f9 · outbound

This paper cites Prompt-to- prompt image editing with cross-attention con- trol,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Prompt-to- prompt image editing with cross-attention con- trol,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.259659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.473034Z digest=sha256:6168c7a6373d0b9a66d2b6c014174f218d137d9cb9e7b41314df1ff196ba3a9e

Observation 51391606-a24e-44bd-a28a-3088e7a2d9f7 · outbound

This paper cites Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.333086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.444796Z digest=sha256:6520eabf87db4b3295467011100fd3158354de143e5ab4e82b1fca307dfaaf44

Observation 13c1e6ef-09d2-4070-974d-b5c4e9ff79cb · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.249345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.477277Z digest=sha256:24b67874025a96bde2df32ed336d5f3f488b92d638925ea2740eb427c84f1790

Observation d163a200-2080-4189-81c4-ec27564fde37 · outbound

This paper cites Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.239226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.481096Z digest=sha256:5ac7c3e445a9d85934a9d12e6b3a8c6b730f26f41bb8edc3951c4bdd014cc58f

Observation 0c747d61-fac2-41a2-9f55-67d3490b89d9 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.228419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.484451Z digest=sha256:93fe0dd41826054737f449df16a7907ca7f8127ed38580eb653b5255fbabfa1c

Observation 05a24ad4-193e-4e78-a85c-39471fa3a39e · outbound

This paper cites Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.323341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.450753Z digest=sha256:a395fc98fae84d6f6bec300fada59dc1f8a1f7220bdf509c173c4c568eba02f5

Observation 094ae7f7-c06b-4836-970a-c16f3a7d3432 · outbound

This paper cites AudioLDM: Text- to-audio generation with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM: Text- to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.213796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.491844Z digest=sha256:1abc2aeb12a7441f541004f5ae19c65db9588b5669894dedfaffe6fc61df0b06

Observation 641e40c4-6c03-4697-812a-0c7803250b6b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.495779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.495779Z digest=sha256:6041a5cfb4e9f09729f5406e1821ee9cf95d568b1458331bfa107a27954ce332

Observation 6f653922-2083-43b9-a3de-f852195c57c6 · outbound

This paper cites AUDIT: Audio editing by following in- structions with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AUDIT: Audio editing by following in- structions with latent diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.186979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.506419Z digest=sha256:860ba7cdb295104f57292e4259ee46f2622d50aa6c5df1f8ddb000b706c5893e

Observation b58c729e-353c-4979-b850-2ae52784f25a · outbound

This paper cites Text-to-audio generation using instruction guided latent diffusion model,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Text-to-audio generation using instruction guided latent diffusion model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.510770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.510770Z digest=sha256:f114bc86a1c3cca9fa953cf20326df5623b6004c0e1e0542d6162e6b6890a909

Observation cb1355dd-1f1a-4735-9bf3-4289fa82414b · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.516295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.516295Z digest=sha256:460de44db2a1ec9e9795e86f83af474beda78785440ded5164f9a71f3939b5d1

Observation 41414649-1fb4-454d-9aa2-7b702a3ff99c · outbound

This paper cites MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.165442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.524168Z digest=sha256:4037704ffc3852e2736bfb3ae34ffd9128734601c6ca6270926b0e4e9d632890

Observation 96f0f70d-1ed0-4b70-9f00-2f55e8cd2515 · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.780086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.527735Z digest=sha256:3e7d175c5d534916a53d8cfeb3445817ca233179f1136307f10eb7d9fad685ab

Observation 27980d42-36ee-4c10-96a4-2befd82c6f23 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing WaveNet: A Generative Model for Raw Audio

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.531204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.531204Z digest=sha256:21ef18b64102a0da208e2de6fc04526c3be2bd8301bf92b0c2193b9c26d3693e

Observation 1b32fa68-7f95-4f25-8b3f-9db7c204b828 · outbound

This paper cites Audiogen: Textually guided audio genera- tion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audiogen: Textually guided audio genera- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.154915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.535045Z digest=sha256:4e92f8848743a8c2a2385695fc8eb1749398ead19579b45ccc3767bd6fa77d67

Observation e8084bfe-9e05-4bb0-86ea-71ba8fef237c · outbound

This paper cites Jukebox: A Generative Model for Music.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Jukebox: A Generative Model for Music

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.540076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.540076Z digest=sha256:458494d11c3cca381d9fd8faf0e3e34ea912f14af92ec12bb16fbe27118c533c

Observation 05075a94-0224-4dae-8a53-2acfa6dd4680 · outbound

This paper cites Neural discrete representation learning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Neural discrete representation learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.140555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.543739Z digest=sha256:371755883c8d3693d5c9bd197962ea2b0b8113ba65f96919a0813ed23f7dacc7

Observation 0f0f49c6-7193-4135-8835-6f6a1965dd3e · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLM: a Language Modeling Approach to Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.547479Z digest=sha256:4e7e3475336fad8658f465a4fb719858eed0beb557a7b1d65ea913cbeaf0d9b8

Observation 757c56d3-8f1f-4fe2-a449-14bc158deee4 · outbound

This paper cites Soundstream: An end-to-end neu- ral audio codec,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Soundstream: An end-to-end neu- ral audio codec,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.123953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.551187Z digest=sha256:170f228100e2a87d74637aa7e89cd5cd8ada14cfbd08ad32d18be1fdb8bd94eb

Observation 430323d8-6cab-48ee-aeb1-1e38f82c512c · outbound

This paper cites End-to-end optimized speech cod- ing with deep neural networks,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing End-to-end optimized speech cod- ing with deep neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.107018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.554641Z digest=sha256:28abf2849195b640a6d4bf4ffeda7928ad23392de1bf04c2088ac815312a52a0

Observation df02e93d-76c6-4a72-a3b7-d9874ff8e559 · outbound

This paper cites Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.096933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.558163Z digest=sha256:ddf4defdc4eea703ead76647c5286201c02a849dd1f48128b3054823ec6e7124

Observation 7dcc803c-55e9-4d6b-9ea3-99ce39c413ab · outbound

This paper cites MusicLM: Generating Music From Text.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLM: Generating Music From Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.562352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.562352Z digest=sha256:1a69e3de54a5a3a37951023742ecf64f6b0fc544cd8be4dbc2b3b910e1d36ee6

Observation 8e362c39-afec-48c0-98cf-0b86c0f3a3d8 · outbound

This paper cites Simple and control- lable music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Simple and control- lable music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.085641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.566201Z digest=sha256:ffe284bc0b1c9ed6282ce77bc864629cd35c13c7af3166732a1060ade0e1575b

Observation 45e2670d-ef47-403c-b9a5-c0e5e8c579ad · outbound

This paper cites High fidelity neural audio compression,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing High fidelity neural audio compression,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.569796Z digest=sha256:bc9fefb54d14bd46f3f8373de1f69c71177f7786316b1800b5c8a29abdc3d689

Observation efb8e301-1609-4cef-8345-274ad718c106 · outbound

This paper cites Null-text inversion for editing real im- ages using guided diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Null-text inversion for editing real im- ages using guided diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.054968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.573681Z digest=sha256:992165d501fae2ade96ee065565a9ba98c9c11d92d17a0909795f46f8c92793f

Observation baaa9fc4-4c3f-49cc-ba3c-acb8cd9b5641 · outbound

This paper cites Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.041200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.578547Z digest=sha256:5fd400606c17c14b737a4815194c3da49f9936a951077894f06a5dac45346b01

Observation 769f635d-1cac-4d9e-8b60-21c208ca9497 · outbound

This paper cites Multi-concept customization of text-to-image diffusion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Multi-concept customization of text-to-image diffusion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.010492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.581943Z digest=sha256:af1d2ff1807467e5570a66b06f1af7c945ee0834216235ff826550c5d761e0c6

Observation 3235656a-5e43-4ab0-b6fa-d678989ffe1c · outbound

This paper cites SVDiff: Compact parameter space for dif- fusion fine-tuning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing SVDiff: Compact parameter space for dif- fusion fine-tuning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.992132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.585641Z digest=sha256:e3cb59eb3028232fce6726bd8397c2d13d9e40fd7ec6596a7454c708c9e1b88c

Observation 7f52f4b6-2a27-4684-9c18-f4a4979c02b9 · outbound

This paper cites Countering language drift via visual grounding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Countering language drift via visual grounding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.979777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.591831Z digest=sha256:ef74926b3be534b95321e84160e1b9b2bd0ee947f1cdc4ed93b1d5e7b635d0be

Observation e7739275-b8fb-4c94-bd53-a58cdba6eccd · outbound

This paper cites Likert scale: Explored and explained,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Likert scale: Explored and explained,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.899844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.637571Z digest=sha256:4f4ca8f98167ca665bf2329ffb3554fb4496d5718d6359a0e8a4a4d9e19ea7b5

Observation 6f6943c1-c394-47c9-a915-fc26e93bb154 · outbound

This paper cites Investigating personaliza- tion methods in text to music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Investigating personaliza- tion methods in text to music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.958897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.599929Z digest=sha256:444954a181b6ccf0ac1641077c4b69dbe289f1bf9e4ceae95566d511bb2209d7

Observation e13327b7-eb7e-4d5a-8ffa-013c22004306 · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.603502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.603502Z digest=sha256:404c6833510c7429f4d381975af49d48a640654a286e30e6b2655dc2a31c76a5

Observation 79b0cdd2-eeb8-4192-88d7-48621f19d2f2 · outbound

This paper cites An Edit Friendly DDPM Noise Space: Inversion and Manipulations.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An Edit Friendly DDPM Noise Space: Inversion and Manipulations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.607581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.607581Z digest=sha256:69d852d0c26957aa68d2a14b247fc0a4ae13f03feefd1f608202b94cf4eb8bdc

Observation e87ace66-a28a-4ddf-a7f4-86419d50b3ce · outbound

This paper cites Photorealistic text-to-image diffu- sion models with deep language understanding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Photorealistic text-to-image diffu- sion models with deep language understanding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.948459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.611446Z digest=sha256:dedbd3244f3b9191f6713f802fa0981dc246a87dbb5a35b3ce3843e456787081

Observation 41af34df-31d2-47da-984d-967b6956125f · outbound

This paper cites Music ControlNet: Multiple Time-varying Controls for Music Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Music ControlNet: Multiple Time-varying Controls for Music Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.697050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.615013Z digest=sha256:fa44c019d71ef7570e69e9c2f241b9a2fa6ab3b4207ff9780e8f2f14e68f5e6c

Observation d1855218-0518-461c-ba1a-da8cc5f7a7b3 · outbound

This paper cites Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T17:21:31.671027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.618387Z digest=sha256:2edf16432f7b798b28695eceaa5d1b6af02e90b616004188f29489eaa9902eeb

Observation 1e2e7b21-c867-484f-a509-356a3ea75b37 · outbound

This paper cites mir_eval: A transparent implementation of common mir metrics,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing mir_eval: A transparent implementation of common mir metrics,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.938083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.621895Z digest=sha256:dbc4a864212a2688eeb13c6ba7466cb6943f31948af5fb3728805fe8393e7967

Observation 149c2df2-8efe-4795-8e0c-12f93b27d828 · outbound

This paper cites An efficient state- space model for joint tempo and meter tracking.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An efficient state- space model for joint tempo and meter tracking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.625142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.625142Z digest=sha256:30f6d6882bcdbede888ee4672b85b55d0cd3ae38f7ea66d028406adca4a207b9

Observation 7928f099-baf0-4901-8ba7-de10090931ff · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.921391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.628136Z digest=sha256:1d01e0a161dca3be785aef73374e48654d7a4a7a496a2c4c5e9a78db39209949

Observation 983330f6-1d59-42aa-94ef-5bd46daeb1dc · outbound

This paper cites HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.911114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.631335Z digest=sha256:b8654d4a9377b4a03084adff19c9ef6d500d164a63a9dfa6f2719ed5cab41237

Observation 8c76b7c9-5998-4a40-a9ba-417eed918ec9 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Representation Learning with Contrastive Predictive Coding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.634601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.634601Z digest=sha256:eaaf6e821b9ebf9e6a5744b0fdb52d3bf1b76fc3c4aad22de50c922d727f380c

Observation 9f26091d-4751-422e-9db3-f3d6d8799954 · outbound

This paper cites Available: https://aclanthology.org/ D19-1447.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Available: https://aclanthology.org/ D19-1447

Reference 4395

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.969764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:21:31.596119Z digest=sha256:38a3efd6943af9b5af657b26a161a6086c4371dfc3567d88ec6c6fe6a6ebd7d2

Pith citing papers

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · inbound

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing cites this paper.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:2667cc41be133a822258569ba3062d1ca92c1fa27f12cef2feca2b98b0cb0105

Observation 0615c238-6ed3-4616-b1d8-e7d03f23a4a6 · inbound

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption cites this paper.

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.494215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T13:17:40.718000Z digest=sha256:2df9dd306c27f2ca36cf20659d83009b206bbfa5785602cbf9bb3f9271e3264d