Pith. sign in

Paper Citation Record · LEDGER

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2507.11096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11096 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.637571Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.422666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.492700Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy36
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1836cb7-56b4-40b5-8a95-1dd7398a3814 · outbound

This paper cites EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.371790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.410643Z digest=sha256:c57fda1d9dbeb255bbb087d2363d4a507a128f9223c0848d750fff7fdf197ee0

Observation 39fb741f-b162-410a-847e-7ee8294c203f · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.362693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.418748Z digest=sha256:bec17a1dba270a0938d46b9269e41baa3d3767f5d9e789835f32ca0a5a9649d2

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · outbound

This paper cites EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:c6f83509c75fe5b107bffcf26543aba32236e67ce47f9603754795386cc08c2a

Observation 62c952b9-1fb4-4883-81ca-5f8d0c33861e · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.352291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.427018Z digest=sha256:c6eb4b34b3b1eb217c70ff335fe4a622810def4cd6f201cd5a313425ec54a883

Observation cd8834d9-09b7-4c01-aebb-3e43448982ef · outbound

This paper cites Yang et al.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Yang et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.342500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.438096Z digest=sha256:1aeb741c08f8972cc9696c9a3cf805abcd4d70f91976059984e0c967e508e98f

Observation 14ee1a2e-467a-4ef6-b21b-050c29796121 · outbound

This paper cites We began with Auffusion, leveraging its existing capabilities for prompt- based editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing We began with Auffusion, leveraging its existing capabilities for prompt- based editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.269463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.468785Z digest=sha256:a85ca3ad4ae81d6e95cbeabbe3cd39a47f86d86f6ebbffb4668dacdc520ab907

Observation 74e85710-6cef-4bc2-8fa4-66c6473c4f52 · outbound

This paper cites Masked autoencoders that listen,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Masked autoencoders that listen,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.197286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.499750Z digest=sha256:1cbc008de134284f5ba4627ce3360c1771ec16d978ce52471057100bbdd8ac05

Observation 8f6c65de-b84c-44fb-be5b-9f5ef05bc44e · outbound

This paper cites The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs).

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.309637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.455289Z digest=sha256:d8d34ca19dcdb972b8a7e0c4d71554c9421b0018f271e711de77fedddf2a4e9d

Observation d2e3bb99-8b24-4daa-9251-504a96eea1a2 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.289769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.459793Z digest=sha256:beef2b9ffafa02f4a6f60bd9c230569933dae0c88bd37c5ee8948bce0ddfda31

Observation 7ffbaf05-ae58-446b-8e7f-cfb9a91a1885 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.279470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.463949Z digest=sha256:dbf9d57353c7073517627749b2065fec274afabbb7370b3ffc8bd4f36faa6b1b

Observation 5ea41673-d7c3-4a12-95d4-bf4abbf531ea · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.175498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.520183Z digest=sha256:df3ec55be9322e2f21e2657f85b093bf105dac05576afd2a48ee59ca96fa88ef

Observation dfbc75aa-9c63-4334-925d-661392fa34f9 · outbound

This paper cites Prompt-to- prompt image editing with cross-attention con- trol,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Prompt-to- prompt image editing with cross-attention con- trol,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.259659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.473034Z digest=sha256:ab98828b01eec2b69b09cf59669a9d036c14aff8b217d02ce7614378fc7f8827

Observation 51391606-a24e-44bd-a28a-3088e7a2d9f7 · outbound

This paper cites Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.333086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.444796Z digest=sha256:897673b7c6860d2a642e06b4225e4d29f3fc717761f4987dc4f2cfa45e0d0d9b

Observation 13c1e6ef-09d2-4070-974d-b5c4e9ff79cb · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.249345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.477277Z digest=sha256:cf667ff8fe5ca41d444590ee64b58a1e707ce32e9642f82816e02510d310f5ea

Observation d163a200-2080-4189-81c4-ec27564fde37 · outbound

This paper cites Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.239226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.481096Z digest=sha256:db6a70141ddfd2568ab23f54ba1449e5923589b485f29796070ef27e459150b7

Observation 0c747d61-fac2-41a2-9f55-67d3490b89d9 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.228419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.484451Z digest=sha256:953a38d7039a60e125dcfea6b975b2ab3b31a87b43593430945bdca6bb5120d9

Observation 05a24ad4-193e-4e78-a85c-39471fa3a39e · outbound

This paper cites Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.323341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.450753Z digest=sha256:2fe0be6ed4709eb2ae9688c799aa55b0bd29fca1b8e1af9d153f969a98ea4c56

Observation 094ae7f7-c06b-4836-970a-c16f3a7d3432 · outbound

This paper cites AudioLDM: Text- to-audio generation with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM: Text- to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.213796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.491844Z digest=sha256:d4adf32c3c1b95cf94c715052fd42b333957d655bd04d1f07fb0351110ae0f0d

Observation 641e40c4-6c03-4697-812a-0c7803250b6b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.495779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.495779Z digest=sha256:1642fb4f2b9699c47440648a17fe01b241ed56508e7e0d56d71b3f308227f8b4

Observation 6f653922-2083-43b9-a3de-f852195c57c6 · outbound

This paper cites AUDIT: Audio editing by following in- structions with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AUDIT: Audio editing by following in- structions with latent diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.186979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.506419Z digest=sha256:f804bdbb7773f3e847c1fb53cc0e0c0ba6bf26002a13c6406c03c30d76049f3a

Observation b58c729e-353c-4979-b850-2ae52784f25a · outbound

This paper cites Text-to-audio generation using instruction guided latent diffusion model,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Text-to-audio generation using instruction guided latent diffusion model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.510770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.510770Z digest=sha256:a1e5e51753f47e2b0ae00da041613ef71f7369259c04ebe4a004613e1a403634

Observation cb1355dd-1f1a-4735-9bf3-4289fa82414b · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.516295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.516295Z digest=sha256:3f92e9b5d81a0b52476570bb0170d3fb75dd4c6a6fe516d2233bc6702f50b63d

Observation 41414649-1fb4-454d-9aa2-7b702a3ff99c · outbound

This paper cites MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.165442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.524168Z digest=sha256:0879f6f36fcd0e8c4be1b4652f2b8c9e90944b9b72c39ca25ece7a5bea9ef189

Observation 96f0f70d-1ed0-4b70-9f00-2f55e8cd2515 · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.780086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.527735Z digest=sha256:606a183ee1d26c24e70b2eb89f4d8e5c1609638f762ed10101e42d76a5caafcb

Observation 27980d42-36ee-4c10-96a4-2befd82c6f23 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing WaveNet: A Generative Model for Raw Audio

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.531204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.531204Z digest=sha256:c2b43174398e142b01e8f6c4dd494011ca4f45e650aec7423286b5c6f709cc71

Observation 1b32fa68-7f95-4f25-8b3f-9db7c204b828 · outbound

This paper cites Audiogen: Textually guided audio genera- tion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audiogen: Textually guided audio genera- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.154915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.535045Z digest=sha256:29a76963ea76a6aa25016b2da1137ea894968c9d120ebb1dca246de4e5d7f73d

Observation e8084bfe-9e05-4bb0-86ea-71ba8fef237c · outbound

This paper cites Jukebox: A Generative Model for Music.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Jukebox: A Generative Model for Music

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.540076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.540076Z digest=sha256:f8f73c073cc56912245293c4eabf47f699d32cee8d3843c8f311b57e6195c8a7

Observation 05075a94-0224-4dae-8a53-2acfa6dd4680 · outbound

This paper cites Neural discrete representation learning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Neural discrete representation learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.140555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.543739Z digest=sha256:78c88401268d3f8a98929a406486342fa2b90f584ce539117b47f4694e69c1c9

Observation 0f0f49c6-7193-4135-8835-6f6a1965dd3e · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLM: a Language Modeling Approach to Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.547479Z digest=sha256:c85db2847b3d7a9fb7f97a6d76431d649e2d97be2e5221c28af120793a428a4b

Observation 757c56d3-8f1f-4fe2-a449-14bc158deee4 · outbound

This paper cites Soundstream: An end-to-end neu- ral audio codec,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Soundstream: An end-to-end neu- ral audio codec,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.123953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.551187Z digest=sha256:2e1c1406bffa14cb541973d5a6f3986b328e1cd1413ed545c44357a7ec78d1c1

Observation 430323d8-6cab-48ee-aeb1-1e38f82c512c · outbound

This paper cites End-to-end optimized speech cod- ing with deep neural networks,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing End-to-end optimized speech cod- ing with deep neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.107018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.554641Z digest=sha256:b82bbc8a5bdaa84fa66e3ab6174c48b60817dd725131d06d0b7e35d15c7f142b

Observation df02e93d-76c6-4a72-a3b7-d9874ff8e559 · outbound

This paper cites Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.096933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.558163Z digest=sha256:715fa7f926a81ad954263e884be085adab60686be34124702e8a1e42c38adc96

Observation 7dcc803c-55e9-4d6b-9ea3-99ce39c413ab · outbound

This paper cites MusicLM: Generating Music From Text.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLM: Generating Music From Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.562352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.562352Z digest=sha256:bdf9a2729d100ab3b5cc1fa58d34b98f6d009ec6eba32e6cf5265b08fe7e8159

Observation 8e362c39-afec-48c0-98cf-0b86c0f3a3d8 · outbound

This paper cites Simple and control- lable music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Simple and control- lable music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.085641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.566201Z digest=sha256:bcf8df45c9e9280fe83ec6be6c9851fe2ea14f8d51a61ae6400ce0f0a99d5e7d

Observation 45e2670d-ef47-403c-b9a5-c0e5e8c579ad · outbound

This paper cites High fidelity neural audio compression,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing High fidelity neural audio compression,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.569796Z digest=sha256:1339510ab57315e111242fa0e4846e108d4db97835e8a8c868018d68701211b6

Observation efb8e301-1609-4cef-8345-274ad718c106 · outbound

This paper cites Null-text inversion for editing real im- ages using guided diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Null-text inversion for editing real im- ages using guided diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.054968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.573681Z digest=sha256:15ca271a524e5f44e5c25b9daefbd3b04b0d1fa287d53014cbd7e424bde54bc8

Observation baaa9fc4-4c3f-49cc-ba3c-acb8cd9b5641 · outbound

This paper cites Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.041200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.578547Z digest=sha256:73d505fe6b2800b22ce3ea07c859b7058e4489000c5c1a28025d4d9c357d0a4b

Observation 769f635d-1cac-4d9e-8b60-21c208ca9497 · outbound

This paper cites Multi-concept customization of text-to-image diffusion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Multi-concept customization of text-to-image diffusion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.010492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.581943Z digest=sha256:bd8068a9d82164a9ee65be3bae537bcf62dd7a40a8a7f9c103b6ffb3a9b5440e

Observation 3235656a-5e43-4ab0-b6fa-d678989ffe1c · outbound

This paper cites SVDiff: Compact parameter space for dif- fusion fine-tuning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing SVDiff: Compact parameter space for dif- fusion fine-tuning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.992132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.585641Z digest=sha256:c6d6e88f16c0088fe1e728f3fc33d0b53e610850654ba41fd91a3bdb97d751d5

Observation 7f52f4b6-2a27-4684-9c18-f4a4979c02b9 · outbound

This paper cites Countering language drift via visual grounding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Countering language drift via visual grounding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.979777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.591831Z digest=sha256:a977ff065999084bbfe135ea354b92cd47e3bd90c1d205ef8f00c008b507225a

Observation e7739275-b8fb-4c94-bd53-a58cdba6eccd · outbound

This paper cites Likert scale: Explored and explained,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Likert scale: Explored and explained,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.899844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.637571Z digest=sha256:0d762df7e80e617137586d76fb17667ed032b59ced267acd25aa9ccf6deea02c

Observation 6f6943c1-c394-47c9-a915-fc26e93bb154 · outbound

This paper cites Investigating personaliza- tion methods in text to music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Investigating personaliza- tion methods in text to music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.958897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.599929Z digest=sha256:184ef038ce7d33440cc4138fa37e006d05fc863d199bb3aa8d61fff5d5f50ebc

Observation e13327b7-eb7e-4d5a-8ffa-013c22004306 · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.603502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.603502Z digest=sha256:4b410d9230281567f616614438151155ea85908328779e7c1b7019cfb1661bf4

Observation 79b0cdd2-eeb8-4192-88d7-48621f19d2f2 · outbound

This paper cites An Edit Friendly DDPM Noise Space: Inversion and Manipulations.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An Edit Friendly DDPM Noise Space: Inversion and Manipulations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.607581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.607581Z digest=sha256:07333c22e0c1ed1090be372c6b31984e89abfe04e94add64f8e856e546befe83

Observation e87ace66-a28a-4ddf-a7f4-86419d50b3ce · outbound

This paper cites Photorealistic text-to-image diffu- sion models with deep language understanding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Photorealistic text-to-image diffu- sion models with deep language understanding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.948459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.611446Z digest=sha256:89a523d0683e64ad78c26ae28c6d75192f33543c4572c0df22f0c498c65ced67

Observation 41af34df-31d2-47da-984d-967b6956125f · outbound

This paper cites Music ControlNet: Multiple Time-varying Controls for Music Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Music ControlNet: Multiple Time-varying Controls for Music Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.697050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.615013Z digest=sha256:49e472a379460a43f96bfc2b9fb25655a98bf31e1fd4b821851d1a9d0a9bc960

Observation d1855218-0518-461c-ba1a-da8cc5f7a7b3 · outbound

This paper cites Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T17:21:31.671027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.618387Z digest=sha256:1d54ab19f8cf471c7df50525d5b4236a4d73c786e239cc08445ceff46b449b04

Observation 1e2e7b21-c867-484f-a509-356a3ea75b37 · outbound

This paper cites mir_eval: A transparent implementation of common mir metrics,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing mir_eval: A transparent implementation of common mir metrics,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.938083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.621895Z digest=sha256:31415623d2e38927fbd327e95f0786f4ded7897aca5c50f505f1dd2571e42312

Observation 149c2df2-8efe-4795-8e0c-12f93b27d828 · outbound

This paper cites An efficient state- space model for joint tempo and meter tracking.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An efficient state- space model for joint tempo and meter tracking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.625142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.625142Z digest=sha256:f9767b42ccf2e9a0e1d17d4afcebc9c7fffd473a2171db52af929b88b0818b90

Observation 7928f099-baf0-4901-8ba7-de10090931ff · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.921391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.628136Z digest=sha256:49196e60ec859b7ed37c7e7b378609ee986735c16c08dc5f6196773c17310608

Observation 983330f6-1d59-42aa-94ef-5bd46daeb1dc · outbound

This paper cites HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.911114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.631335Z digest=sha256:ed5ea3ff028a133d223f67b8d29caf72604f479564770927fe3d0eb456edebcf

Observation 8c76b7c9-5998-4a40-a9ba-417eed918ec9 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Representation Learning with Contrastive Predictive Coding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.634601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.634601Z digest=sha256:e9e27fe85a651e7c3feb2a4ea073ba5e8bc48ef74fdc64af95e171a35782f8c7

Observation 9f26091d-4751-422e-9db3-f3d6d8799954 · outbound

This paper cites Available: https://aclanthology.org/ D19-1447.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Available: https://aclanthology.org/ D19-1447

Reference 4395

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.969764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T17:21:31.596119Z digest=sha256:f6a7aa2558a41523cf26a296d83c7c157d8251d56f5ce45de23bb9f936e5cf18

Pith citing papers

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · inbound

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing cites this paper.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:c6f83509c75fe5b107bffcf26543aba32236e67ce47f9603754795386cc08c2a

Observation 0615c238-6ed3-4616-b1d8-e7d03f23a4a6 · inbound

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption cites this paper.

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.494215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T13:17:40.718000Z digest=sha256:7a427f7552be1ec1b5dadbf64232fb81fcbce4623c4cb1a52a284ee23bdfb247