Pith. sign in

Paper Citation Record · LEDGER

Efficient Training of Audio Transformers with Patchout

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2110.05069.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.05069 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:05:47.857768Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:02.620646Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 607c1373-cf4d-4ae3-9802-db6a3ac4bef6 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Efficient Training of Audio Transformers with Patchout

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.857768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.857768Z digest=sha256:ca6591d2371701ce081b51d327bad0bf2f67b266685634c4e7c6c2c6c938ba93

Observation 547da7d8-38ba-4613-8b33-23dff43c8410 · inbound

Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification cites this paper.

Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification Efficient Training of Audio Transformers with Patchout

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:42.788432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:42.788432Z digest=sha256:927b620d32a7c923460bda900b773c35d5b56696ff32093ca812adcf8aa14a54

Observation d4f56042-7847-425a-9a6b-5a373a4f283a · inbound

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation cites this paper.

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation Efficient Training of Audio Transformers with Patchout

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:06.847274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T08:20:02.986562Z digest=sha256:f6db4ad7a35d7852cfebfa656017cb80f077b21ae302ece45ad5c8ca2ef466be

Observation d88f42e7-2290-47d5-b122-5dc5d3a2c19f · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Efficient Training of Audio Transformers with Patchout

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:56.830276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:56.830276Z digest=sha256:6a240904ee2bdaad26fbe186003e0625716e59fe2a473367c6b29767a6af1b1d

Observation 43296ace-ee09-4fb4-b75e-52caf66cff98 · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation Efficient Training of Audio Transformers with Patchout

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.082250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:1e4405ad3a0d8dac0cc2f7ee91e6ecf0cb68096f5cc311c6f3b8381c6a052011

Observation 4bb54ab7-7168-4529-8112-4002e8e33a81 · inbound

Conditional Flow Matching for Visually-Guided Acoustic Highlighting cites this paper.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Efficient Training of Audio Transformers with Patchout

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.142515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.142515Z digest=sha256:2df8f9ec33e24b4288bbaf77450a288a8d7f63b96f09890213807dd43fcd0406

Observation decb39f0-df24-489b-8a5c-d4fa99bd6956 · inbound

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models cites this paper.

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models Efficient Training of Audio Transformers with Patchout

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.493830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T19:53:18.200223Z digest=sha256:5451d86435343a8cbda118ff13838f69a08312635d6afad147b234ab994b4cde

Observation c1577142-cf88-47c2-b439-badde0b4144e · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling Efficient Training of Audio Transformers with Patchout

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.047377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:e72df41474cb4ad0c672d8549fac25638cb3ca4eb308ee374bb13d5aa2989881

Observation d52a8f66-5a48-4a3e-b5c7-9c16034491a3 · inbound

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing cites this paper.

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing Efficient Training of Audio Transformers with Patchout

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:55.879342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T03:27:33.149625Z digest=sha256:e8fa915ee8b89c1a13d1b19ea205352233908c1382e2c98a77e8f94f897da9ba

Observation 8d311385-866b-4523-b396-c7e46adc7ea9 · inbound

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification cites this paper.

SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification Efficient Training of Audio Transformers with Patchout

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:57.043063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:31:49.239866Z digest=sha256:eb5c460e4f18222a02c2b49fa3c4d300e344d533c67f208a9f916abf8728cfba

Observation bd50c096-80ff-447d-bcf8-12c7bf936eb8 · inbound

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators cites this paper.

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators Efficient Training of Audio Transformers with Patchout

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:25:59.043307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T03:24:50.604019Z digest=sha256:886ddfaa264d92195c5f166d322d1f4ac37be90abbdacc6d944ae0dfb1beaade

Observation bcdb4ba8-a655-45f7-9fb2-02b094bf5848 · inbound

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation cites this paper.

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation Efficient Training of Audio Transformers with Patchout

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.459909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T07:26:28.527338Z digest=sha256:6d4b5fb5a4ebfa5776302d7274be1b73bee9ae31f426ed37eb622f6d7ebd1d76

Observation 842dbdcd-230f-4d5f-adec-9ccdaf3e1684 · inbound

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation cites this paper.

Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation Efficient Training of Audio Transformers with Patchout

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:02.622264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T22:42:57.448084Z digest=sha256:7110949e03a6f7929a0aa1793da9567bd9da7abb1baa005f2d69783f2948e6d8

Observation 3f25018d-04b1-4c8d-957b-3f56b043c333 · inbound

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation cites this paper.

FdAudio: MeanFlow-Anchored Fr\'echet-Distance Post-Training for One-Step Text-to-Audio Generation Efficient Training of Audio Transformers with Patchout

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T11:52:50.598080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:52:50.598080Z digest=sha256:6e6b52571041998bab66df3be88fd5351d2d7a6c97b024b977976acb350ec065