Pith. sign in

Paper Citation Record · LEDGER

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

As of 9 August 2026, this Paper Citation Record lists 7 of 7 outbound references and 3 inbound Pith citation observations for arXiv:2602.10230.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.10230 v2

Coverage vector

measured 7 of 7 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:17:47.362069Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:43:32.422846Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

7 of 7 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 739419b8-9968-432a-b95a-06f712fcb33e · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Sequence Transduction with Recurrent Neural Networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:46.970140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:46.970140Z digest=sha256:fed06326c720de6dd517386b8079e9edf7d3f7241a6044c292525d4f249a6ea6

Observation 756e342d-7bf4-4be8-b1ed-6d74b418cf49 · outbound

This paper cites Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.237653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.237653Z digest=sha256:606ffc6aa1593f4f51e747f88a5773abde6d13349a9c6b82e5a1777f05fcad8e

Observation 6f92203b-6391-4f61-aa9f-9cc912f31ddc · outbound

This paper cites word": "much.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization word": "much

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.362069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.362069Z digest=sha256:727c08a6eb60f068cbf43a1b17c561f1855e08690e2f34de3664335c90ef6aa0

Observation 2b8c1b47-2992-4ede-99ee-0329898dec86 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.065638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.065638Z digest=sha256:084c0928c883814e30ca4d24b8270acf57b6b90d28afa0a6b3986e4d192a6fdc

Observation bc4856d6-0694-4807-94b7-1f02e9a70a05 · outbound

This paper cites cc/paper_files/paper/2023/file/ d842425e4bf79ba039352da0f658a906-Paper-Conference.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization cc/paper_files/paper/2023/file/ d842425e4bf79ba039352da0f658a906-Paper-Conference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:47.268297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.268297Z digest=sha256:01bc5f57982c39717f4bd1f51634b1b2a8863312789686cafaf3df692532ffd8

Observation 6f18f1dd-31ad-4a52-824c-bc3794f08c4b · outbound

This paper cites Voxtral.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Voxtral

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-03T01:17:47.177241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:47.177241Z digest=sha256:5cf77968bcc3652967eb5950769f4fbf4b41b1dda1037b5a0f24ec82fddf6abc

Observation 7df5fd3d-4755-4212-87af-cdc4cbb74137 · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T01:17:46.910143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:17:46.910143Z digest=sha256:d0f8b82f1aeab09adc0972e734b069d5d28c38e6c596b26cf416485e0ea6efa7

Pith citing papers

Observation 4ad3990f-7093-488d-86af-0f21858899a4 · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:32.422846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:43:32.422846Z digest=sha256:dbc95473126e83b0d0e3c099d5fba0cb85a2f80b57c440b3a8030591edf5765b

Observation 5109a179-b385-46fc-b755-48b774a3b227 · inbound

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision cites this paper.

SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:28:47.970632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:28:47.970632Z digest=sha256:63a677159f72bdc6ff3040af32732386411a59fad6fa28a1198d513a62e17f73

Observation a227c080-866e-436d-b888-9ae5bb68591b · inbound

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding cites this paper.

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:46:57.963783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:46:57.963783Z digest=sha256:98cdd1742dd002cc5d3aa01b407321d746cb7690e5d0ad2b76950e10019c2f88