Pith. sign in

Paper Citation Record · LEDGER

Direct Simultaneous Translation Activation for Large Audio-Language Models

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2509.15692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.15692 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T16:30:39.575148Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact9
  • verified fuzzy14
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c843852a-3bee-44dc-b5b5-17405dfc9ab4 · outbound

This paper cites Can neural machine translation do simultaneous translation?.

Direct Simultaneous Translation Activation for Large Audio-Language Models Can neural machine translation do simultaneous translation?

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.823363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:1fed3b4b363befece56543761fbd7c0a611207f8e3c098a0502e761e895e5be4

Observation 5b36bc1d-782a-4d35-91ee-7fa9a60cb686 · outbound

This paper cites Online and linear-time attention by enforc- ing monotonic alignments.

Direct Simultaneous Translation Activation for Large Audio-Language Models Online and linear-time attention by enforc- ing monotonic alignments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.926646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:26f6c80d827c1979caae9fecd08d4b2b43547eee3d644a3b3593ce4e8344f807

Observation 6ab5d3ee-e123-45b8-ab31-960f6d728406 · outbound

This paper cites Monotonic infinite lookback attention for simul- taneous machine translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models Monotonic infinite lookback attention for simul- taneous machine translation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.920539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:780205ecde55663ca5dadcd9b6404294708d76508d4f51467a755ee35c7e3d6d

Observation 37548497-bc25-48ac-a701-967a05de8dd2 · outbound

This paper cites SiLLM: Large Language Models for Simultaneous Machine Translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models SiLLM: Large Language Models for Simultaneous Machine Translation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:36.842019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:751d13bd677911b8a6765bca6f046a96ff70c2178bed42067c13ad2c31507bb9

Observation 1dbb89a1-2a4d-4731-8360-684a038f5dd3 · outbound

This paper cites Direct segmentation models for streaming speech translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models Direct segmentation models for streaming speech translation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.923442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:0676b0fe2df02076bb2b276cfecb33fca4dbc052bfd99efd88a6058d4d40dfff

Observation 90cdea54-c8d8-4ce0-8125-0d063c756d4a · outbound

This paper cites Direct simultaneous speech-to-text translation assisted by syn- chronized streaming asr.

Direct Simultaneous Translation Activation for Large Audio-Language Models Direct simultaneous speech-to-text translation assisted by syn- chronized streaming asr

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.930069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:bc8641c56ef454e0ed900f49ae197a13c864f67b24bc6824c5521ddb2a3b40c1

Observation 0d6042c5-28dd-4faa-9b2b-27721b76bb52 · outbound

This paper cites Re- cent advances in end-to-end simultaneous speech translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models Re- cent advances in end-to-end simultaneous speech translation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.963171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:300e997c80ff51beacd8508b793eaebc42fa6f2fe05db6386f1e59e71bb9253a

Observation bb553127-5668-449d-a4fc-53389ee7699f · outbound

This paper cites Learn- ing when to translate for streaming speech.

Direct Simultaneous Translation Activation for Large Audio-Language Models Learn- ing when to translate for streaming speech

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.948997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:441d0219afb21c4d3eeb6b05ca7f6f04001b533369e0234e93660587536138ab

Observation 52efd8da-7533-4a91-80de-676e371f1620 · outbound

This paper cites End-to-end simultaneous speech translation with differentiable segmentation.

Direct Simultaneous Translation Activation for Large Audio-Language Models End-to-end simultaneous speech translation with differentiable segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.933369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:4c94c92f29bc769033cbd208418dcb86b6d9c4e26c007fe18298a560c44944d3

Observation b12c2c22-b541-4d44-99e0-b69b290bb478 · outbound

This paper cites Decision Attentive Regularization to Improve Simultaneous Speech Translation Systems.

Direct Simultaneous Translation Activation for Large Audio-Language Models Decision Attentive Regularization to Improve Simultaneous Speech Translation Systems

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:36.847157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:1185859d69e67c337481632f8d8a643fa680823919bfc6ab265f39c0466d3912

Observation f1d534ab-5f6f-4389-b757-8df64a980a76 · outbound

This paper cites Cross attention augmented transducer networks for simultane- ous translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models Cross attention augmented transducer networks for simultane- ous translation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.952525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:01be83901db44c856e26912fdcda44126bce808f9aef2187c005d3bcb522a905

Observation 25dbe68d-7a18-42e0-ae39-7bce12e87ed0 · outbound

This paper cites Incremental Blockwise Beam Search for Simultaneous Speech Translation with Controllable Quality-Latency Tradeoff.

Direct Simultaneous Translation Activation for Large Audio-Language Models Incremental Blockwise Beam Search for Simultaneous Speech Translation with Controllable Quality-Latency Tradeoff

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:36.836102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:2b52aceefc4156d217bc14f10cbcfe9b561451f412dbeb3d196a8da0acc4026a

Observation bfe18e1a-eb72-417e-8548-7c9c41e6a4b9 · outbound

This paper cites Attention as a guide for simultaneous speech translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models Attention as a guide for simultaneous speech translation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.945992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:e50d39a9262f2efd7bcdee695143f9896e51f203c3ba087b4ab5255687309efc

Observation dd84d51e-7157-4d6f-a6f6-17f81ad08e06 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Direct Simultaneous Translation Activation for Large Audio-Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.817162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:2996c3d52e837627d3c763aac61bdbe184723af3e875247e07f21613d0992394

Observation dad0e699-91aa-4189-9fd3-e402380beab9 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Direct Simultaneous Translation Activation for Large Audio-Language Models SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.829565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:8072ebd7db0672ff9f711ac25960f3d9567311c2290f06b567780aa0e19c7e40

Observation 24cce4b3-79af-445d-9dcd-fb45577f3dec · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

Direct Simultaneous Translation Activation for Large Audio-Language Models SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T16:31:36.857609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:bb657bc8ace53224be803c91dbb6a92d2dd65f20665df1752ed1ff7e76adbea3

Observation 68f794ce-c0db-41c7-84ec-95c5ff9184f5 · outbound

This paper cites Qwen2-Audio Technical Report.

Direct Simultaneous Translation Activation for Large Audio-Language Models Qwen2-Audio Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.811496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:a079e1a0145a0a1d3859bbef19654d4f9bb8b0d31a55e0e69940e6afa4114954

Observation dccab124-2081-4da8-a6c9-7aef4a5e5b57 · outbound

This paper cites Alignsum: Data pyramid hierarchical fine-tuning for aligning with human summarization preference.

Direct Simultaneous Translation Activation for Large Audio-Language Models Alignsum: Data pyramid hierarchical fine-tuning for aligning with human summarization preference

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.959778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:6df86bed8189e10e4600e893570590a2e5463222544e9ccb74e57c50175076a8

Observation ed27091e-7587-41cb-9ab0-3eb6594f678d · outbound

This paper cites A generalization of the beta distribution with applications.Journal of Econometrics, 66(1-2):133–152.

Direct Simultaneous Translation Activation for Large Audio-Language Models A generalization of the beta distribution with applications.Journal of Econometrics, 66(1-2):133–152

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.956392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:c6f016bff5ae255fe1afe639ba713560e1c431e182d849ae299dee5c24b9d4b5

Observation a0b788c3-36ca-41ae-8cef-57088ee23901 · outbound

This paper cites Fast infer- ence from transformers via speculative decoding.

Direct Simultaneous Translation Activation for Large Audio-Language Models Fast infer- ence from transformers via speculative decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.942974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:7d5d826788c42765dd251f215672640fcda4fad9503914a47cbe7ce0c058da73

Observation 875d1873-ea61-4487-9dbe-57d79dfc656e · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:31:36.805785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:29e706f6869693899f39ec6a9eee05613f28e005f11da40de99a3ac1d6184614

Observation 9d2b1a91-b90d-4eee-b643-2644c93f22e5 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Direct Simultaneous Translation Activation for Large Audio-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:36.852025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:155f8e6a229f6bee8b4e2b32fad3b9896e17e2d97561576f74533252427e5972

Observation 3aecd6b4-c54f-439b-9159-5802e76e9c34 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Direct Simultaneous Translation Activation for Large Audio-Language Models Bleu: a method for automatic evaluation of machine translation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.936665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:490e2d09e64da721b9fd44ca1755c446f15ca4b4d7aef59621154190ab18ecdd

Observation 5643e2c3-670f-4397-9b32-9885366aefdd · outbound

This paper cites xcomet: Transpar- ent machine translation evaluation through fine-grained error detection.Transactions of the Association for Computational Linguistics, 12:979–995.

Direct Simultaneous Translation Activation for Large Audio-Language Models xcomet: Transpar- ent machine translation evaluation through fine-grained error detection.Transactions of the Association for Computational Linguistics, 12:979–995

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T16:31:37.940011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T16:30:39.575148Z digest=sha256:a6f3dc1d12b32f2669acef48011448b9150374af2fac38762565d5d8feefc550

Pith citing papers

No inbound Pith citation observations are available.