Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T12:07:39.817307Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2607.10387.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T12:07:39.817307Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T12:07:39.817307Z
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f168dcb9-4262-47a4-a1f4-11ec0dcf65f5 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33842f2d-04d1-407b-b38c-889770613c08 · outbound
GigaChat Audio: Time-aware Large Audio Language Model GigaChat Audio: Time-aware Large Audio Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b0536fe-e30e-4ccc-a413-ea01700fdf23 · outbound
GigaChat Audio: Time-aware Large Audio Language Model audio tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5957e482-b7cc-406f-88d5-a7ba160f0804 · outbound
GigaChat Audio: Time-aware Large Audio Language Model We then obtain word-level timestamps and time-aligned transcripts us- ing WhisperX; the same alignment is also used to estimate the silence ratio [9]
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19afe744-d780-4268-9e35-1c05d1bbffed · outbound
GigaChat Audio: Time-aware Large Audio Language Model when was this phrase spoken?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fcfa865-2e5b-41ff-9676-d7d5ea0ddc21 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Our results show that temporal ground- ing in long recordings remains a major bottleneck for exist- ing multimodal models, which often degrade sharply beyond a few minutes of audio
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d5a6be-3f2b-462a-b693-dec2037c2a68 · outbound
GigaChat Audio: Time-aware Large Audio Language Model All outputs were carefully reviewed, edited, and validated by the authors, who take full responsibility for the fi- nal content
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ee9c40-d3e6-4413-9091-9f65e56682f5 · outbound
GigaChat Audio: Time-aware Large Audio Language Model GPT-4o System Card
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4edbcfc4-8f24-4d2f-9deb-0d23d01d9a9f · outbound
GigaChat Audio: Time-aware Large Audio Language Model Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382579c1-70e4-4e99-9485-5309bbe4cacb · outbound
GigaChat Audio: Time-aware Large Audio Language Model Voxtral
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181f658a-4ba8-4757-9d4c-c87aa2351753 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Qwen3-Omni Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23072848-deb8-4ca2-822c-398e9065bf9e · outbound
GigaChat Audio: Time-aware Large Audio Language Model Qwen2-Audio Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33910b4-a0e7-41e4-be13-d211839ecf5c · outbound
GigaChat Audio: Time-aware Large Audio Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0ec7e5-e52a-43af-9672-b2ebe74c0948 · outbound
GigaChat Audio: Time-aware Large Audio Language Model VIBEVOICE-ASR technical report,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956adcb4-f2cb-4e1f-991c-50c6418120c3 · outbound
GigaChat Audio: Time-aware Large Audio Language Model MOSS Transcribe Diarize Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d20c44b-129d-4788-8e73-21dec973487c · outbound
GigaChat Audio: Time-aware Large Audio Language Model WhisperX: Time- Accurate Speech Transcription of Long-Form Audio,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64470285-107b-4ffa-905e-a9d43e62a6f4 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Text-to-audio grounding: Building correspondence between captions and sound events,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bddfef2b-46c1-47e4-aa6f-8dd1f15e28a0 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Multi-domain audio question answering toward acoustic content reasoning in the dcase 2025 challenge,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f734a812-218c-4b59-8c57-19d28005e72b · outbound
GigaChat Audio: Time-aware Large Audio Language Model Listening between the frames: Bridging temporal gaps in large audio-language models,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f9c7df-9395-4267-ad9c-5e4646563e15 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Language-based audio moment retrieval,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4936fe5-da6b-43c1-9f01-6b0d1f8faf95 · outbound
GigaChat Audio: Time-aware Large Audio Language Model CASTELLA: Long audio dataset with captions and temporal boundaries,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2967d62b-6b03-4505-9015-7fb120dc3c0a · outbound
GigaChat Audio: Time-aware Large Audio Language Model Localizing moments in video with temporal language,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd481eea-dbef-44cf-822a-b25a6070bf61 · outbound
GigaChat Audio: Time-aware Large Audio Language Model TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fde91ef-89e5-45eb-9c5e-fe6bed15d4b4 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Longcat- flash technical report,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d056bb-fd5f-4617-8f28-3514bbe4cb5e · outbound
GigaChat Audio: Time-aware Large Audio Language Model G-eval: NLG evaluation using gpt-4 with better human alignment,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd975aea-fda3-4cb8-b2b1-a0fdb6c440ad · outbound
GigaChat Audio: Time-aware Large Audio Language Model GigaChat3-10B-A1.8B,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71ceff2-d8d6-4a2f-bd46-c29657062437 · outbound
GigaChat Audio: Time-aware Large Audio Language Model FlashAttention-2: Faster attention with better parallelism and work partitioning,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b446f7d-8518-464d-9423-86fdcdd879f6 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd718bb8-c590-4123-a3d2-88cf791bee38 · outbound
GigaChat Audio: Time-aware Large Audio Language Model GigaAM: Efficient Self-Supervised Learner for Speech Recognition,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95994720-37d4-4408-a38a-7a12968103e2 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfcd6b06-7d83-41d5-a102-c0c0ac659c87 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Clotho: an audio cap- tioning dataset,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7915f50a-b080-4559-bcef-400df359bc7e · outbound
GigaChat Audio: Time-aware Large Audio Language Model AudioCaps: Generating captions for audios in the wild,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b981f732-048d-4360-9f6d-2323a0a1bf93 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Yodas: Youtube-oriented dataset for audio and speech,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c1a9d04-01cb-4ef9-8643-d964f6798f57 · outbound
GigaChat Audio: Time-aware Large Audio Language Model SpeechBrain: A General-Purpose Speech Toolkit,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 808dd49e-80f2-4d31-90a6-256d5f8f1a9e · outbound
GigaChat Audio: Time-aware Large Audio Language Model V oxlingua107: A dataset for spoken language recognition,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ca7311-6920-44be-962a-65090d8088f6 · outbound
GigaChat Audio: Time-aware Large Audio Language Model gpt-oss-120b & gpt-oss-20b Model Card
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6537162a-083f-4e8b-8cf3-93cee8c186db · outbound
GigaChat Audio: Time-aware Large Audio Language Model A Comparative Study of Quality Evaluation Methods for Text Summarization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5f9441-2ec7-436e-b905-7dff63a249fd · outbound
GigaChat Audio: Time-aware Large Audio Language Model Summeval: Re-evaluating summarization evaluation,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f064eb1b-c133-4a0b-acaa-f5bda8e28628 · outbound
GigaChat Audio: Time-aware Large Audio Language Model Unleashing the killer corpus: experiences in creating the multi-everything ami meeting corpus,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33842f2d-04d1-407b-b38c-889770613c08 · inbound
GigaChat Audio: Time-aware Large Audio Language Model GigaChat Audio: Time-aware Large Audio Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.