Pith. sign in

Paper Citation Record · LEDGER

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2503.03983.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.03983 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:05:27.317435Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ad9a24e6-3bfd-487b-8adb-81255a67fd56 · inbound

BLAB: Brutally Long Audio Bench cites this paper.

BLAB: Brutally Long Audio Bench Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:27.317435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:05:27.317435Z digest=sha256:ac51920ced6b2f49cda082729162a50da13e9e1ce746a413ac9941f1766223ca

Observation 87703216-962f-4ed3-be43-2d13eac1b752 · inbound

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix cites this paper.

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:12.593547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:12.593547Z digest=sha256:95fab4907f6ab5b72ff5598fc4757c6dca50b149f26469758dea4fb2fdfce94e

Observation faaf1430-0060-4ed1-99b9-2e4d8b354f35 · inbound

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding cites this paper.

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:06.028974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:06.028974Z digest=sha256:462ac5277634da868f7096a5759209e564041adf43e5b310d73cd200f29a1357

Observation f78fe07d-4bdf-4cf2-8357-5353d91130da · inbound

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following cites this paper.

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:50.218454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:50.218454Z digest=sha256:736a172776a889b8141979ce1b768adb523dbea6039664394d9d74cc07a61635

Observation d6eb32c1-0faf-4999-86af-f60a5b54c463 · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:34.161298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:34.161298Z digest=sha256:78f6f1eaf278677f5bfaf11ac075d6098a375898da7b790e0823162449ffc049

Observation 9daf461b-eab3-46d9-8e7c-7a68f1a3d7a1 · inbound

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding cites this paper.

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:24.546999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:24.546999Z digest=sha256:6c774e61d1f47f4b2c2c0fce0da3909fa5ece9b12cf9965ebc929dbb34c32da1

Observation 2482effc-92ad-4225-a722-cc2ace70b7e1 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:02.450734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:02.450734Z digest=sha256:53781c9dce6accf5c251cb6f9bc89794365620cc84d2ea801f0ed5f071eca818

Observation cc18c3e4-6860-40e5-82f7-e582062a4df1 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.015617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:407d7de9c778cca4fabf273ea1faa2312c29353760ba1a8fbe3979764bd3013a

Observation 790b06dd-25e5-44b1-8942-25f3c8f43102 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.261844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.261844Z digest=sha256:6025c2418bc24b61c61daffed6fcbaca1b838cda00d86cf5ec3c38f1f23f1f8d

Observation ce6d2a7f-0377-48f8-9a32-3893dc6f4dcb · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.971766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.971766Z digest=sha256:25df9fc42fe875b31ed38e971906bbd613cf58a0e8d0ffe584c6932bcdddc0b4

Observation 1ca8e5ab-f5dc-433b-90a8-29a0b028f2ad · inbound

WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations cites this paper.

WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:43:27.210402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:43:27.210402Z digest=sha256:b3a0d6822adb0cc6033d5b2ec01530c454f0cc90de7004182cd9202bd7c49538

Observation 1da3cfc8-93d9-40e5-9552-ac757a222577 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:00.050889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:00.050889Z digest=sha256:66d40a7c0b7634c5133ae097102e5898bfa753aa4dc2965951b8f31eeab0c33a

Observation 9e029be1-4592-4b6f-a584-2dbfc5d5dded · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:10:43.108380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:f48dae5c63addaee7bd1e130d1b4c655ca684a614ac7d70644853fea1036a9ee

Observation 2652559e-6df6-4a18-abe9-e3b7c8d37e5c · inbound

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding cites this paper.

SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.631705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:54:52.275011Z digest=sha256:7d6e16589bb31a84d842d8a88101c5a46f38968c929eba05addf4579f2396a44

Observation 59f72fc5-5d63-4c36-8b5c-8dbe497526bb · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.051834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:6a140accaace6f51a04e3b8112a0884f73bbc3cc69f60f714e026b64a8a9187a

Observation bf4b9018-0016-4edd-9b2e-362ec6250864 · inbound

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models cites this paper.

HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:31:14.370153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T05:20:30.823304Z digest=sha256:8c62a3d383c7f08584b2b9e01ccb076d5de69488136373c4647a6630a0961b30

Observation 79924e11-1dd8-426b-ba6e-5a07000b7e4f · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.072605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:0431a19305b9fdf0c83466dd22bc13740827dc058c267975604c0dbc687d8dfc

Observation 8016a153-88ae-481b-8b36-36e61735ea75 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.998577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:34faeacdb3d170885e3ceb9fba889a166b74a74fd938b8037a5bbbf6e095f53f

Observation e8ae7566-cd02-4b84-b8da-3f3f3ebf7645 · inbound

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization cites this paper.

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:46.697224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T17:45:10.339950Z digest=sha256:f5035638a7941a6bef9196e7b41285fb4c67fb666353963555a94fa046b35a76

Observation d9c56024-a811-499c-bbbc-f8ed687dbe10 · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.388730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:b37a50426473cd6834471b76cea2539ed7dd08e4a102d3fed8d602a387c44cc6

Observation d46a9922-240c-492b-a61c-4259fec8c112 · inbound

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints cites this paper.

TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.827782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T18:18:41.816342Z digest=sha256:7514d111213f36120f12902f156f333ef9a6dc04ea7f4c8a7086a8997e366708

Observation 3956a836-13a2-4a95-a7af-7bd398394a8c · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T22:15:05.308750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:60549fd23cb4c5b36a000a674aaf5daccb6f54658d3a733ff17221fcf1bf0f1d

Observation 178befd7-2047-465e-a124-f93262f31937 · inbound

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark cites this paper.

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:17:43.954076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T12:11:56.629002Z digest=sha256:fcafcedef7318438661b294ad1678f150c91f70f4343c06368818c40a581c1d6

Observation a152d5c4-cb10-4a61-a202-8841013b14ac · inbound

Continuous Audio Thinking for Large Audio Language Models cites this paper.

Continuous Audio Thinking for Large Audio Language Models Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:18.002241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T21:52:49.839901Z digest=sha256:733d472891956a8cbd73a7b14f6015a49b2f74b8fa18804ed817934e07af9824

Observation 24985778-cf4a-4593-9557-8599f6dd6b6d · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:29.155417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:57d9f7409400484d4696e219f700cc93800e94c5298eb6060a2446315b844cee

Observation b147afc1-d365-467a-aac4-435a39917103 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.332502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T15:51:48.672086Z digest=sha256:aeb1c32d1257899a936d4b861f503ed8164fedbdc5a0fb6b3242b51056a34058

Observation 42d05f29-ac5f-4c73-a39c-3b055e3b27a1 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.004309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T23:36:39.368933Z digest=sha256:2bece3e1b4d5b67ff462e1467610dd13ada675208287023da0e764fdf8b3e5e2

Observation 85c4457d-8f24-4970-8c7b-1551134ddaa4 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-15T10:41:34.347336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:41:34.347336Z digest=sha256:6f12b74dfb6801cfbda7433c2b345b354db2790c3d9e786a840a1f79754f14de

Observation c6db14df-ebe8-4c4c-8a2c-7ec07b92283f · inbound

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models cites this paper.

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:37.193075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T06:47:53.308909Z digest=sha256:ca4a237a568c3175eaa376ea034c897f1448c5f3dcde8d01076b7b4aaeec4806

Observation a17a7f5c-b979-4e1d-bbc6-f5190ad4fc7f · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.631058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:c0c180918adfe8867b67656f08d14028a934aee5c57932674a471b6d5662e278

Observation 87ae1241-5a59-47df-934c-0d21baebe401 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:aab221c5f82bfd741ead033fa91eb66360151a60226b0c0c7e07800663ba6cc2

Observation ea274fe0-9202-4966-b8b5-ad83df298189 · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:8188f90aea7b9d344bca2938f4e933fe793593550bc2be2784d123be04c2358a

Observation ada37a26-5187-4f53-b72a-338eaae07299 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.567045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.567045Z digest=sha256:3c61a8876ae67f2ec8056d97e82e4e3e1fd2c2baadf2092a48ab492d5cb0ccc2

Observation c6893a8c-bff4-45e4-9398-5ea127d52293 · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:25.777254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:25.777254Z digest=sha256:db2d678f0a5c3bd9aae8a8223acdda481f85b2111605fb4b99bb49af5fd637e2

Observation ed269b75-4fc8-4ee2-8213-14e5013f8b44 · inbound

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning cites this paper.

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:19.576273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:03:19.576273Z digest=sha256:4bdd4608f3827889f482ede56e65a89fe20deab6b81a691a52780240d7d1c11b

Observation a6654286-5e25-4ae2-b20f-3e67995c7063 · inbound

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference cites this paper.

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:59.434263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:59.434263Z digest=sha256:e67405e6b0786237760759fb255db54e9e54ad2bf3f062cde72136528ed6f6b4