Pith. sign in

Paper Citation Record · LEDGER

Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2406.10082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10082 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:14.631987Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T15:16:18.164748Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation efe3c1e0-9f9d-4815-b1e4-db2a7d387273 · inbound

Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications cites this paper.

Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:09.668071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:09.668071Z digest=sha256:2ca052811e464dde3e9cdb1dbae1c900d3b88726abf88bc03a479c0677b8f8fb

Observation 14646b13-963a-4262-9c08-d723ac268e36 · inbound

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization cites this paper.

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:14.631987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:14.631987Z digest=sha256:3753aee64031e1fff4bd14b06482a1ef0f48bf97bebb78e7051d3b77b09cfe21

Observation 282af02d-6efd-4ae8-acd3-9cbcfad85f63 · inbound

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses cites this paper.

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:54.187906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:54.187906Z digest=sha256:4d17ec519f0dfa50b549a742ad92c1e0b2cc42cb2f8f922f4ae11763b666339b

Observation a334828b-cd9f-4cb2-a8e6-a0c0914eb90d · inbound

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition cites this paper.

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:12.049729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:12.049729Z digest=sha256:92e004056ad16c080ce10868e9c6c49d19d4052dad1b3fa946f85db2619a0ff2

Observation ac1134e4-820d-41c1-8371-fd135c4d2cd0 · inbound

HumanOmni-Speaker: Identifying Who said What and When cites this paper.

HumanOmni-Speaker: Identifying Who said What and When Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.527913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T00:57:12.239390Z digest=sha256:f2e163e808d5692cab3890e8208071661208658fb0e5216983d0f4b3123a4595

Observation f39153ad-0790-4acc-a295-d007ff5aaa2a · inbound

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders cites this paper.

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-09T15:16:18.166295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T15:07:52.715880Z digest=sha256:8a5441a4bf48872335d62c0cbc185ee04d451523395b3c95e7cd2587608e34d7

Observation 7331b4a1-1972-4a67-badd-7c70241832a2 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.598277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.598277Z digest=sha256:512c46f976cfe9d6c0472b0037aaed5894420f496ed0a4fd6c997675d62f516c