Pith. sign in

Paper Citation Record · LEDGER

Ola: Pushing the Frontiers of Omni-Modal Language Model

As of 22 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 35 inbound Pith citation observations for arXiv:2502.04328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04328 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.389717Z

measured 115 of 115 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:04.401662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.266999Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36cc8bf1-f639-4d93-9b12-f766ef5737cc · outbound

This paper cites MusicLM: Generating Music From Text.

Ola: Pushing the Frontiers of Omni-Modal Language Model MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.033803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.033803Z digest=sha256:3d967ccd6619d4c7446d86513ae48adf40158442001bc6a9958260202fe19f5c

Observation 96c0d92a-1af9-4948-8e4b-d2c433f5928c · outbound

This paper cites Pixtral 12B.

Ola: Pushing the Frontiers of Omni-Modal Language Model Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.039633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.039633Z digest=sha256:876356719d0351612ed53296c38faf8f18cb5a73699cd3cc75dbb156f03acb1d

Observation e4eb3386-1687-4e99-a13f-255dd36cfdba · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 35: 23716–23736, 2022.

Ola: Pushing the Frontiers of Omni-Modal Language Model Flamingo: a visual language model for few-shot learning.NeurIPS, 35: 23716–23736, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.044339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.044339Z digest=sha256:ea53ae6cf7763dc381068a1b3c52902d7ce86652d459728e91c120321ea3651e

Observation c298e9bf-9dda-47d5-b750-d6aa875e20e2 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Ola: Pushing the Frontiers of Omni-Modal Language Model X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.048721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.048721Z digest=sha256:28ae79be99fc91425f34eab5fb969aefe7f28030ccdaf62da40553ca1a1d927e

Observation 339dcc89-68b2-4e8b-ab24-4856fa4e70ca · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Ola: Pushing the Frontiers of Omni-Modal Language Model GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.053243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.053243Z digest=sha256:6586c0f2d2944a0d451b60654571ba4450a1ebb072ee3bbca411f512d603eafa

Observation 1d05a51c-ac15-48bd-bfb8-255fc99a0ebd · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Ola: Pushing the Frontiers of Omni-Modal Language Model Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.057732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.057732Z digest=sha256:c778c2a6f0422546abde44ca1a0f0acbaf66f10ab985275b8d99731c35b2fdca

Observation 84909031-5f6e-4f8d-8289-8991b2a172e7 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Ola: Pushing the Frontiers of Omni-Modal Language Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.062734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.062734Z digest=sha256:112e61c2c9fc7442c1afd8c599c7a100befff63a3b07e7af196725217d6b9a89

Observation b69afcea-1c75-4a4f-8d2d-068d06171ca0 · outbound

This paper cites Beats: Audio pre- training with acoustic tokenizers.

Ola: Pushing the Frontiers of Omni-Modal Language Model Beats: Audio pre- training with acoustic tokenizers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.548206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.067887Z digest=sha256:505946b0d0e19f937af5eaedc36cc6ef3cd0e4fe1f4913c02344fc967a8d6204

Observation d3d61bf3-5cef-4116-b793-1bcec186b665 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Ola: Pushing the Frontiers of Omni-Modal Language Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.073045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.073045Z digest=sha256:35173367d87310a15fff240a271634d914cbb6d2b1708dcf73d63189f7c7a50c

Observation aa5fd178-274f-479b-b61a-10d16037da86 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Ola: Pushing the Frontiers of Omni-Modal Language Model Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.533798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.077800Z digest=sha256:05e1fb01fc271f6bf60412d948eae5c874fdfcb368cf80e615c2eb6b7f580c82

Observation 0d6deb94-f2a1-45f5-a16b-c5b2bd9df8e6 · outbound

This paper cites Qwen2-Audio Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2-Audio Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.082294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.082294Z digest=sha256:9349e1b3cd90967c34cb05b6328254214ad2039a5967ac37206a1b9110a8b98b

Observation a8095656-7376-4254-90e7-aa126aa79d0f · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.087369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.087369Z digest=sha256:32d57a0921a92de8f06be002cb1674cde206697bfe91f687762178b4a4f1bec9

Observation be010433-2ee2-4dd2-8caf-8680846187d5 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.092001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.092001Z digest=sha256:3c34fffcff3fd5cd091750b44ed0136f9edf56763c7a121f2985449f0e34b2ed

Observation c4ac83af-cfa6-4765-bf09-a7c46045218f · outbound

This paper cites Clotho: An audio captioning dataset.

Ola: Pushing the Frontiers of Omni-Modal Language Model Clotho: An audio captioning dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.519585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.096858Z digest=sha256:62930d27dc303faf02c3c9b4c54574220c066b8869d6d271c6c55c01c6766e13

Observation b2c86b57-3462-43f1-a363-34c80cbe1f48 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.101465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.101465Z digest=sha256:5e0df08edbadbaccdf9d3ac34c58554204a8250de59d515db8025c469ca38c61

Observation 7dc3286d-0748-4cc8-888e-601fe4d49c74 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.105733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.105733Z digest=sha256:ce83085fc6777f27ea04cdf8689fe11035d29a89ebc28fb9222cdc59961c73b5

Observation 9ec2ccbc-8a75-4cb3-9c5c-e8732075f457 · outbound

This paper cites Finevideo.

Ola: Pushing the Frontiers of Omni-Modal Language Model Finevideo

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.495231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.110748Z digest=sha256:9ec465dd0298ebe7312b9647b2795b63650859e3fde910ea9dd1f2a22cc62f6c

Observation b6edf24a-58d0-40ea-83be-7dbb610912a8 · outbound

This paper cites Prompting large language models with speech recognition abilities.

Ola: Pushing the Frontiers of Omni-Modal Language Model Prompting large language models with speech recognition abilities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.480758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.115161Z digest=sha256:18c3ecc692145748301d0920151fd19041fa9b97ab88d55b1284f0c35600914b

Observation 748aef16-b096-4671-98f1-b15f73cda2a0 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.119516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.119516Z digest=sha256:0bdb722b3db983bdc21a5bbd4f2257087a48eb44340bf7afe0762312840fb98d

Observation b20dffb3-3b4d-48fb-95ed-2a4a6244f084 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.124203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.124203Z digest=sha256:9138207e1ca79c1ae9b480540199102f805def175de43c834a95d591ad651f52

Observation 8cc93ee0-cec6-44eb-9cc5-923a6683089f · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Ola: Pushing the Frontiers of Omni-Modal Language Model VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.128649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.128649Z digest=sha256:f829a9d0be4c9b77ea8eba9d2986dbe564b769d8e78bb8e8c922099b108a65ae

Observation 284e8d0c-f632-4704-9b80-307f553280a2 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Ola: Pushing the Frontiers of Omni-Modal Language Model VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.133157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.133157Z digest=sha256:52f29576eb2353874b87526325780c148c1850194ed5444e51f235011496186c

Observation 9190eed9-e584-49e2-b687-6d0b3c4bd1f7 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Ola: Pushing the Frontiers of Omni-Modal Language Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.138130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.138130Z digest=sha256:50ed25d34861ba526d8b9cbb19cd804305a93f4fcf66f21afdfd4010c941209a

Observation 32799e57-afca-408e-97c5-16b96a1b73a1 · outbound

This paper cites Hallusionbench: an advanced diag- nostic suite for entangled language hallucination and visual illusion in large vision-language models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Hallusionbench: an advanced diag- nostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.142824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.142824Z digest=sha256:0f548481612d0de5dd429f9b42925c865197b7f45d91bfbae024ea905a6dd784

Observation f02c1b66-f08c-445d-a481-2881b210d20a · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Ola: Pushing the Frontiers of Omni-Modal Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.147822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.147822Z digest=sha256:eeb97455efa2f104cdc4984a49801454b5b0a92921eb541ea4d260a866cf20fd

Observation 68d3c0cc-b976-4704-b1ed-ae04ece4bc1e · outbound

This paper cites 3d-llm: Injecting the 3d world into large language models.NeurIPS, 36:20482– 20494, 2023.

Ola: Pushing the Frontiers of Omni-Modal Language Model 3d-llm: Injecting the 3d world into large language models.NeurIPS, 36:20482– 20494, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.456893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.152244Z digest=sha256:5e9e7d37a1e36948f9f09528b4673205d4c35e383e09d3bcd7f6ebeb551f44c0

Observation 3286a960-6612-4779-9a7d-7be08835858f · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Ola: Pushing the Frontiers of Omni-Modal Language Model Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.442436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.156237Z digest=sha256:6af6886eb7a37e7a6e91c24dfca31e6c74fd222cde84905f859897677ebefb81

Observation 915809bb-967a-49ec-9b7b-587b34b1c388 · outbound

This paper cites A diagram is worth a dozen images.

Ola: Pushing the Frontiers of Omni-Modal Language Model A diagram is worth a dozen images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.428020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.160152Z digest=sha256:8ffc91fa813ebf6706a8e779c2ac1ddda828cefc3dd4a6787a7572b851621169

Observation cd595f54-6dfd-477a-be1f-5417575e5ac8 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Ola: Pushing the Frontiers of Omni-Modal Language Model Audiocaps: Generating captions for audios in the wild

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.413912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.164136Z digest=sha256:16515f0ca81bde3c1081b642893833734041f43f4aa8d0ac763f161f8482c8d9

Observation 094c9579-4d98-4f7c-ac56-95046e30840c · outbound

This paper cites What matters when building vision-language models?,.

Ola: Pushing the Frontiers of Omni-Modal Language Model What matters when building vision-language models?,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.168130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.168130Z digest=sha256:985ad4b4864986319c9942a0fc42c620b716fc40074428bbd382e55036d15aba

Observation 2ccad04d-6854-46c4-9f16-3d646fad8786 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.172390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.172390Z digest=sha256:a2830e97230fd4e5e01b3fdcee4067737b1faeab7f88c0d650b9e41dd75f4568

Observation c7a3cc94-b33c-4581-ab86-4b75029c9ab9 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.390244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.176335Z digest=sha256:b2268ef45b52cee3a64801411156a09f01d00f9882794d5955cfe569be53d21c

Observation e2c42e2c-b04c-4a94-8f70-42b3e1f71714 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.180712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.180712Z digest=sha256:7ba8f9b7901055c99d01894c7970e2f67420a2228611134229ec23785178661d

Observation aa6285d0-f875-4fac-b0b9-9ac1cdc58209 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.185155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.185155Z digest=sha256:b6ede1043bf548bc4b258d373afdec7b118be4eb21594c4b0a57bb7e388cfacd

Observation e1c9e7d7-3850-434b-b9d2-bda94cdf279d · outbound

This paper cites Improved baselines with visual instruction tuning.

Ola: Pushing the Frontiers of Omni-Modal Language Model Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.376428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.189638Z digest=sha256:2f3d6aee2c8caa4356b64014e9aec69a519688d103d794a90ea641a24e1f90c8

Observation 42a524a5-c017-416e-b7d9-00b89c4ce79c · outbound

This paper cites Visual instruction tuning.NeurIPS, 36, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Visual instruction tuning.NeurIPS, 36, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.362427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.193766Z digest=sha256:eb0b2eaabda033dfe08dcfee109c284ce446f39d5757f8dfd8830c7f9f1ef291

Observation 6a7e792b-09d0-461e-8e87-0be6e8c963b9 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Ola: Pushing the Frontiers of Omni-Modal Language Model MMBench: Is Your Multi-modal Model an All-around Player?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.197952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.197952Z digest=sha256:f379ae85f4d9f0ec5b30183f65f307297dfee985a261d4a74ab4a28030aa7f4e

Observation 3f1ea753-738a-42b7-9dd4-3a7ad0fe66d2 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.202509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.202509Z digest=sha256:7af4c9be6c9d65be59e26a858d86148535df8063d1f05a57007cc18714eaf9c3

Observation 0d28cb45-166f-4728-9025-ba01d2f44225 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Ola: Pushing the Frontiers of Omni-Modal Language Model Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.207697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.207697Z digest=sha256:d4846bd356f8323120d2fe8e110701b44f27345135eb31f85d19328b8438e63b

Observation 12ae62c6-8899-485a-86ff-e0ce6deaca68 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.212186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.212186Z digest=sha256:dcc8c2503879855f77bac623638803da66e839fbee5710972981412857334db5

Observation 2ef09257-3459-4409-ae06-ddae9877efd8 · outbound

This paper cites Efficient Inference of Vision Instruction-Following Models with Elastic Cache.

Ola: Pushing the Frontiers of Omni-Modal Language Model Efficient Inference of Vision Instruction-Following Models with Elastic Cache

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:47:39.762001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.216914Z digest=sha256:c6814a4047d8458e4aa6e16219d08f1147eb0c2547d8f907968136ba07020f46

Observation 285adbf9-99ec-436d-88a5-531da218938b · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Ola: Pushing the Frontiers of Omni-Modal Language Model DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.221310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.221310Z digest=sha256:3c5b5e450a4b615765837dd2e6e37454b6313798a361b47ff4a4963f86cdde23

Observation 7bbe32ce-e38a-4d57-af59-e3947f23078f · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Ola: Pushing the Frontiers of Omni-Modal Language Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.225870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.225870Z digest=sha256:aae3eedc01e2f30447655ce974176b7b566f57c0076d470eeec7bd480d2da708

Observation 73c510a9-e742-4f9e-b557-48e913ef9a77 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.230344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.230344Z digest=sha256:2662fc81795e61702c46f0f7c2c09de1d4d5ffb77f807cc60b59cb9f6744519e

Observation c7000482-891b-43e8-9b24-80a146338ddc · outbound

This paper cites The million song dataset challenge.

Ola: Pushing the Frontiers of Omni-Modal Language Model The million song dataset challenge

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.348049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.234810Z digest=sha256:c24ee7f0f6add839d970d09e9875d3ec91cd5366c91557e80247846ad8eab1e3

Observation d33f4b2f-646c-4215-9fd7-429c655d951f · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multi- modal research.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multi- modal research.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.333894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.239231Z digest=sha256:26df9826230563a264b0c385be50d84e7096161609ce45916a90fe0c59b1f245

Observation 468d8b70-41a2-4d10-9a9e-cf29947a2a4f · outbound

This paper cites Openai gpt-3.5 api.OpenAI API, 2023.

Ola: Pushing the Frontiers of Omni-Modal Language Model Openai gpt-3.5 api.OpenAI API, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.319172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.243492Z digest=sha256:8de46cc8a2075af91b5d025f9b23f44eed28044f3bbd50babb678d13fdaa16f6

Observation f5ac6ca2-80f7-4bab-b80c-c1ac36009b3b · outbound

This paper cites Gpt-4v(ision) system card.OpenAI Blog, 2023.

Ola: Pushing the Frontiers of Omni-Modal Language Model Gpt-4v(ision) system card.OpenAI Blog, 2023

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.305194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.247874Z digest=sha256:fccfcdc9bcf20c4f291b8f785ec4b7149577cdd760c01d386db55b58f1d6d450

Observation b14035cc-a05b-41da-aa4a-11279cf1e9eb · outbound

This paper cites GPT-4 Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.252511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.252511Z digest=sha256:227ebbe613972f219579e9adb43034b5f3851497a0bec90b4be26578243f44bb

Observation f2a69b28-10c8-4593-8d99-0192b7704f5c · outbound

This paper cites Hello gpt-4o — openai.OpenAI Blog, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Hello gpt-4o — openai.OpenAI Blog, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.289737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.257343Z digest=sha256:c0eb8916ebc1acdff2cd2fd84c11d7d299ec76a079f4a23fb37ba7f5844c0abb

Observation 2d958a66-1976-49d6-a5c8-c63f236b158d · outbound

This paper cites Librispeech: an asr corpus based on public do- main audio books.

Ola: Pushing the Frontiers of Omni-Modal Language Model Librispeech: an asr corpus based on public do- main audio books

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.273985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.261725Z digest=sha256:303cb348d6580049ca4cc33a9d87ae71ff40ad8f4d6f5bad7aec186e966560a1

Observation ebde3cab-2cd6-4a00-bb8a-fb189d31bb3e · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model Streaming Long Video Understanding with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.266066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.266066Z digest=sha256:e6b523286cf15968daafdbf4ef75b5832eb592a3bd06fe73bfde847a3b541aa3

Observation e6e01e05-6c44-437d-b3ea-3e9c669289ca · outbound

This paper cites Qwen2 Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.270605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.270605Z digest=sha256:c3c84268080bd260e39557a69a395bdb381a145756ae50b0f53e82e4dcea6df9

Observation dbf09d75-de8a-41a8-a55e-6ba4d5f982f0 · outbound

This paper cites Qwen2-vl: To see the world more clearly.Wwen Blog, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2-vl: To see the world more clearly.Wwen Blog, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.259014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.274734Z digest=sha256:8993d1717ca11dfd0134d13a811918a8a6562c1683a99ceef102e74249f9e384

Observation 5135a1fb-ca51-431e-90c1-90af14ff915c · outbound

This paper cites Robust speech recog- nition via large-scale weak supervision, 2022.

Ola: Pushing the Frontiers of Omni-Modal Language Model Robust speech recog- nition via large-scale weak supervision, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.243530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.278590Z digest=sha256:bfb42c5e6f5cb0ece3613343841e1318f92f2527b02868266902e88a3db3d53b

Observation 6bc40731-e56a-4617-ab74-3b1a1ff2a2bd · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Ola: Pushing the Frontiers of Omni-Modal Language Model Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.228446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.282567Z digest=sha256:218c39f152916eb04a8218364d712e9095c003573863af3418ae87aa1cd909fa

Observation ba1cb269-c287-4983-b8a5-afd2497a9b36 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Ola: Pushing the Frontiers of Omni-Modal Language Model CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.286819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.286819Z digest=sha256:030702e007f30a7ca244c99022f1323b1e6e5b52a9d59478cd162753d46fb2c1

Observation 6a7c2c05-3021-4866-989e-068aa7196854 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Ola: Pushing the Frontiers of Omni-Modal Language Model AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.290979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.290979Z digest=sha256:65536565b30a9b88d120d43cfd004627a293a35770253aa547e20c509c662e9f

Observation 92bffe21-3310-495a-91bf-cb568b7dbb45 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

Ola: Pushing the Frontiers of Omni-Modal Language Model MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.295397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.295397Z digest=sha256:24a4d86cb01ab2c6675292dce8b2045a33ba17c6b02b1c0d3b1af5b0f56a47d4

Observation f4e70091-ef1b-4385-8c25-f8fa97886be9 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Ola: Pushing the Frontiers of Omni-Modal Language Model LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.299864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.299864Z digest=sha256:c4f9ca80cb2b44db41b817d50fcfc00ef88b8f0828da82968cc9e5fc2578d61d

Observation b62ffc24-67a2-48c9-8536-aabd360bf92c · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Ola: Pushing the Frontiers of Omni-Modal Language Model SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.304398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.304398Z digest=sha256:85b53bb9593ecd4dbc9d07209abe137b99c95096c48c43c62eddf3914ec56e00

Observation 8f29f496-73dd-4629-a25a-71cfb294925d · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2.5: A party of foundation models, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.213910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.308845Z digest=sha256:b4e8efc6d5188b86fd93eb164726efdb7b316d2baa46b74a55acf2b216ec3a91

Observation f02690d6-cb74-4241-99e1-3fd1f20758c7 · outbound

This paper cites Qwen2.5-vl, 2025.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2.5-vl, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.312940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.312940Z digest=sha256:bf8680a0dfa5ff77e390ac9be975256c1a3f64a8ca4510e66bef72b9c9d5f7aa

Observation 9655bb93-6ecd-4ba6-a0df-4dc8aec363b0 · outbound

This paper cites Learning Features of Music from Scratch.

Ola: Pushing the Frontiers of Omni-Modal Language Model Learning Features of Music from Scratch

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.317222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.317222Z digest=sha256:569c01222408d56052229f3aef4aee00fea03c4b00168cee69b6581b363c923f

Observation dcf6dc59-a89d-43bf-9aa5-a1cf5dd525de · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Ola: Pushing the Frontiers of Omni-Modal Language Model Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.322045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.322045Z digest=sha256:5be8f40769932fa39bcd5b9bd2e0be26902afc4e29bc75210b095f2d12a74317

Observation 21dda9e6-26c1-4c46-ab0d-76e8d652c64f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Ola: Pushing the Frontiers of Omni-Modal Language Model LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.326631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.326631Z digest=sha256:6ce7f62116bd5fbb808b0df93cae0705c1699e9aae3905488e3c643f2ec1c40a

Observation b5dbb05a-f09a-4431-9b3f-2c114c7bc473 · outbound

This paper cites On decoder-only architecture for speech-to-text and large language model integration.

Ola: Pushing the Frontiers of Omni-Modal Language Model On decoder-only architecture for speech-to-text and large language model integration

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.186955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.331408Z digest=sha256:01ce9665c5e8420bab26cfc52c7e36c6f647724e907f4dee673e3356e3629ecf

Observation e74f88dd-b47d-41e1-bc18-36eb5ea138e0 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.335713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.335713Z digest=sha256:d315f2bd3eb69139ebe415d07a2d88661c9b048b38947773783965084dc72dad

Observation b9c0e938-8242-4b36-a5e8-c13da6e66357 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Ola: Pushing the Frontiers of Omni-Modal Language Model Qwen2.5-Omni Technical Report

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.340402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.340402Z digest=sha256:bbab008492db5f83a39f247c1ca4a06e18a2b195f4876a0693b3b6a2b8249002

Observation 8d754839-e1c2-444b-b95f-24a45f959baa · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

Ola: Pushing the Frontiers of Omni-Modal Language Model AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.344979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.344979Z digest=sha256:ca6911549866dc190b8ceb64a70679414e8249a93bba7bce7a97da37f5ee238c

Observation 1e41759d-1964-4f45-a720-32c7192f653a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Ola: Pushing the Frontiers of Omni-Modal Language Model MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.349659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.349659Z digest=sha256:c18193cbc43e1d8d477ec24d277d345b46efe0ecc5858d7baff404ec8c8bae7a

Observation 4fa3a959-f049-4aa8-8f20-891d7b801c66 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Ola: Pushing the Frontiers of Omni-Modal Language Model Yi: Open Foundation Models by 01.AI

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.354193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.354193Z digest=sha256:ed08450cfe71a024bc24e73f1211774405b6fd576856edd5a34c5d4ceb108cef

Observation 8b956f44-8267-4cab-af4c-f303836b5dbb · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Ola: Pushing the Frontiers of Omni-Modal Language Model Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.172036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.358992Z digest=sha256:9a7322c7bc37536d577bc2e385cc07e5c80057a13635cbbef906026d9be073af

Observation 4024c18d-223e-408b-bbd8-b86158815af2 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Ola: Pushing the Frontiers of Omni-Modal Language Model LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.363415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.363415Z digest=sha256:922c0a27b94ca598a56d1c1592aedd3669587d2d410e8a69bfc7eb5ccda5437d

Observation 15082f6f-e3ea-42c3-b2b2-4273c8c5f284 · outbound

This paper cites Sigmoid loss for language image pre-training.

Ola: Pushing the Frontiers of Omni-Modal Language Model Sigmoid loss for language image pre-training

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.156938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.368080Z digest=sha256:62cee899b2b40469d328a18689c5b8b019bb5869d731a1e684eac95ab2e804b8

Observation a5d9ac97-e599-4e74-ab34-e903bb88d8a9 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Ola: Pushing the Frontiers of Omni-Modal Language Model SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.372296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.372296Z digest=sha256:e30cb3184e0963e30fb18feaee031b3b41d8ba82c95ae0ca33051f08e6fee2b9

Observation c4b67586-0fda-4dfd-a13e-849b0c11f5c8 · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

Ola: Pushing the Frontiers of Omni-Modal Language Model InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.376946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.376946Z digest=sha256:4cfa4fcdb2211e4918079fbeaea843dcd1c4aed6d306f62e9736b489250de4e4

Observation 005b69e8-1a87-462e-a221-1724ddd9ccd9 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Ola: Pushing the Frontiers of Omni-Modal Language Model Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.381215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.381215Z digest=sha256:77461179763afd2fd19715f22cd481b1be7984b42d5170135d2fbf9f64a35e16

Observation e7907262-478e-44d1-b798-74a35c42a3f6 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Llava- next: A strong zero-shot video understanding model, 2024

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.142589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.385607Z digest=sha256:9c92883f3063a59518489564c13c2a188a5fd0936c8e23c65f734c2172a4bd30

Observation 77d877b2-5f39-43b3-9717-b46f71729b2a · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

Ola: Pushing the Frontiers of Omni-Modal Language Model Video instruction tuning with synthetic data, 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:47:40.126837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T22:47:39.389717Z digest=sha256:f1ef3f8431b084d2dabc331b344768495ecec4015c4506d3d72d644d4404e772

Pith citing papers

Observation 28d1de46-e4ed-4a6e-9d03-3056d57a8c90 · inbound

Compositional Generative Model of Unbounded 4D Cities cites this paper.

Compositional Generative Model of Unbounded 4D Cities Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-10T20:17:26.561510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:17:26.561510Z digest=sha256:b00863f0ef6fdc3a94261501bd84669e52fdab32fcab16d9d25a9869c3f8ecf0

Observation 72593941-1480-4b2f-8cae-5df4ae47d81a · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:34:36.856851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:745ef4ab3df2bfe1611b763a87073e759ac34b07d2f1b21b5f877f0f50f38481

Observation f4be1992-5bfe-4944-9db6-92afa7afd6d1 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:00:51.428555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:a1fecbf1a9ba8ff0cf334cec0990a8a9708a6d8f8ac1c733c003d16afcc2e3ee

Observation 79da6943-9a46-463a-9bf6-86ecd9b7955a · inbound

Is Extending Modality The Right Path Towards Omni-Modality? cites this paper.

Is Extending Modality The Right Path Towards Omni-Modality? Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:40.820699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:40.820699Z digest=sha256:427d59c5943c52227a8a3711bf37df11b37b50450ad9ae861fb6f21ba7d20d65

Observation be396407-bbd2-4d64-bb62-6ae72aab8163 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.533023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.533023Z digest=sha256:76056a4a4cbd406112fb25855c65986631df0573e50040184c5af988de4a4b06

Observation bf4ec20d-c45d-4fb2-90a6-085f09707196 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.562512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.562512Z digest=sha256:8d1183a100bd1b5701f497b0d3d531c402bb916fd1b8e0ca2f993895bfa56433

Observation 757897fa-3c40-4072-bca9-bb4530bc2848 · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.719441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.719441Z digest=sha256:042547e5a0cf73456e37d3e2ce1a73ede5e07abcc5b5d0366bef47b489bc2e6e

Observation 6a229ab0-3baa-4764-bed8-d65f74e078af · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:32.454824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:32.454824Z digest=sha256:db9fe38ae231da4ec5201d2383d0fafb0c8236fd23a23ef428e4381518ec9d26

Observation b21fc665-8563-4dac-84d1-1b4f15f5c7a2 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.250680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.250680Z digest=sha256:7dbd6f7eaf00e354a3426736241ae36f6b3060efd9ac921775a32455517762b4

Observation 891797d4-fed9-483a-87f6-fb8f24729b7f · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:12:07.011738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:96297588175f85fb952ddbd423ff8120ca77637eda06e63c635dcd816b868648

Observation 37c6c68c-dfa0-458c-aa6f-d438fb959ff9 · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:33.446078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:33.446078Z digest=sha256:7fe474b9cd719ce1a801e6943513df8ea0ffc1331681a271fc6e6d8ef1054e34

Observation 96ea2f47-d8aa-495b-8e66-fb699670c22b · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:07.832466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:07.832466Z digest=sha256:ae415d390a6fd0d77727f65251645ab59ca2f60ae1e1f5a5d7797dd890034f7c

Observation 93a144e0-8404-42a2-b380-4c3252808c8b · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.059602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.059602Z digest=sha256:390576ee3df34f8bfcee7494cbc31a457fe393c2a4d6d8e7a27310ec9937f68e

Observation 1b28e4ed-9481-4ad9-9b68-ffef3d04fb0f · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:40.028245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:40.028245Z digest=sha256:786d6411970dffb096150478cafef73e7b8ba0afb0583db8e374c539088b3869

Observation 52c170a9-6577-4502-9730-ac4b1ecd25c6 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.361938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.361938Z digest=sha256:c1317512a9669dacb7ad96e6cb259b423d7ede55210713613a9171ada0f37834

Observation feb94f13-ae33-40bc-a428-480a624c18e5 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.170122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:f0a075b06652457db8f2d0ca6f255f0cf5cf70dcadb77c49f4a48eb2571b73f3

Observation 2c6d5228-e79f-41a6-952e-3860e9100e21 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.227284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:da1e5a9f6c529215ee10aa84f4b3f2792d4c5bde16d50d883a0e5419975bc8ea

Observation 23330239-3260-46ab-bcb0-7c3a5a358b14 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 281

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.196437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:6ded044607cdde4bfccb31fa4e848ab1631a413d1df664d5771f71986d099e6f

Observation 1f256362-ffdb-4a6b-bfac-9af589abdda7 · inbound

Valley3: Scaling Omni Foundation Models for E-commerce cites this paper.

Valley3: Scaling Omni Foundation Models for E-commerce Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:05.976588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T14:53:55.160230Z digest=sha256:37dfe1fa33d87535f2e6f867254c239b11d1cfb3cc9607aedcf242f68e19844c

Observation bea1c93c-b9e3-4145-bf8c-0242fa2c665f · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:55.936975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:8847b5e435d7a5531b965fb470d1f3b4e576201bdab42b96634d9f53833d2620

Observation 576aaaa5-9692-4b5a-bcf5-e9d11540dc1b · inbound

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs cites this paper.

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:07:33.673297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T18:06:27.962891Z digest=sha256:0e021d754d0bd07b4a2d6aa4354478a452d49f134ff8c9f2923beaedce523500

Observation e310f04a-6116-4a6b-a3ad-cf50d7278c41 · inbound

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing cites this paper.

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:53:24.776257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:53:11.761843Z digest=sha256:2753d872ee119e92c3d1da4f58a97b982af49747f77c73821dea748983946741

Observation e046fc4b-1d3e-439a-8325-8d14bb80a7b7 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.901284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:dfe288b0b17d40a71ee9f7c17a88e43c3e36637e5652c34a7c3b92996c4fcdb1

Observation 3843e84b-bbca-4776-8646-153ed1d55dc8 · inbound

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain cites this paper.

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:23:17.992169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T10:20:12.860927Z digest=sha256:bec4bd40cc3962857d2b71e87fd99ae60768d2b429fd51b4b65c2847c18c809e

Observation 1ed0ee24-4c98-4e50-97f0-f325d22baf69 · inbound

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation cites this paper.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.177260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:91775bbcc81bf9ef94711df03b7ec5135bf82a75978fd5f3d845a4b7f28e78a3

Observation 65eca983-78dc-43d8-af42-ee4c9ea5a895 · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.391014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a3ab46b7a47fa7b5d4e077c0763fb3dabc98024db84041a834ef3303c1768448

Observation 0a3d6fd9-e5aa-4f55-851a-05d55517200a · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.365632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:27fac6ebda1fd8086a1460b77405b726b2a99af268a7bc6b7025e51b47fd64dd

Observation ba0ac130-32b2-431e-a8a1-0f689da9c32a · inbound

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning cites this paper.

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:03.433035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T09:43:56.299302Z digest=sha256:6249bdd9f0f53537ad59ca23b5e761710cbc94617c1b38443a0a69e933545b36

Observation 7c30ffa1-b114-4e1e-a1ce-e3adfb50705d · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.386746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:6d444e8d44284a7116b352a2c7944260f9f847d0fa0758f6133b8758c2622a58

Observation ae2729d0-56ea-4491-963f-4bd682948047 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.269489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:2df28ab932bea5fa1252002fc9d68821b21cad8522cedb3fd96e02168512bf5c

Observation 53305da3-9d94-4f90-9e0b-b2b6ea9cba5f · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.872965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:2f8176bda1bed1716b403fb3acefd764b4599c554653fdb5921b16cb5c36c4ae

Observation 367c2df0-13f8-499a-b7d2-0f7a77a9917a · inbound

Conversational Human Audio-visual Talking Dialogue Generation cites this paper.

Conversational Human Audio-visual Talking Dialogue Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T06:59:35.258176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:59:35.258176Z digest=sha256:bf467dce7897950ec829c306fff5a15b68aa946747fa5f72dff80ff330c54d72

Observation b27d6c40-0dd5-4988-85f7-3028d1c79fa7 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:49c86dbc482cd2d0827f5ac32e1516c9b85739d81c39e3000fb421d61ca61cf0

Observation 66981f95-7734-4775-aed4-d246dfee7d91 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:50.700260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:50.700260Z digest=sha256:a1342d2992904c148bad614c98c9978530df921da885b26f4b02237baa60d36a

Observation 94ddffae-aa2f-47ee-83fa-78e024b07ac4 · inbound

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs cites this paper.

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:04.401662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:04.401662Z digest=sha256:573280d87f82c5f9448ff99cf45400d445f08cb636adaf2bed1cd23805c95c75