Pith. sign in

Paper Citation Record · LEDGER

Ola: Pushing the Frontiers of Omni-Modal Language Model

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2502.04328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04328 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:03:07.832466Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:49:57.266999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 72593941-1480-4b2f-8cae-5df4ae47d81a · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:34:36.856851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:3f600879ad99e7b8c2ce68510f43207da3c3dcb49cb422d9b88483859dc81953

Observation f4be1992-5bfe-4944-9db6-92afa7afd6d1 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:00:51.428555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:126b2dc8ac2390fa8427107cc703a70d69010ba48fb8715ad0e486052f46a487

Observation 891797d4-fed9-483a-87f6-fb8f24729b7f · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:12:07.011738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:e7f48bf41445c53f1a4f90a4732b1cd9910c8b241310f70a07976a7308d5fe92

Observation 96ea2f47-d8aa-495b-8e66-fb699670c22b · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:07.832466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:07.832466Z digest=sha256:d78abeb53613e31c505356542dbe672b8751f33ca574b2b454cdc8dad965afda

Observation 93a144e0-8404-42a2-b380-4c3252808c8b · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.059602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.059602Z digest=sha256:5a768a31fce909c1e755aeba91af527cc9b36960f1f2aa984ce238c1e1f23c51

Observation 1b28e4ed-9481-4ad9-9b68-ffef3d04fb0f · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:40.028245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:40.028245Z digest=sha256:0e8e982b68bf58c7cb8639ea9d2e897963d32ff6a69acc7c670287fcf72cdeb7

Observation 52c170a9-6577-4502-9730-ac4b1ecd25c6 · inbound

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs cites this paper.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.361938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.361938Z digest=sha256:3d462b10edd6c094ceb6f3aa762c144c78e091d299350a5d0efb041e081fe7c5

Observation feb94f13-ae33-40bc-a428-480a624c18e5 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.170122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:b5754722c73c3a2df43eea9827a535dd61f44393e358429ee180ca10d5734060

Observation 2c6d5228-e79f-41a6-952e-3860e9100e21 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.227284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:75e9649ba5536acfdfe819587357b96244c1dbecd3fbc93396b480520b52003d

Observation 23330239-3260-46ab-bcb0-7c3a5a358b14 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 281

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.196437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:19fa2402927d3c94c9c719c9e77d826a508900bcad747efde0a5c24613ab5f4e

Observation 1f256362-ffdb-4a6b-bfac-9af589abdda7 · inbound

Valley3: Scaling Omni Foundation Models for E-commerce cites this paper.

Valley3: Scaling Omni Foundation Models for E-commerce Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:05.976588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T14:53:55.160230Z digest=sha256:42aff28afeb757c54ab68db5d9694e3cd561924ec65e3b2cf3a5170bde9febcc

Observation bea1c93c-b9e3-4145-bf8c-0242fa2c665f · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:55.936975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:43d1b4f743e7e1b10b9a86c09cc076e24c295eacc2c2dfcd236fb08d576dc3d9

Observation 576aaaa5-9692-4b5a-bcf5-e9d11540dc1b · inbound

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs cites this paper.

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:07:33.673297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T18:06:27.962891Z digest=sha256:ad907bd7546ab99fa0b67e98c7fe2c78edee425198ebe592d5e234c449fc81fe

Observation e310f04a-6116-4a6b-a3ad-cf50d7278c41 · inbound

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing cites this paper.

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:53:24.776257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T07:53:11.761843Z digest=sha256:bdee137334b9443496c22bf8a29a4a65a66467e533e61d60fa594d1edcd5a14d

Observation e046fc4b-1d3e-439a-8325-8d14bb80a7b7 · inbound

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding cites this paper.

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:55:24.901284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T05:51:49.390597Z digest=sha256:bc660050b7ff59fabc8c879b1c7d16de7defef107230ed9fe7a02f0619f26799

Observation 3843e84b-bbca-4776-8646-153ed1d55dc8 · inbound

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain cites this paper.

Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:23:17.992169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T10:20:12.860927Z digest=sha256:b96d89eb3fd6e2d4721274f8c0345c3adc2cb86a4db4e2b58bc7323e070254cd

Observation 1ed0ee24-4c98-4e50-97f0-f325d22baf69 · inbound

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation cites this paper.

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:26.177260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T12:58:56.555335Z digest=sha256:12d61356d055da197b46c7881bb2304b4e2f5a3d5a7502a69957cc26f11a6abf

Observation 65eca983-78dc-43d8-af42-ee4c9ea5a895 · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.391014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:32b671eaa3600cca6c2fdc656873f97695c910b44e68d89c836f9d5bd579bc96

Observation 0a3d6fd9-e5aa-4f55-851a-05d55517200a · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.365632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:73f0e689aaf9b23148fe9226cb4a1ddfbca02410cb5734af39443a224c458f9d

Observation ba0ac130-32b2-431e-a8a1-0f689da9c32a · inbound

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning cites this paper.

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:03.433035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T09:43:56.299302Z digest=sha256:ab22db789403e1a6420c221ee45e226c8340a7e2d280b615eb2aea841273eb99

Observation 7c30ffa1-b114-4e1e-a1ce-e3adfb50705d · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.386746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:20b073b83930ad1944269c66b847f49fbd77ea0baad788195d0ecc34ed00677f

Observation ae2729d0-56ea-4491-963f-4bd682948047 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.269489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:2b598143770a7b42811b98c38efed29debb81ef769f5637a71bd8d7f6688762d

Observation 53305da3-9d94-4f90-9e0b-b2b6ea9cba5f · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.872965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:f945f7ecfbdd7037572711e3017a1cd902021d425c3df19b85e9dd1b40ceb73c

Observation 367c2df0-13f8-499a-b7d2-0f7a77a9917a · inbound

Conversational Human Audio-visual Talking Dialogue Generation cites this paper.

Conversational Human Audio-visual Talking Dialogue Generation Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T06:59:35.258176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:59:35.258176Z digest=sha256:45cb6a4976972bdee6fd12b0032ebc7e40e1c969260dda720868148be7ff75ba

Observation b27d6c40-0dd5-4988-85f7-3028d1c79fa7 · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:8aa58169a0b640c44d862947d9eac2a7ff95591ffa415fc909c175994a3f5810

Observation 66981f95-7734-4775-aed4-d246dfee7d91 · inbound

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos cites this paper.

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-01T21:22:50.700260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:22:50.700260Z digest=sha256:abff49ba9a15be1cfe0ca93184ce812746ca8f1f835a423873753109df299325